<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Q. Hu);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Eficient Plagiarism Detection via Sentence Embeddings and FAISS-based Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>JiaCheng Tang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>QingBiao Hu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ZhongYuan Han</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Foshan University</institution>
          ,
          <addr-line>33 Guangyun Road, Shishan Town, Nanhai District, Foshan, Guangdong</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This work presents an eficient and scalable framework for detecting plagiarism in large document collections using sentence embedding models and fast approximate nearest neighbor search. Each document is segmented into overlapping chunks using a sliding window approach and encoded into dense semantic vectors using the pretrained intfloat/e5-base-v2 model. To accelerate semantic comparison, we first apply document-level ifltering using global embeddings, followed by chunk-level matching via FAISS for GPU-accelerated top-k retrieval. The proposed system significantly reduces runtime through embedding reuse and candidate pruning, while maintaining strong detection performance on large-scale benchmark datasets. It is fully modular, supports both CPU and GPU execution, and is compatible with the TIRA evaluation platform. Source code is publicly available at https://github.com/koppen777/plagiarism-detectio.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;PAN 2025</kwd>
        <kwd>Plagiarism Detection</kwd>
        <kwd>sentence embeddings</kwd>
        <kwd>FAISS</kwd>
        <kwd>sliding window</kwd>
        <kwd>top-k retrieval</kwd>
        <kwd>chunk matching</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Automatic plagiarism detection has become increasingly important in the era of large-scale digital
content and machine-generated text. The PAN shared task provides a benchmark for evaluating systems
on challenging plagiarism scenarios involving paraphrasing, translation, and AI-generated rewriting.
Given a suspicious document and a pool of source documents, the goal is to accurately locate plagiarized
passages and align them with their original sources.</p>
      <p>While traditional approaches often rely on lexical overlap or character-based similarity, such methods
are limited when facing paraphrased or semantically modified text. Recent work has explored the use
of sentence embeddings to capture semantic similarity; however, many of these approaches are either
computationally expensive or require exhaustive pairwise comparisons. We observed that models like
all-MiniLM-L6-v2 were not robust enough for this task, producing low recall. These limitations call for
a more eficient and scalable solution.</p>
      <p>To address this, we propose a lightweight and eficient two-stage plagiarism detection system that
combines sentence embeddings with FAISS-based retrieval. We segment suspicious and source documents
into overlapping chunks using a sliding window strategy and embed them using the intfloat/e5-base-v2
model, which we found to outperform MiniLM in this task. Document-level embeddings are used to
iflter unrelated sources, followed by chunk-level approximate nearest neighbor search via FAISS to
detect plagiarism. We further optimize performance by tuning key parameters such as window size,
stride, similarity threshold, and the number of top-K candidates. Our system is compatible with TIRA
and supports both CPU and GPU execution, making it suitable for large-scale evaluation environments
such as Kaggle. This work was carried out by the Foshan University Artificial Intelligence Laboratory,
which focuses on natural language processing and computer vision research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The PAN plagiarism detection shared task has evolved into a standard benchmark for evaluating text
reuse systems under realistic and adversarial scenarios. In recent years, top-performing systems have
shifted from lexical fingerprinting methods toward semantic-aware techniques. For instance, the
topranked system in PAN 2023 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] employed cross-encoder transformers to directly classify sentence pairs,
achieving strong performance but at the cost of high computational overhead. In contrast, the runner-up
system [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] adopted a retrieval-based approach using MiniLM embeddings and FAISS indexing to balance
eficiency and efectiveness.
      </p>
      <p>
        The 2025 edition of the PAN shared task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] introduces new challenges in generative plagiarism
detection, supported by a large-scale manually annotated dataset and a unified evaluation platform. The
subtask overview [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] highlights the dificulty of detecting semantically rewritten AI-generated content.
Evaluation is conducted using the TIRA platform [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which provides a reproducible and containerized
benchmarking infrastructure used across PAN tasks.
      </p>
      <p>
        Broadly, existing methods for semantic plagiarism detection fall into three categories. The first category
uses traditional lexical features (e.g., n-gram overlap, edit distance), which are fast but vulnerable to
paraphrasing. The second category leverages supervised classifiers (e.g., BERT or RoBERTa fine-tuned
on similarity datasets), which ofer strong performance but require labeled training data and are slow
at inference. The third category, which our method belongs to, applies sentence embedding models
such as Sentence-BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] or E5 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to encode textual chunks, enabling eficient approximate matching
via ANN methods like FAISS [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>A major limitation of many prior systems is their reliance on exhaustive pairwise comparison
between all suspicious and source chunks, which becomes infeasible at scale. Furthermore, many public
implementations either neglect document-level filtering or use unoptimized thresholds, leading to
suboptimal trade-ofs between precision and recall. Our work addresses these issues by incorporating
document-level embedding filtering, parameter tuning, and embedding reuse to maximize detection
quality while minimizing runtime.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>Our plagiarism detection system adopts a two-stage semantic retrieval pipeline, as shown in Figure 1.
The system is designed to eficiently identify semantic reuse across a large number of document pairs
by leveraging sentence embeddings and FAISS-based approximate nearest neighbor (ANN) search.
The method consists of the following key components:</p>
      <sec id="sec-3-1">
        <title>3.1. Document Preprocessing and Chunking</title>
        <p>Each suspicious and source document is preprocessed through sentence splitting using regular
expressions. Specifically, we apply a heuristic rule that splits the text at punctuation followed by
whitespace (e.g., period, exclamation mark, or question mark), which corresponds to the regular expression
(?&lt;=[.!?])\s+. This lightweight rule performs reasonably well on English text and avoids the
overhead of full syntactic parsing.</p>
        <p>After sentence segmentation, we apply a sliding window mechanism to generate overlapping textual
chunks of  sentences (we set  = 6 with stride 2), which are used as the basic comparison units in the
downstream embedding and retrieval process.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Embedding with Pretrained Sentence Transformers</title>
        <p>We encode each chunk using a pretrained Sentence Transformer model. After evaluating
several candidates, we selected intfloat/e5-base-v2 as it provides stronger performance than
all-MiniLM-L6-v2 in semantic detection tasks. Each document also has a full-text embedding
to support coarse filtering.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Top-K Candidate Source Filtering</title>
        <p>To reduce the number of comparisons, we compute a global embedding for each suspicious document
and retrieve its Top-K most similar source documents based on cosine similarity. This document-level
ifltering reduces the chunk comparison space from thousands to just a dozen documents per query.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Chunk-Level Semantic Matching with FAISS</title>
        <p>
          We use FAISS [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] to perform approximate nearest neighbor search between suspicious and source chunks.
Both query and index vectors are normalized, and the inner product metric is used to approximate
cosine similarity. We retain the top 5 matches per chunk and apply a similarity threshold (e.g., 0.83) to
select candidate matches.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Span Aggregation and XML Output</title>
        <p>Matched chunks are mapped back to character ofsets using sentence-level ofsets. Overlapping or
adjacent segments are merged if they are within a certain character distance (e.g., 30 chars). The final
predicted plagiarized spans are then saved in the XML format required by the PAN evaluation system.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment</title>
      <sec id="sec-4-1">
        <title>4.1. Experimental Setup</title>
        <p>We evaluated our system on the pan25-generated-plagiarism-detection-validation dataset
released by the PAN 2025 shared task organizers. The dataset contains suspicious documents and source
documents, with manually annotated plagiarism cases provided in XML format.</p>
        <p>Each document is preprocessed using sentence tokenization based on regular expressions. We use a
sliding window of 6 sentences with a stride of 2 to generate overlapping textual chunks. Sentence
embeddings are computed using the intfloat/e5-base-v2 model, chosen for its strong semantic
retrieval performance. FAISS is configured with inner product indexing (normalized vectors), and top-5
chunk matches are retrieved for each suspicious chunk. To reduce search space, we first compute a
document-level embedding and filter the top-5 candidate source documents.</p>
        <p>All experiments are conducted on the Kaggle platform using GPU for embedding and CPU for FAISS
retrieval. Embeddings are cached and reused to reduce runtime overhead. Table 1 summarizes the key
hyperparameters.
Evaluation is performed using the oficial PAN evaluation script, which computes precision, recall, and
F1-score based on the overlap of predicted and ground-truth spans. All XML outputs follow the PAN
format and are submitted through TIRA for validation.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Results and Analysis</title>
        <p>We observed that increasing the window size and lowering the similarity threshold improved recall
without sacrificing much precision. Document-level filtering using global embeddings significantly
reduced runtime, from 30+ hours to under 3 hours on the same hardware. Compared to a naive baseline
with no filtering and smaller embedding model, our approach ofers better scalability and accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We presented an eficient two-stage plagiarism detection system based on sentence embeddings and
FAISS-based retrieval. By combining document-level filtering and chunk-level approximate matching,
our method achieves a strong balance between accuracy and computational eficiency. The use of
pretrained sentence encoders such as intfloat/e5-base-v2, along with optimized hyperparameters
and embedding reuse, significantly improves detection quality while reducing runtime. Our system
is fully compatible with the TIRA evaluation platform and can be deployed on resource-constrained
environments such as Kaggle. Our implementation code is publicly available at https://github.com/
koppen777/plagiarism-detection to foster reproducibility and further research in plagiarism detection.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>
        This work is supported by the Social Science Foundation of Guangdong Province, China (No.
GD24CZY02), and by the organizers of the PAN 2025 shared task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the CLEF initiative. We
also thank the TIRA team [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] for providing the evaluation infrastructure and helpful documentation.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used ChatGPT (GPT-4) in order to: grammar and
language polishing. After using this tool, the author(s) reviewed and edited the content as needed and
take full responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Althobaiti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <source>Plagiarism detection at pan</source>
          <year>2023</year>
          , in: CLEF 2023 Labs,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Liu,
          <article-title>Eficient cross-language plagiarism detection with minilm</article-title>
          ,
          <source>in: Working Notes of CLEF</source>
          <year>2023</year>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Greiner-Petter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2025:
          <article-title>Voight-Kampf Generative AI Detection, Multilingual Text Detoxification, Multi-Author Writing Style Analysis, and Generative Plagiarism Detection</article-title>
          , in: J.
          <string-name>
            <surname>C. de Albornoz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piroi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Sixteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Greiner-Petter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Wahle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ruas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aizawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <source>Overview of the Generative Plagiarism Detection Task at PAN</source>
          <year>2025</year>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , D. Spina (Eds.),
          <source>Working Notes of CLEF 2025 - Conference and Labs of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kolyada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Grahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elstner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Loebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <article-title>Continuous Integration for Reproducible Shared Tasks with TIRA.io</article-title>
          ,
          <source>in: Advances in Information Retrieval. 45th European Conference on IR Research (ECIR</source>
          <year>2023</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2023</year>
          , pp.
          <fpage>236</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          ,
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , et al.,
          <article-title>Text embeddings by weakly-supervised contrastive pre-training</article-title>
          ,
          <source>arXiv preprint arXiv:2212.09741</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , M. Douze,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jégou</surname>
          </string-name>
          ,
          <article-title>Billion-scale similarity search with gpus</article-title>
          ,
          <year>2019</year>
          . ArXiv preprint arXiv:
          <volume>1702</volume>
          .
          <fpage>08734</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kolyada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Grahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elstner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Loebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <article-title>Continuous Integration for Reproducible Shared Tasks with TIRA.io</article-title>
          , in: J.
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Maistro</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Kruschwitz</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Caputo (Eds.),
          <source>Advances in Information Retrieval. 45th European Conference on IR Research (ECIR</source>
          <year>2023</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2023</year>
          , pp.
          <fpage>236</fpage>
          -
          <lpage>241</lpage>
          . URL: https://link. springer.com/chapter/10.1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>20</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -28241-6_
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>