<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A. R. Basani);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>DivEye at PAN 2025: Diversity Boosts AI-Generated Text Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Advik Raj Basani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pin-Yu Chen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Birla Institute of Technology and Science, KK Birla Goa Campus</institution>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IBM Research</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Detecting AI-generated text is increasingly important to prevent misuse in education, journalism, and social media, where synthetic fluency can obscure misinformation. This paper presents our solution for the Generative AI Authorship Verification Task at PAN 2025, where the objective is to distinguish machine-generated text from human-written content. We propose DivEye, a novel detection framework that leverages surprisal-based features to capture fluctuations in lexical and structural unpredictability, a signal more prominent in human-authored text. Our method performs competitively across diverse text domains and models, especially on challenging cases where model-generated text closely resembles human writing, and also outperforms the four oficial baselines of the PAN 2025 task.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;PAN 2025</kwd>
        <kwd>Voight-Kampf AI Detection Sensitivity</kwd>
        <kwd>Generative AI Authorship Verification Task</kwd>
        <kwd>llms</kwd>
        <kwd>ai text detection</kwd>
        <kwd>interpretability</kwd>
        <kwd>zero-shot</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Large Language Models (LLMs) are widely used in tasks from personal assistance to content creation
[
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">1, 2, 3, 4, 5</xref>
        ]. While their fluency enhances utility, it also enables seamless insertion of AI-generated
text into essays, articles, legal briefs, and social media, often without detection [
        <xref ref-type="bibr" rid="ref6 ref7 ref8 ref9">6, 7, 8, 9</xref>
        ].
      </p>
      <p>Reliable AI-text detection is vital for combating risks like misinformation, academic dishonesty,
professional misconduct, and the suppression of genuine human writing [10, 11, 12]. Traditional
supervised detectors [13, 14, 15] rely on labeled datasets but often fail to generalize to unseen models
or domains [16, 11], especially as new LLMs emerge. Zero-shot detectors [17, 18, 19, 20] address this
by leveraging statistical signals or LLMs at inference time, ofering scalable, model-agnostic detection
critical for maintaining platform integrity.</p>
      <p>Contributions. In this work, we present DivEye, a zero-shot framework for AI-generated text
detection submitted to the PAN@CLEF 2025 Generative AI Authorship Verification task [ ? ? ]. The
challenge centers on identifying the human-written text when presented with a pair comprising one
human and one machine-authored sample. Our method leverages diversity-based statistical features
computed over token-level surprisal [21] sequences from a reference language model. Unlike
finetuned classifiers or signature-based detectors, DivEye captures distributional irregularities inherent
in AI-generated content by measuring surprisal variance, entropy, and other diversity features. These
metrics are grounded in linguistic theory and require no access to a specific text-generation LM. DivEye
is model-agnostic, scalable, and can operate without retraining, making it suitable for real-world
deployment. Notably, it complements existing detectors by revealing statistical signals often missed by
black-box or fine-tuned approaches. Our results show strong generalization across domains and model
families, achieving competitive performance in this challenging verification setting.</p>
      <p>Computation of Surprisal</p>
      <p>via a Frozen LLM</p>
      <p>LLM
Input Text</p>
      <p>Frozen LLM</p>
      <p>Extraction of Surprisal per Token</p>
      <p>Split into tokens &amp;
compute per token</p>
      <p>DivEye Feature</p>
      <p>Extraction
Distribution Features
Mean, Variance, Skew, Kurtosis
First-Order Features
Second-Order Features
Variance-II, Entropy, Autocorrelation</p>
      <p>External AI-Detection
Framework (if any)</p>
      <p>Feature
Concatenation</p>
      <p>Feature Concatenation
+ Classification
Binary Classifier</p>
      <p>(XGBoost)
Framework to be</p>
      <p>Boosted</p>
      <p>Human Text</p>
      <p>AI Text</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background &amp; Preliminaries</title>
      <p>The rise of LLMs has enabled machine-generated text that closely mimics human writing by
approximating the true conditional distribution of natural language, human( | &lt;), through training on
large human-written corpora [22, 23]. The LLM’s learned distribution, LLM( | &lt;), is used to
sequentially generate tokens during inference via sampling [24]. Despite their fluency, LLMs imperfectly
approximate human language (LLM ̸= human) [25, 26], and this subtle diference is the crux of AI text
detection.</p>
      <p>Existing detection methods fall into three categories: watermarking, supervised / fine-tuned and
zero-resource detection. Watermarking [27, 28, 29] embeds patterns in generated text but requires
model access or fine-tuning, limiting use in black-box or adversarial settings. Zero-resource methods
need no model knowledge and rely on statistical or learned diferences between human and AI text,
further divided into statistical and training-based approaches.</p>
      <p>Supervised / Fine-tuned detection methods [30, 31, 32] train classifiers, such as fine-tuned
transformers on a labeled corpora of human and AI text. While these models can be accurate, they
often fail to generalize across domains or against adversarial paraphrasing, especially when trained
on specific generators or prompts. Statistical / Zero-shot detection methods refers to identifying
AI-generated text without task-specific training, either by leveraging LLM probability cues or prompting
LLMs directly as detectors. For example, methods like Entropy [33], LogRank [34], DetectGPT [17, 19],
and Binoculars [35] use of-the-shelf LLMs to evaluate the consistency of token predictions under
masked or perturbed inputs.</p>
      <p>Despite progress, AI-text detection remains unsolved. We move beyond individual token probabilities
to measure statistical diversity across token sequences, capturing variation in surprise and predictability.</p>
    </sec>
    <sec id="sec-3">
      <title>3. DivEye: Methodologies</title>
      <sec id="sec-3-1">
        <title>3.1. Design Hypothesis</title>
        <p>One of the main challenges in detecting AI-generated text [34, 36] lies in the fact that while modern
LLMs excel at generating fluent and coherent text, they often do so at the expense of variability and
diversity [37].</p>
        <p>We hypothesize that human-authored text naturally displays greater stylistic diversity and
unpredictability than text produced by AI. Human writing tends to include creative and impulsive
choices that introduce unexpected shifts, whereas large language models prioritize high-probability
sequences [38], resulting in more uniform and predictable outputs. This hypothesis is supported by
both intuitive reasoning and empirical findings (see Remark 1).</p>
        <p>Remark 1: Proof Sketch
Consider a text sequence  = (1, 2, . . . , ) generated either by a human or by a language model
 . The language model defines a probability distribution  () = ∏︀
=1  ( | &lt;) where
each token is chosen to maximize overall likelihood.</p>
        <p>Humans, however, produce language through a complex, multi-layered cognitive process that
balances informativeness, creativity, and contextual appropriateness, rather than strictly maximizing
statistical likelihood. Formally, the surprisal of token  under model  is defined as:
 () = − log  ( | &lt;)
Since  is trained to assign high probability to plausible continuations, its outputs tend to minimize
surprisal on average, implying that maximum likelihood generation compresses diversity:</p>
        <p>E∼  [ ()] ≤ E∼  [ ()]
where  denotes the distribution of human-generated text.</p>
        <p>Similarly, human language exhibits higher variance in surprisal due to spontaneous creative choices,
idiomatic expressions, and stylistic variation, causing:</p>
        <p>Var∼  [ ()] &lt; Var∼  [ ()]
We validate this theoretical intuition through empirical experiments detailed below, which confirm
statistically significant diferences in surprisal and diversity metrics between human-written and
AI-generated texts.</p>
        <p>We empirically validate these theoretical claims using a dataset of 200 human-written essays and
200 GPT-4-Turbo-generated essays on matched topics, sourced from BiScope [39]. For each essay,
we compute token-level surprisal scores using a fixed language model evaluator (GPT-2), then
calculate the mean and variance of surprisal per essay. Figure 2a presents the distribution of mean
surprisal scores across both groups, while Figure 2b shows the corresponding variance distributions.
Human-written texts demonstrate a broader spread and heavier tails in both metrics, indicating
greater unpredictability and stylistic richness. In contrast, AI-generated texts are more tightly
clustered with lower mean surprisal and significantly reduced variance. These findings empirically
corroborate our hypothesis: human language inherently reflects higher diversity and surprise,
whereas AI-generated language, optimized for likelihood, tends toward more predictable
and homogeneous patterns.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Foundations of DivEye</title>
        <p>DivEye computes higher-order statistical features over surprisal sequences, enabling the capture of
structural patterns that go beyond aggregate likelihood. More detailed theoretical foundations and
experimental results are presented in the original DivEye paper [40].</p>
        <p>Surprisal. Human language balances consistency with creative bursts, introducing novel expressions
and stylistic variation. This diversity can be quantified using surprisal [ 41], the negative log-probability
of a token given its context () = − log  ( | 1, 2, . . . , − 1). For a sequence  = 1, . . . , ,
surprisal ofers a principled measure of local unpredictability based on model log-probabilities.</p>
        <p>Rather than examining individual token surprisals in isolation, we summarize their behavior through
aggregate metrics. The mean surprisal ( ) serves as a coarse indicator of how “expected” a text is on
average: Lower values suggest closer conformity to the model’s distribution, whereas higher values
signal greater unpredictability. Moreover, human writing also exhibits fluctuations in predictability due
to stylistic shifts, topic changes, or bursts of creativity, motivating the use of surprisal variance ( 2)
2.0
2.5
3.0
3.5
4.0
5
6
7
8
9
10
11
12
Mean Surprisal
Surprisal Variance
(a)</p>
        <sec id="sec-3-2-1">
          <title>Mean Surprisal Distribution</title>
          <p>(b)</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Surprisal Variance Distribution</title>
          <p>1 ∑︁(() −   )2
Mean and Variance are not suficient.</p>
          <p>Mean and variance capture surprisal’s central tendency and
spread but miss deeper structural signals distinguishing human from AI text. Human writing often shows
asymmetric surprisal distributions with bursts of creativity, causing occasional spikes in unpredictability.
AI-generated text, optimized for consistency, tends toward more symmetrical distributions centered
on high-probability tokens [25]. Skewness ( 1) measures this asymmetry, positive values indicate
rare, surprising tokens typical of human writing, while kurtosis ( 2) reflects the frequency of extreme
deviations, signaling stylistic diversity. These higher-order moments enable DivEye to detect subtle
irregularities overlooked by methods focusing only on average behavior.</p>
          <p>1 =
1 ∑︁ (︂ () −   )︂ 3
 =1
 
;  2 =
1 ∑︁ (︂ () −   )︂ 4
 =1
 
abrupt shifts in topic or tone common in human writing.</p>
          <p>Static metrics still miss temporal structure.</p>
          <p>While static surprisal metrics (mean, variance,
skewness, kurtosis) summarize overall unpredictability, they miss how it evolves across a sequence, a
key trait separating human from AI text. To model these dynamics, we compute the first-order diference
Δ = () − (− 1), with its mean (Δ ) and variance (Δ 2) capturing stylistic volatility, such as
(
rate of surprisal change. From this, we extract: (1) variance ( Δ22 ) for erratic transitions; (2) entropy
ℋΔ2 ) for irregularity; and (3) autocorrelation ( (Δ2)) for clustering of unpredictability bursts. These
metrics uncover rhythmic, non-stationary patterns typical of human text but rare in the smoother, more
uniform outputs of LLMs, ofering a richer signal for detection. These have been formally defined as:
We also compute the second-order diference
Δ2 = Δ −
Δ− 1 to track fluctuations in the
Δ = () − (− 1),
Δ2 = Δ −
Δ− 1,
Δ =
 Δ22 =
 (Δ2) =
1
1
 − 1
 − 2

=2

=3
∑︁ Δ,</p>
          <p>Δ 2 =
∑︁(Δ2 −  Δ2 )2,</p>
          <p>ℋΔ2 = −
E ︀[ (Δ2 −  Δ2 )(Δ2+1 −  Δ2 )</p>
          <p>︀]
∑︁(Δ −  Δ)</p>
          <p>2
∑︁  log ,

(1)
(2)
(3)
(4)
(5)
where  Δ2 is the mean of second-order diferences, and  is the empirical probability of a value
falling into bin  after discretizing Δ2 for entropy computation. We provide empirical validation of
these temporal features and their individual contributions to detection performance in Appendix A.</p>
          <p>Combinations. Collectively, DivEye, formalized as () in Equation (6), encapsulates critical aspects
of text generation that distinguish human creativity from algorithmically generated predictability,
thereby serving as a robust basis for our detection framework.
 is a 9-dimensional vector of distributional, first-order, and second-order statistics, derived by passing
text through an autoregressive LLM. These features feed a binary classifier, optionally combined with
existing detector outputs. See Algorithm 1 for details.</p>
          <p>DivEye as a booster. Existing detectors often fail against high-quality adversarial text that mimics
human writing. DivEye provides a complementary signal by capturing statistical and temporal patterns
of token-level unpredictability, orthogonal to traditional features. We enhance detectors by appending
DivEye ’s feature vector to their outputs and training a lightweight meta-classifier (e.g., XGBoost [ 42],
Random Forest [43]) on the combined representation. This fusion significantly improves performance
on adversarial and out-of-distribution text, without retraining or altering the base model.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>
        Datasets. We evaluate our zero-shot DivEye framework on a diverse suite of datasets that span a
wide range of generative models, domains, and adversarial strategies. To substantiate the methodology
described above1, our primary benchmark is the MAGE benchmark [44]. MAGE comprises eight distinct
testbeds covering multiple domains (e.g., Yelp [45], XSum [46], SciXGen [
        <xref ref-type="bibr" rid="ref10">47</xref>
        ], CMV [
        <xref ref-type="bibr" rid="ref11">48</xref>
        ]) and a range of
text generator families (e.g., GPT [
        <xref ref-type="bibr" rid="ref12">49</xref>
        ], OPT [
        <xref ref-type="bibr" rid="ref13">50</xref>
        ], Bloom [
        <xref ref-type="bibr" rid="ref14">51</xref>
        ]).
      </p>
      <p>This fine-grained evaluation setup enables us to isolate and analyze the contribution of
diversitybased metrics across diferent domains and model architectures. Each testbed includes predefined
training and evaluation splits, which we use accordingly. For full implementation details and extended
experimental results on MAGE, we request the reader to refer to our original paper [40], where all
MAGE-related experiments are presented and discussed in depth.</p>
      <p>
        PAN Dataset. The PAN@CLEF 2025 Generative AI Author Verification Task [
        <xref ref-type="bibr" rid="ref15 ref16">52, 53</xref>
        ] provides a
dataset comprising both human-authored and machine-generated texts. We utilize the training and
supplementary evaluation splits from this dataset to train a binary classifier enhanced with features
from DivEye. All of our submissions are exclusively trained on this provided dataset.
      </p>
      <p>
        Implementation Details &amp; Metrics. Unless stated otherwise, we use GPT-2 to compute all DivEye
feature vectors. In score-only detection scenarios, predictions are based solely over concatenated
DivEye features. For both standalone and boosted setups, we train a lightweight XGBoost [42] classifier
as a meta-model, using only DivEye features in the former, and concatenating them with the original
detector’s prediction scores in the latter. We use an XGBoost classifier for binary classification as a
preliminary choice, without extensive comparison to other classifiers, leaving exploration of alternative
models for future work. We evaluate our method using the oficial PAN@CLEF 2025 evaluation platform
(TIRA [
        <xref ref-type="bibr" rid="ref17">54</xref>
        ]), which reports the following metrics:
• AUROC: The conventional Area Under the Receiver Operating Characteristic Curve.
• c@1: A metric that rewards systems for leaving uncertain cases unanswered.
• 0.5: A variation of the F-score that emphasizes correctly identifying same-author cases.
• F1-score: The harmonic mean of precision and recall, capturing balanced model performance.
1Note: These models were not submitted for any tasks in PAN 2025; results are reported solely to empirically validate our
approach.
      </p>
      <p>• Brier Score: Measures the accuracy of probabilistic predictions by computing the mean squared
error between predicted probabilities and true labels.</p>
      <p>Baselines. We compare DivEye against a diverse set of baselines under two evaluation settings. As
detailed in our original paper [40], for the MAGE benchmark, we evaluate both traditional statistical
detectors and recent fine-tuned models, including RADAR [ 32], FastDetectGPT [17], Binoculars [35],
and BiScope [39].</p>
      <p>Although we do not explicitly report quantitative results in this manuscript for MAGE, we kindly
refer readers to our original paper [40] for full empirical comparisons. Nevertheless, we summarize and
discuss the key findings and core observations here to provide insight into the comparative performance
of DivEye.</p>
      <p>
        For the PAN@CLEF 2025 task, we follow the oficial evaluation protocol and compare against the
provided baselines: Linear SVM with TF-IDF features, Binoculars [35], and a PPMd compression-based
cosine similarity method [
        <xref ref-type="bibr" rid="ref18">55</xref>
        ]. These lightweight, model-agnostic baselines highlight the advantage of
incorporating statistical diversity features even in constrained, zero-shot scenarios.
4.1. DivEye in PAN 2025
To evaluate the efectiveness of DivEye in the PAN 2025 authorship verification task, we follow the
oficial protocol and train exclusively on the dataset provided by the organizers. Given the consistent
performance gains observed with DivEye across diverse settings, we focus our submission on the
enhanced variant DivEye + BiScope [39], which demonstrated superior results in earlier evaluations.
This combination leverages the complementary strengths of BiScope’s decision boundary with DivEye ’s
diversity-based signal, leading to improved robustness and generalization across authorship verification
cases. The final results are summarized in Table 1.
      </p>
      <sec id="sec-4-1">
        <title>4.2. Robustness, Eficiency &amp; Boosting Efectiveness of</title>
      </sec>
      <sec id="sec-4-2">
        <title>DivEye</title>
        <p>We evaluate DivEye across a wide range of challenging testbeds to assess its robustness and adaptability
under both domain and model distribution shifts. Our experiments span multiple testbeds from the
MAGE benchmark [44], including both in-distribution and out-of-distribution scenarios. Across
all settings, DivEye consistently outperforms existing zero-shot and fine-tuned baselines in terms
of AUROC and average accuracy, demonstrating strong generalization to both familiar and novel
generation patterns. In our original paper [40], we report detailed performance metrics for DivEye
across these testbeds. The results highlight high AUROCs (e.g., 0.98 and 0.93 across domains and
generator families), along with strong accuracy, underscoring the stability and robustness of our
approach across diverse evaluation conditions.</p>
        <p>To further assess robustness, we evaluate DivEye under adversarial conditions such as paraphrasing
attacks. Even in these challenging scenarios, DivEye outperforms strong fine-tuned baselines by
notable margins in both AUROC and average accuracy. In addition to accuracy, DivEye is highly
eficient, processing each input in approximately 0.01 seconds due to its lightweight GPT-2 backbone
and fast statistical feature computations. This makes it particularly well-suited for deployment in
real-time or resource-constrained environments.</p>
        <p>Finally, we demonstrate that DivEye’s diversity-based surprisal features substantially enhance the
performance of existing detectors when used in combination. Fusing DivEye with other diverse
detectors results in AUROC and accuracy gains exceeding 18.7%, showing that these features ofer
complementary signals to traditional methods. For full experimental results, including quantitative
comparisons and boosting analyses, we refer reviewers to our original paper [40].</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>We successfully participated in the PAN@CLEF2025 Generative AI Authorship Verification task using
our proposed framework, DivEye, which leverages surprisal diversity for robust zero-shot detection.
By integrating DivEye with BiScope, we achieved strong performance in distinguishing human-written
from machine-generated texts, overperforming all given baselines. Our method shows high adaptability
across domains and paraphrased inputs, indicating its efectiveness in real-world authorship verification
scenarios.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We would like to thank the Data, Systems and High Performance Computing (DaSH) Lab2 and the PI,
Prof. Arnab K. Paul, for providing the computational resources necessary to conduct our experiments.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>The authors have not employed any Generative AI tools.
[10] S. Abdali, R. Anarfi, C. Barberan, J. He, Decoding the ai pen: Techniques and challenges in detecting
ai-generated text, 2024. URL: https://arxiv.org/abs/2403.05750. doi:https://doi.org/10.1145/
3637528.3671463. arXiv:2403.05750.
[11] H. D. S. Gameiro, A. Kucharavy, L. Dolamic, Llm detectors still fall short of real world:
Case of llm-generated short news-like posts, 2024. URL: https://arxiv.org/abs/2409.03291.
arXiv:2409.03291.
[12] J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, D. F. Wong, A survey on llm-generated text detection:
Necessity, methods, and future directions, Comput. Linguistics 51 (2025) 275–338. URL: https:
//doi.org/10.1162/coli_a_00549. doi:10.1162/COLI\_A\_00549.
[13] S. M. Shukla, C. Magoo, P. Garg, Comparing fine tuned-lms for detecting llm-generated text, in:
2024 3rd Edition of IEEE Delhi Section Flagship Conference (DELCON), IEEE, 2024, pp. 1–8.
[14] I. Tolstykh, A. Tsybina, S. Yakubson, A. Gordeev, V. Dokholyan, M. Kuprashevich, Gigacheck:</p>
      <p>Detecting llm-generated content, arXiv preprint arXiv:2410.23728 (2024).
[15] R. Wang, H. Chen, R. Zhou, H. Ma, Y. Duan, Y. Kang, S. Yang, B. Fan, T. Tan, Llm-detector:
Improving ai-generated chinese text detection with open-source llm instruction tuning, arXiv
preprint arXiv:2402.01158 (2024).
[16] J. Doughman, O. M. Afzal, H. O. Toyin, S. Shehata, P. Nakov, Z. Talat, Exploring the limitations of
detecting machine-generated text, 2024. URL: https://arxiv.org/abs/2406.11073. arXiv:2406.11073.
[17] G. Bao, Y. Zhao, Z. Teng, L. Yang, Y. Zhang, Fast-detectgpt: Eficient zero-shot detection of
machinegenerated text via conditional probability curvature, 2024. URL: https://arxiv.org/abs/2310.05130.
arXiv:2310.05130.
[18] S. Gehrmann, H. Strobelt, A. M. Rush, Gltr: Statistical detection and visualization of generated
text, 2019. URL: https://arxiv.org/abs/1906.04043. arXiv:1906.04043.
[19] E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, C. Finn, Detectgpt: Zero-shot
machinegenerated text detection using probability curvature, 2023. URL: https://arxiv.org/abs/2301.11305.
arXiv:2301.11305.
[20] H. Wang, X. Luo, W. Wang, X. Yan, Bot or human? detecting chatgpt imposters with a single
question, 2024. URL: https://arxiv.org/abs/2305.06424. arXiv:2305.06424.
[21] E. G. Wilcox, T. Pimentel, C. Meister, R. Cotterell, R. P. Levy, Testing the predictions of surprisal
theory in 11 languages, 2025. URL: https://arxiv.org/abs/2307.03667. arXiv:2307.03667.
[22] B. Chen, X. Wang, S. Peng, R. Litschko, A. Korhonen, B. Plank, “seeing the big through the
small”: Can LLMs approximate human judgment distributions on NLI from a few explanations?,
in: Y. Al-Onaizan, M. Bansal, Y.-N. Chen (Eds.), Findings of the Association for Computational
Linguistics: EMNLP 2024, Association for Computational Linguistics, Miami, Florida, USA, 2024, pp.
14396–14419. URL: https://aclanthology.org/2024.findings-emnlp.842/. doi: 10.18653/v1/2024.
findings-emnlp.842.
[23] Y. Lu, J. Huang, Y. Han, B. Bei, Y. Xie, D. Wang, J. Wang, Q. He, Llm agents that act like us: Accurate
human behavior simulation with real-world data, 2025. URL: https://arxiv.org/abs/2503.20749.
arXiv:2503.20749.
[24] Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y. Lou, L. Wang, Z. Yuan, X. Li, S. Yan, G. Dai, X.-P.</p>
      <p>Zhang, Y. Dong, Y. Wang, A survey on eficient inference for large language models, 2024. URL:
https://arxiv.org/abs/2404.14294. arXiv:2404.14294.
[25] D. Ippolito, D. Duckworth, C. Callison-Burch, D. Eck, Automatic detection of generated text is
easiest when humans are fooled, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of
the 58th Annual Meeting of the Association for Computational Linguistics, Association for
Computational Linguistics, Online, 2020, pp. 1808–1822. URL: https://aclanthology.org/2020.acl-main.164/.
doi:10.18653/v1/2020.acl-main.164.
[26] C. R. Jones, S. Trott, B. Bergen, Comparing humans and large language models on an experimental
protocol inventory for theory of mind evaluation (epitome), Transactions of the Association for
Computational Linguistics 12 (2024) 803–819. URL: https://doi.org/10.1162/tacl_a_00674. doi:10.
1162/tacl_a_00674.
[27] J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, T. Goldstein, A watermark for large language
models, 2024. URL: https://arxiv.org/abs/2301.10226. arXiv:2301.10226.
[28] Y. Liang, J. Xiao, W. Gan, P. S. Yu, Watermarking techniques for large language models: A survey,
2024. URL: https://arxiv.org/abs/2409.00089. arXiv:2409.00089.
[29] A. Liu, L. Pan, Y. Lu, J. Li, X. Hu, X. Zhang, L. Wen, I. King, H. Xiong, P. S. Yu, A survey of text
watermarking in the era of large language models, 2024. URL: https://arxiv.org/abs/2312.07913.
arXiv:2312.07913.
[30] Y. Chen, H. Kang, V. Zhai, L. Li, R. Singh, B. Raj, Token prediction as implicit classification to
identify LLM-generated text, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023
Conference on Empirical Methods in Natural Language Processing, Association for Computational
Linguistics, Singapore, 2023, pp. 13112–13120. URL: https://aclanthology.org/2023.emnlp-main.810/.
doi:10.18653/v1/2023.emnlp-main.810.
[31] C. Mao, C. Vondrick, H. Wang, J. Yang, Raidar: generative ai detection via rewriting, 2024. URL:
https://arxiv.org/abs/2401.12970. arXiv:2401.12970.
[32] X. Hu, P.-Y. Chen, T.-Y. Ho, Radar: Robust ai-text detection via adversarial learning, 2023. URL:
https://arxiv.org/abs/2307.03838. arXiv:2307.03838.
[33] T. Lavergne, T. Urvoy, F. Yvon, Detecting fake content with relative entropy scoring, in: Proceedings
of the 2008 International Conference on Uncovering Plagiarism, Authorship and Social Software
Misuse - Volume 377, PAN’08, CEUR-WS.org, Aachen, DEU, 2008, p. 27–31.
[34] S. S. Ghosal, S. Chakraborty, J. Geiping, F. Huang, D. Manocha, A. S. Bedi, Towards possibilities
impossibilities of ai-generated text detection: A survey, 2023. URL: https://arxiv.org/abs/2310.15264.
arXiv:2310.15264.
[35] A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, T.
Goldstein, Spotting llms with binoculars: Zero-shot detection of machine-generated text, 2024. URL:
https://arxiv.org/abs/2401.12070. arXiv:2401.12070.
[36] V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, Can ai-generated text be reliably
detected?, 2025. URL: https://arxiv.org/abs/2303.11156. arXiv:2303.11156.
[37] C. Yang, A. Holtzman, How alignment shrinks the generative horizon, 2025. URL: https://arxiv.</p>
      <p>org/abs/2506.17871. arXiv:2506.17871.
[38] B. Park, J. Choi, Identifying the source of generation for large language models, 2024. URL:
https://arxiv.org/abs/2407.12846. arXiv:2407.12846.
[39] H. Guo, S. Cheng, X. Jin, Z. ZHANG, K. Zhang, G. Tao, G. Shen, X. Zhang, Biscope: AI-generated
text detection by checking memorization of preceding tokens, in: The Thirty-eighth Annual
Conference on Neural Information Processing Systems, 2024. URL: https://openreview.net/forum?
id=Hew2JSDycr.
[40] A. R. Basani, P.-Y. Chen, Diversity boosts AI-generated text detection, in: Data in Generative Models
- The Bad, the Ugly, and the Greats, 2025. URL: https://openreview.net/forum?id=QuDDXJ47nq.
[41] T. Kuribayashi, Y. Oseki, S. B. Taieb, K. Inui, T. Baldwin, Large language models are human-like
internally, 2025. URL: https://arxiv.org/abs/2502.01615. arXiv:2502.01615.
[42] T. Chen, C. Guestrin, Xgboost: A scalable tree boosting system, in: Proceedings of the 22nd ACM
SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, ACM, 2016,
p. 785–794. URL: http://dx.doi.org/10.1145/2939672.2939785. doi:10.1145/2939672.2939785.
[43] L. Breiman, Random forests, Machine Learning 45 (2001) 5–32. URL: https://doi.org/10.1023/A:
1010933404324. doi:10.1023/A:1010933404324.
[44] Y. Li, Q. Li, L. Cui, W. Bi, Z. Wang, L. Wang, L. Yang, S. Shi, Y. Zhang, Mage: Machine-generated
text detection in the wild, 2024. URL: https://arxiv.org/abs/2305.13242. arXiv:2305.13242.
[45] X. Zhang, J. Zhao, Y. LeCun, Character-level convolutional networks for text classification, in:
C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, R. Garnett (Eds.), Advances in Neural Information
Processing Systems, volume 28, Curran Associates, Inc., 2015. URL: https://proceedings.neurips.
cc/paper_files/paper/2015/file/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf.
[46] S. Narayan, S. B. Cohen, M. Lapata, Don‘t give me the details, just the summary! topic-aware
convolutional neural networks for extreme summarization, in: E. Rilof, D. Chiang, J. Hockenmaier,
J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language</p>
    </sec>
    <sec id="sec-8">
      <title>A. Motivation Behind Temporal Features</title>
      <p>While static surprisal statistics such as mean, variance, skewness, and kurtosis provide useful
summaries of token-level unpredictability, they overlook the evolution of this unpredictability over time, a
dimension critical to distinguishing human and AI-generated text. Human authors naturally embed
stylistic variability through temporal fluctuations, such as abrupt topic shifts, tonal changes, and bursts
of creativity, which manifest as distinctive temporal dynamics in surprisal sequences. Intuitively, these
temporal features, as listed in Section 3, expose rhythmic and non-stationary patterns characteristic of
human creativity and coherence, typically absent in the more uniform output of large language models.</p>
      <p>Furthermore, through an ablation study on Testbed 4 of the MAGE benchmark, we empirically show
that augmenting static surprisal features with temporal metrics leads to a measurable improvement in
classification accuracy. This highlights the complementary value of temporal dynamics in enhancing
the robustness of AI-generated text detection. Moreover, an analysis of feature importance reveals
that temporal features collectively contribute more than static features, consistently ranking among
the most informative signals for distinguishing between human and AI-generated text. Both these
experiments are thoroughly detailed in our original paper.</p>
      <p>Overall, these findings motivate the inclusion of temporal surprisal features as integral components
of our DivEye framework.</p>
      <p>is human-written (ℓ = 1) or machine-generated (ℓ = 0)
Require: Pretrained auto-regressive language model  (e.g., GPT-2)
Require: XGBoost classifier with hyperparameters
Θ (Appendix C)
Algorithm 1 DivEye: Algorithm for Feature Extraction &amp; Training
Require: Text dataset  = {(, ℓ)}=1, where  is a text input and ℓ ∈ {0, 1} indicates whether it
Ensure: Trained binary classifier</p>
      <p>for each (, ℓ) ∈  do
Initialize an empty feature matrix ℱ ←
Compute token-level log-likelihoods:  ←
Convert to token-level surprisals:  ← −</p>
      <p>()
Append (DivEye(), ℓ) to ℱ</p>
      <p>Compute diversity features DivEye() ∈ R9 as described in Equation (6) using 
end for
return</p>
      <p>Train binary classifier  on feature set ℱ using XGBoost with hyperparameters Θ</p>
    </sec>
    <sec id="sec-9">
      <title>B. Additional Results.</title>
      <p>For further analysis, we refer readers to our main paper [40], which includes: (1) domain-specific
performance of DivEye, (2) model-specific performance of
DivEye, (3) the relative importance of
DivEye when used in a boosted ensemble, (4) DivEye’s efectiveness across diferent base detectors,
and (5) a breakdown of feature importance within DivEye. These additional evaluations further support
the generality, complementarity, and interpretability of our approach.</p>
    </sec>
    <sec id="sec-10">
      <title>C. Hyperparameter Settings</title>
      <p>classifier with standard but tuned settings to handle class imbalance and optimize detection performance.
For our proposed method DivEye, we set the number of bins for entropy computation to 20 and truncate
input sequences at a maximum length of 1024 tokens. All experiments were run on a two NVIDIA RTX
4060Ti (16 GB each), and reported results reflect the median of three runs.</p>
      <p>XGBoost Hyperparameter Value
42
(len(train) − ∑︀ train)/ ∑︀ train
12
200
0.8
0.7
5
1.0
Value
20
1024 + Truncation</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Alahdab</surname>
          </string-name>
          ,
          <article-title>Potential impact of large language models on academic writing</article-title>
          ,
          <source>BMJ evidence-based Medicine 29</source>
          (
          <year>2024</year>
          )
          <fpage>201</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Meyer</surname>
          </string-name>
          , R. J.
          <string-name>
            <surname>Urbanowicz</surname>
            ,
            <given-names>P. C.</given-names>
          </string-name>
          <string-name>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. O'Connor</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>T. J.</given-names>
          </string-name>
          <string-name>
            <surname>Bright</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Tatonetti</surname>
            ,
            <given-names>K. J.</given-names>
          </string-name>
          <string-name>
            <surname>Won</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalez-Hernandez</surname>
          </string-name>
          , et al.,
          <article-title>Chatgpt and large language models in academia: opportunities and challenges</article-title>
          ,
          <source>BioData mining 16</source>
          (
          <year>2023</year>
          )
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Lund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. R.</given-names>
            <surname>Mannuru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shimray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Chatgpt and a new academic reality: Artificial intelligence-written research papers and the ethics of the large language models in scholarly publishing</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>74</volume>
          (
          <year>2023</year>
          )
          <fpage>570</fpage>
          -
          <lpage>581</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yuan</surname>
          </string-name>
          , G. Shi,
          <article-title>Dynamic content generation in large language models with real-time constraints (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Coenen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Reif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ippolito</surname>
          </string-name>
          ,
          <article-title>Wordcraft: story writing with large language models</article-title>
          ,
          <source>in: Proceedings of the 27th International Conference on Intelligent User Interfaces</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>841</fpage>
          -
          <lpage>852</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>A. De Giorgio</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Matrone</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mafei</surname>
          </string-name>
          ,
          <article-title>Detecting large language models in exam essays</article-title>
          ,
          <source>in: 2025 IEEE Engineering Education World Conference (EDUNINE)</source>
          , IEEE,
          <year>2025</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Papageorgiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chronis</surname>
          </string-name>
          , I. Varlamis,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Himeur</surname>
          </string-name>
          ,
          <article-title>A survey on the use of large language models (llms) in fake news</article-title>
          ,
          <source>Future Internet</source>
          <volume>16</volume>
          (
          <year>2024</year>
          )
          <fpage>298</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Telenti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Auli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. L.</given-names>
            <surname>Hie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Maher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Saria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Ioannidis</surname>
          </string-name>
          ,
          <article-title>Large language models for science and medicine</article-title>
          ,
          <source>European journal of clinical investigation 54</source>
          (
          <year>2024</year>
          )
          <article-title>e14183</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Törnberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Valeeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uitermark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bail</surname>
          </string-name>
          ,
          <article-title>Simulating social media using large language models to evaluate alternative news feed algorithms</article-title>
          ,
          <source>arXiv preprint arXiv:2310.05984</source>
          (
          <year>2023</year>
          ). Processing, Association for Computational Linguistics, Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>1797</fpage>
          -
          <lpage>1807</lpage>
          . URL: https://aclanthology.org/D18-1206/. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D18</fpage>
          -1206.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Takamura</surname>
          </string-name>
          , H. Nakayama,
          <article-title>SciXGen: A scientific paper dataset for context-aware text generation</article-title>
          , in: M.
          <article-title>-</article-title>
          <string-name>
            <surname>F. Moens</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , S. W.-t. Yih (Eds.),
          <source>Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2021</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Punta Cana, Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>1483</fpage>
          -
          <lpage>1492</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .findings-emnlp.
          <volume>128</volume>
          /. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .findings-emnlp.
          <volume>128</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>C.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Niculae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Danescu-Niculescu-Mizil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Winning arguments: Interaction dynamics and persuasion strategies in good-faith online discussions</article-title>
          ,
          <source>in: Proceedings of the 25th International Conference on World Wide Web, WWW '16, International World Wide Web Conferences Steering Committee</source>
          ,
          <year>2016</year>
          . URL: http://dx.doi.org/10.1145/2872427.2883081. doi:
          <volume>10</volume>
          .1145/2872427.2883081.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI</source>
          (
          <year>2019</year>
          ). URL: https://cdn.openai.
          <article-title>com/better-language-models/ language_models_are_unsupervised_multitask_learners</article-title>
          .pdf, accessed:
          <fpage>2024</fpage>
          -11-15.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Roller,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. V.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mihaylov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shleifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shuster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Simig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Koura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sridhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          , Opt: Open pre-trained
          <source>transformer language models</source>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2205.01068. arXiv:
          <volume>2205</volume>
          .
          <fpage>01068</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [51]
          <string-name>
            <surname>B. W.</surname>
          </string-name>
          et al.,
          <article-title>Bloom: A 176b-parameter open-access multilingual language model</article-title>
          ,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2211.05100. arXiv:
          <volume>2211</volume>
          .
          <fpage>05100</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Greiner-Petter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2025:
          <article-title>Voight-Kampf Generative AI Detection, Multilingual Text Detoxification, Multi-Author Writing Style Analysis, and Generative Plagiarism Detection</article-title>
          , in: J.
          <string-name>
            <surname>C. de Albornoz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piroi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Sixteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsivgun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abassy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mansurov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Ta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Elozeiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Tomar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Artemova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Habash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Overview of the “Voight-Kampf” Generative AI Authorship Verification Task at PAN</article-title>
          and
          <article-title>ELOQUENT 2025</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , D. Spina (Eds.),
          <source>Working Notes of CLEF 2025 - Conference and Labs of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kolyada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Grahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elstner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Loebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <article-title>Continuous Integration for Reproducible Shared Tasks with TIRA.io</article-title>
          ,
          <source>in: Advances in Information Retrieval. 45th European Conference on IR Research (ECIR</source>
          <year>2023</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2023</year>
          , pp.
          <fpage>236</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>O.</given-names>
            <surname>Halvani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Graner</surname>
          </string-name>
          ,
          <article-title>On the usefulness of compression models for authorship verification</article-title>
          ,
          <source>in: Proceedings of the 12th International Conference on Availability, Reliability and Security</source>
          , ARES '17,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2017</year>
          . URL: https://doi.org/10.1145/3098954.3104050. doi:
          <volume>10</volume>
          .1145/3098954.3104050.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>