<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Generative AI Detection Using Simple Feature Selection and SVM</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Joseph Larson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indiana University</institution>
          ,
          <addr-line>Bloomington, IN.</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>Generative AI detection has been of interest for at least the past decade, but especially since the emergency of transformer powered LLMs. This paper treats the task as a binary classification problem, where  ∈ [0, 1]. I chose to use a traditional Support Vector Classifier (SVC) with sets of features chosen from examination of the training data set to determine what features human authors, as opposed to AI, are more likely to employ. I found the top 40 unigram and bigrams, along with the top 15 punctuation features, to be the most informative. When combined and input into my SVC, I achieved a mean (the mean of all scores used for the task) score of 94.89.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;PAN 2025</kwd>
        <kwd>AI Detection</kwd>
        <kwd>Feature Selection</kwd>
        <kwd>SVC</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        In terms of impressionistic feature diferences between human and A.I. authors, it has been stated
that overall LLMs are more focused (that is to say, they never leave the subject matter at hand), more
objective and highly formal. In contrast, their human counterpoints employ more subjective language,
less formal and demonstrate an increased propensity to stray from the topic. Linguistically, humans
employ less nouns and conjunctions and LLMs employ less punctuation and adverbs. Dependency
relations for humans are also much shorter. Lastly, humans show more ‘creativity in terms of word
choices’, therefore human texts on average show a higher type to token ratio. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
To this last point, it has been demonstrated that LLMs sample from a limited amount of tokens to
generate natural looking text, e.g. through mass sampling [18] or k-max sampling [19]. For this reason,
[20] found relative success using n-gram tf-idf features (unigrams and bigrams) to distinguish between
human and GPT2 redacted web pages. They created a dataset using three sampling methods: k-sampling
(sampling the highest probability tokens until a threshold of specified tokens is reached), p-sampling
(sampling from the smallest possible set of words until a cumulative probability is reached) and so-called
‘pure’ sampling (a.k.a temperature sampling, where lower ‘temperatures’ are associated with higher
probabilities for tokens). Since k samples overproduce common words, they were the easiest to detect.
[21] use three statistical features: probability of each word, absolute rank of each word and the entropy
of the word’s distribution. It was found that GPT2 oversamples certain words, allowing the model (in
this case, BERT) to easily detect LLM generated text.
      </p>
      <p>
        In terms of models, [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] has claimed that the best detector of LLM generated text are LLMs themselves. The
aforementioned Gehrmann et al. study found that finetuning GPT2 did not yield better results. Inspired
by this, my goal was to develop an LLM text detector that relied on purely statistical models rather
than transformers; this type of method would be computationally inexpensive and easily employable
on a person’s local machine.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task Overview</title>
      <p>The present task involves a binary classification task, whereby documents in the dataset are classified as
either being human authored or machine authored. I experiment with systems that return both binary
labels and probabilities, where scores approaching 0 indicate probable human authorship and scores
approaching 1 indicate probable machine authorship. Scores of 0.5 indicate that the system is unsure.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>The training set used had a total of 23,707 documents while the validation set had a total of 3,589.
The human class was lower for both splits, with it representing roughly two fifths of training set and
roughly a third of the validation set. Various models were used to create the machine documents. Table
1 contains a more detailed summary.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model Selection</title>
        <p>As stated in section 2, my goal for this project was specifically to not use a transformer model. I wanted
to use a model that could be easily employed on more traditional, less computationally expensive model.
SVC’s and SVM’s have both proved to be successful in the domain of authorship attribution, so I chose
to use an SVC as my model. In terms of hyperparameters, I experimented with diferent kernels and
class weights. I found RBF to be the most dynamic kernel and ‘balancing’ (see equation 1:  is the
class weight,  is the number of documents,  is the number of class and () is the bin count of
the classes).</p>
        <p>=  ()

(1)</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Feature Selection</title>
        <p>My initial approach to feature selection was examining the dataset for anomalies. I wanted to first
analyze the claim that human texts haver higher type to token ratios. Figure 1 shows a density plot for
the training dataset of type to token ratios. This graph shows that although the human texts show a
normally distributed curve and the machine generated text is much more irregular distributed, there is
a lot of overlap, meaning this feature probably would not be helpful for classification.
Next, I examined the top unigrams, bigrams and trigrams for both classes, to see if certain n-grams were
more common in one class. I came to the conclusion that the top forty unigrams and bigrams were the
most informative in distinguishing between the two classes. After this, since the literature had stated
AI uses less punctuation, I considered of using this as a feature for the SVC. After careful examination
of the dataset, the first 15 punctuation patterns proved to be the most informative. A summary of the
features used for my model can be found in Table 2.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Evaluation</title>
        <p>To evaluate the performance of the model, the metrics used by the task were:
• AUC (a.k.a. AUC Roc score): The area under the curve score.
• C@1 (a.k.a. Classification at 1): The percentage of instances where the top score was the correct
one.
• F0.5: The harmonic mean of precision and recall, with  = 0.5.
• F1: The harmonic mean of precision and recall, with  = 1.
• Brier: The complement of the Brier Score Loss, which is the mean squared diference between
the predicted class and actual class.
• Mean: The arithmetic mean of all previous scores.</p>
        <p>All metrics naturally produce a score between 0 and 1 and are multiplied by 100.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>To obtain my results, I experimented with diferent numbers of features. While experimenting with
just unigrams and bigrams, I ran models for up to 1,000 count features. My conclusion was that only
the top 40 impacted the document class. For the punctuation features, I found there to be 153 unique
punctuation patterns within the dataset. I ran experiments with varying numbers of these patterns and
found the top 15 to be the most distinctive. My hypothesis was then that combining the features would
yield an improvement in the model; I was proven correct. The best model I was able to train was with
the top 15 punctuation features and the top 40 unigram and bigram features. Figure 2 shows a PCA of
the SVM’s feature space. Table 3 shows my final results on the validation set and table 4 shows my
results on the final test set compared to the other baselines.
86.47
93.53
95.49</p>
      <p>F1
89.36
94.02
96.09</p>
      <p>Brier
85.48
92.23
94.90</p>
      <p>Mean
84.97
92.42
94.89</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Conclusion</title>
      <p>With this short paper, I have shown that feature selection is still an efective method in AI detection.
LLM models clearly still produce text that over sample certain words, as found to the case by [21, 20].
In addition, punctuation patterns prove to be a distinguishing factor: humans use wider ranges of
punctuation patterns and use them with higher frequency. I have been able to demonstrate all of this
without finetune a State-of-the-Art transformer model, with experiments I have run on my local machine</p>
      <p>F0.5u</p>
      <p>Mean
(when doing feature selection, I did at certain points have to use my university’s super computer1,
however one feature selection was finished training for my models required less than a minute). While
transformer models may still yield higher performance than the model I present in this paper, my work
serves as a reminder that feature selection is still a powerful method within the domain of AI text
detection.</p>
      <p>My work, is of course, not without its limitations. I could have done more hyperparameter fine-tuning
to possibly improve my model even more. I also could have carried a more profound analysis of what
features were actually relevant in distinguishing the two classes. The top baseline for this task reports
TF-IDF features as having achieved a mean of 0.98; while I experimented with TF-IDF features, my
ifndings were that they did not outperform raw count features. Future work should focus upon creating
more robust datasets that not only contain a diversity of diferent models, but also a diversity of sampling
models, as demonstrated by [20] to be an important factor in AI detection.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Declaration on Generative AI</title>
      <p>During the preparation of this work, the author used not a single AI tool for any purpose.
1I acknowledge the Indiana University Pervasive Technology Institute for providing supercomputing and storage resources
that have contributed to the research results reported within this paper [22]
USA, 2018, pp. 2814–2822. URL: https://aclanthology.org/C18-1238/.
[13] E. Ferracane, S. Wang, R. Mooney, Leveraging discourse information efectively for authorship
attribution, in: G. Kondrak, T. Watanabe (Eds.), Proceedings of the Eighth International Joint
Conference on Natural Language Processing (Volume 1: Long Papers), Asian Federation of Natural
Language Processing, Taipei, Taiwan, 2017, pp. 584–593. URL: https://aclanthology.org/I17-1059/.
[14] Y. Seroussi, I. Zukerman, F. Bohnert, Authorship attribution with Latent Dirichlet Allocation,
in: S. Goldwater, C. Manning (Eds.), Proceedings of the Fifteenth Conference on Computational
Natural Language Learning, Association for Computational Linguistics, Portland, Oregon, USA,
2011, pp. 181–189. URL: https://aclanthology.org/W11-0321/.
[15] Y. Seroussi, F. Bohnert, I. Zukerman, Authorship attribution with author-aware topic models,
in: H. Li, C.-Y. Lin, M. Osborne, G. G. Lee, J. C. Park (Eds.), Proceedings of the 50th Annual
Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Association
for Computational Linguistics, Jeju Island, Korea, 2012, pp. 264–269. URL: https://aclanthology.
org/P12-2052/.
[16] Y. Seroussi, I. Zukerman, F. Bohnert, Authorship attribution with topic models, Computational
Linguistics 40 (2014) 269–310. URL: https://aclanthology.org/J14-2003/. doi:10.1162/COLI_a_
00173.
[17] M. Fröbe, M. Wiegmann, N. Kolyada, B. Grahm, T. Elstner, F. Loebe, M. Hagen, B. Stein, M. Potthast,
Continuous Integration for Reproducible Shared Tasks with TIRA.io, in: Advances in Information
Retrieval. 45th European Conference on IR Research (ECIR 2023), Lecture Notes in Computer
Science, Springer, Berlin Heidelberg New York, 2023, pp. 236–241.
[18] J. Gu, K. Cho, V. O. Li, Trainable greedy decoding for neural machine translation, in: M. Palmer,
R. Hwa, S. Riedel (Eds.), Proceedings of the 2017 Conference on Empirical Methods in Natural
Language Processing, Association for Computational Linguistics, Copenhagen, Denmark, 2017,
pp. 1968–1978. URL: https://aclanthology.org/D17-1210/. doi:10.18653/v1/D17-1210.
[19] A. Fan, M. Lewis, Y. Dauphin, Hierarchical neural story generation, in: I. Gurevych, Y. Miyao
(Eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics
(Volume 1: Long Papers), Association for Computational Linguistics, Melbourne, Australia, 2018,
pp. 889–898. URL: https://aclanthology.org/P18-1082. doi:10.18653/v1/P18-1082.
[20] I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W.</p>
      <p>Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGufie, J. Wang, Release strategies and
the social impacts of language models, 2019. arXiv:1908.09203.
[21] S. Gehrmann, H. Strobelt, A. Rush, GLTR: Statistical detection and visualization of generated
text, in: M. R. Costa-jussà, E. Alfonseca (Eds.), Proceedings of the 57th Annual Meeting of the
Association for Computational Linguistics: System Demonstrations, Association for Computational
Linguistics, Florence, Italy, 2019, pp. 111–116. URL: https://aclanthology.org/P19-3019/. doi:10.
18653/v1/P19-3019.
[22] C. A. Stewart, V. Welch, B. Plale, G. Fox, M. Pierce, T. Sterling, Indiana University Pervasive
Technology Institute, Technical Report, Indiana University, 2017. URL: https://doi.org/10.5967/
K8G44NGB. doi:10.5967/K8G44NGB.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <article-title>Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions</article-title>
          .,
          <source>Technology Science</source>
          <volume>2019121801</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. I.</given-names>
            <surname>Adelani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Echizen</surname>
          </string-name>
          ,
          <article-title>Generating sentimentpreserving fake online reviews using neural language models and their human-</article-title>
          and
          <source>machine-based detection</source>
          ,
          <year>2019</year>
          . arXiv:
          <year>1907</year>
          .09177.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zellers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Holtzman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rashkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bisk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Roesner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Choi</surname>
          </string-name>
          , Defending against Neural Fake News, Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Herbert-Voss</surname>
          </string-name>
          , G. Krueger,
          <string-name>
            <given-names>T.</given-names>
            <surname>Henighan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hesse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          , E. Sigler,
          <string-name>
            <given-names>M.</given-names>
            <surname>Litwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Berner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>McCandlish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learners</article-title>
          , in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          . URL: https://proceedings.neurips.cc/paper_files/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Uchendu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Authorship attribution for neural text generation</article-title>
          , in: B.
          <string-name>
            <surname>Webber</surname>
            , T. Cohn,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
          </string-name>
          , Y. Liu (Eds.),
          <source>Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>8384</fpage>
          -
          <lpage>8395</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .emnlp-main.
          <volume>673</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .emnlp-main.
          <volume>673</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>How close is chatgpt to human experts? comparison corpus, evaluation, and detection</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2301</volume>
          .
          <fpage>07597</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2025:
          <article-title>Generative AI Authorship Verification, Multi-Author Writing Style Analysis, Multilingual Text Detoxification, and Generative Plagiarism Detection, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsivgun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abassy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mansurov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Ta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Elozeiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Tomar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Artemova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Habash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Overview of the “VoightKampf” Generative AI Authorship Verification Task at PAN</article-title>
          and
          <article-title>ELOQUENT 2025</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , D. Spina (Eds.),
          <source>Working Notes of CLEF 2025 - Conference and Labs of the Evaluation Forum, CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stevenson</surname>
          </string-name>
          ,
          <article-title>Continuous n-gram representations for authorship attribution</article-title>
          , in: M.
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Blunsom</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Koller (Eds.),
          <source>Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>2</volume>
          ,
          <string-name>
            <surname>Short</surname>
            <given-names>Papers</given-names>
          </string-name>
          , Association for Computational Linguistics, Valencia, Spain,
          <year>2017</year>
          , pp.
          <fpage>267</fpage>
          -
          <lpage>273</lpage>
          . URL: https://aclanthology.org/ E17-2043/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nandan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ralhan</surname>
          </string-name>
          ,
          <article-title>An investigation of supervised learning methods for authorship attribution in short hinglish texts using char word n-grams</article-title>
          ,
          <year>2018</year>
          . URL: https://arxiv.org/abs/
          <year>1812</year>
          . 10281. arXiv:
          <year>1812</year>
          .10281.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zečević</surname>
          </string-name>
          ,
          <article-title>N-gram based text classification according to authorship</article-title>
          , in: I.
          <string-name>
            <surname>Temnikova</surname>
          </string-name>
          , I. Nikolova, N. Konstantinova (Eds.),
          <source>Proceedings of the Second Student Research Workshop associated with RANLP</source>
          <year>2011</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hissar, Bulgaria,
          <year>2011</year>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>149</lpage>
          . URL: https://aclanthology.org/R11-2023/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sundararajan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Woodard</surname>
          </string-name>
          ,
          <article-title>What represents “style” in authorship attribution?</article-title>
          , in: E. M.
          <string-name>
            <surname>Bender</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Derczynski</surname>
          </string-name>
          , P. Isabelle (Eds.),
          <source>Proceedings of the 27th International Conference on Computational Linguistics</source>
          , Association for Computational Linguistics, Santa Fe, New Mexico,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>