<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Xiv:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.48550/arXiv.2306.04751</article-id>
      <title-group>
        <article-title>Overview of the “Voight-Kampf ” Generative AI Authorship Verification Task at PAN and ELOQUENT 2024</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Janek Bevendorf</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matti Wiegmann</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jussi Karlgren</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luise Dürlich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evangelia Gogoulou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aarne Talman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Efstathios Stamatatos</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Potthast</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benno Stein</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silo AI</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Helsinki</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RISE Research Institutes of Sweden</institution>
          ,
          <addr-line>Stockholm</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Helsinki</institution>
          ,
          <addr-line>Helsinki</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Kassel</institution>
          ,
          <addr-line>hessian.AI, and ScaDS.AI</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <volume>14165</volume>
      <fpage>74764</fpage>
      <lpage>74786</lpage>
      <abstract>
        <p>The “Voight-Kampf” Generative AI Authorship Verification task aims to determine whether a text was generated by an AI or written by a human. As in its fictional inspiration, 1 the Voight-Kampf task structures AI detection as a builder-breaker challenge: The builders, participants in the PAN lab, submit software to detect AI-written text and the breakers, participants in the ELOQUENT lab, submit AI-written text with the goal of fooling the builders. We formulate the task in a way that is reminiscent of a traditional authorship verification problem, where given a pair of texts, their human or machine authorship is to be inferred. For this first task installment, we further restrict the problem so that each pair is guaranteed to contain one human and one machine text. Hence the task description reads: Given two texts, one authored by a human, one by a machine: pick out the human. In total, we evaluated 43 detection systems (30 participant submissions and 13 baselines), ranging from linear classifiers to perplexity-based zero-shot systems. We tested them on 70 individual test set variants organized in 14 base collections, each designed on diferent constraints such as short texts, Unicode obfuscations, or language switching. The top systems achieve very high scores, proving themselves not perfect but suficiently robust across a wide range of specialized testing regimes. Code used for creating the datasets and evaluating the systems, baselines, and data are available on GitHub.2</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Generative AI is undoubtedly a disruptive technology in the information ecosystem. In particular,
large language models (LLMs) have many desirable applications in writing assistance and information
access. Although often welcome, unlimited text generation raises concerns in many areas of creation;
examples include education and assessment, academic articles and reviews, synthetic misinformation and
disinformation, and social bots that influence public discourse. These troubling applications undermine
trust in (written) information. Recognizing the fingerprint of AI-generated text thus becomes a promising
element for a healthy future information ecosystem, which will become increasingly sophisticated as
the fluency and naturalness of the generated text increases.</p>
      <p>
        The Voight-Kampf task investigates the feasibility of identifying whether text is written by a human
author or generated by a language model. We recognize that the detection of AI-generated text is
closely related to the identification of human authorship where, prospectively and with increasing
ifdelity and diversification of models, each AI model can be considered an author that exhibits particular
and identifiable characteristics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this task, we adapt this idea and formulate AI detection as an
authorship problem. This allows us not only to draw upon experience from previous LLM detection
shared tasks [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2, 3, 4, 5</xref>
        ], but also to adapt decades of theoretical and engineering work on author
identification, including past work at PAN [
        <xref ref-type="bibr" rid="ref10 ref11 ref6 ref7 ref8 ref9">6, 7, 8, 9, 10, 11</xref>
        ].
1 In the movie “Blade Runner”, the eponymous oficers use the Voight-Kampf machine to test whether a subject is a replicant.
2 Code and data: https://github.com/pan-webis-de/pan24-generative-ai-authorship-verification
CLEF 2024: Conference and Labs of the Evaluation Forum, September 09–12, 2024, Grenoble, France
      </p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>Input / Task
1. { ? , ? }
2. { ? , ? }
3. { ? , ? }
4. { ? , ? }
5. { ? , ? }
6. { ? , ? }
7.</p>
      <p>?
→−</p>
      <p>Possible Assignment Patterns</p>
      <p>A ,</p>
      <p>M</p>
      <p>Identifying a single disputed document as AI-generated without reference is an open problem and the
most dificult formulation of the AI detection task. Although the literature suggests limited success in
solving this problem with the current generation of LLMs, it is questionable whether this will remain the
case as the technology improves. Aside from mixed human-machine authorship, we have broken down
the relevant formulations of the problem with increasing levels of dificulty to gain a more fundamental
understanding of the task at hand and the feasibility of potential solutions. Figure 1 visualizes the
cascade of all problem variants from the easiest (Task 1) to the most dificult (Task 7). In the easiest
case, two documents of unknown authorship are given, but we guarantee that exactly one of them
was created by a human A and the other by a machine M . This constraint is relaxed in the following
variants, where e.g., both texts can be generated by a machine, { M , M }. In the hardest case, a single
text is given, which could be either A or M .</p>
      <p>The Voight-Kampf task follows the easiest formulation of the task to establish a feasibility baseline.
The task description is: Given two texts, one authored by a human, one by a machine: pick out the human.</p>
      <p>Our evaluation campaign was organized in a builder-breaker setup in collaboration between the
ELOQUENT and PAN labs at CLEF. In this setup, both systems and evaluation resources are built by
the participants and compete in an adversarial setting. The builders at PAN create analysis tools, in
this case for classifying texts as human-authored vs. machine generated. The breakers at ELOQUENT
provide the data points to break the analysis tools, in this case to fool the classifiers into believing
that a machine-generated text was authored by a human. This adversarial design is intended to focus
participants’ eforts by providing increasingly challenging baselines for both classifiers and generative
language models.</p>
      <p>In total, our evaluation campaign includes 14 breaker dataset collections consisting in total of 70
individual conditions (or variants), eight provided by PAN as baselines and six submitted by ELOQUENT
participants. We evaluate 43 builder systems, 13 of which were provided by PAN as baselines and 30 that
were submitted by PAN participants. We find that the strongest systems are highly efective and robust
across dataset variants in separating machine-generated and human-authored text, at least given the
latest generation of large language models. The most dificult evaluation settings across all submitted
systems are unexpected languages and very short texts, closely followed by Unicode obfuscations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Detecting AI-generated Text</title>
      <p>In total, 30 valid system were submitted to the builder task in addition to the baseline systems we
provided. Of these, 28 teams submitted descriptions of their systems in the form of notebook papers.
Table 4 shows the final system ranking and Table 1 shows an overview of the systems.</p>
      <sec id="sec-2-1">
        <title>2.1. Baselines</title>
        <p>We provided implementations of six baseline systems to compare submitted systems against four
state-of-the-art zero-shot LLM detection baselines and two adapted authorship verification baselines.</p>
        <p>The zero-shot LLM detection baselines are: (1) Binoculars [12], (2) DetectLLM [13] (both NPR and
LRR scoring mode), (3) DetectGPT [14], and (4) Fast-DetectGPT [15]. All three were provided in two
variants using either Falcon 7B [16] or Mistral 7B [17] to estimate text perplexities. The required text
perturbations for DetectGPT and DetectLLM-NPR were generated with T5-3B [18].</p>
        <p>The two authorship verification baselines were adapted to the LLM detection task by splitting each
text in half and comparing the two halves against each other under the assumption that LLM texts
are stylistically more self-similar than human texts. The baselines provided are a compression model
(PPMd Compression-based Cosine) [19, 20] and short-text authorship unmasking [21, 22].</p>
        <p>As an additional seventh baseline, we measured and compared the text lengths in characters. This
baseline serves as both a quasi-random baseline and as a data sanity check.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Builder Submissions to PAN</title>
        <p>While our baseline systems reproduce established methods in either authorship verification or intrinsic,
zero-shot LLM detection, the participant systems cover a broad range of approaches. Table 1 shows
an overview of the essential elements across all systems. The most frequent approach is to train or
ifne-tune a (neural) classifier with term-based features, in most cases (BERT) embeddings. Many systems
either apply some training regime modification (such as PU loss or R-Drop), use ensembles, and / or
expand the given training data with other LLM detection datasets. Some systems use engineered
features like perplexity (PPL), term frequencies (TF), stylometric features (text length, properties of
token distributions), or a combination of these.</p>
        <p>The systems use a broad range of classification models (SVM, XGBoost, LSTM, CNN) but in most
cases in combination with BERT. Most systems only classify if one given document is machine generated
and decide which of the two input documents is human-written by comparing the predicted probability,
similar to how the provided baselines do it. Some participants, however, also trained models to directly
discriminate between the pairings as in the traditional author verification setting. In some cases,
participants also utilized LLMs for the detection task, often using Low-Rank adaptation (LoRA).</p>
        <sec id="sec-2-2-1">
          <title>2.2.1. Term-based Systems</title>
          <p>
            Huang et al. [24], the second-ranked system (0.921), derives a method named Tri-Sentence Analysis
(TSA) from the multi-scale positive-unlabeled (MPU [
            <xref ref-type="bibr" rid="ref35">51</xref>
            ]) LLM-detector. TSA dissects the documents
into small (3-sentence) chunks and trains a BERT classifier for binary AI detection on the chunks, where
each chunk inherits the class label of the document. The chunk scores are then averaged to to estimated
the prediction on the original documents.
          </p>
          <p>Lorenz et al. [25], the third-ranked system (0.886), uses feature-based, supervised classification: Naive
Bayes, Logistic Regression, and linear SVM based on the top 1,000 TF-IDF term count features.</p>
          <p>Guo et al. [26] (0.884) use a supervised hybrid method that utilizes various types of BERT embeddings
and Google Books word frequency features embedded with an LSTM.</p>
          <p>
            Several other mostly well-performing models fine-tuned a BERT (variant) with simple modifications:
Lin et al. [27] (0.851) with R-Drop [
            <xref ref-type="bibr" rid="ref36">52</xref>
            ] regularization, Yadagiri et al. [
            <xref ref-type="bibr" rid="ref14">30</xref>
            ] (0.806) with augmentation
through linguistic features (vocabulary count, a word-density measure, and POS counts), Lv et al. [
            <xref ref-type="bibr" rid="ref15">31</xref>
            ]
(0.804, DeBERTa) with Reptile meta-learning [
            <xref ref-type="bibr" rid="ref37">53</xref>
            ], Cao et al. [
            <xref ref-type="bibr" rid="ref17">33</xref>
            ] (0.778) with dataset augmentation,
Huang et al. [
            <xref ref-type="bibr" rid="ref19">35</xref>
            ] (0.735) with R-Drop in combination and additional training data from Kaggle [
            <xref ref-type="bibr" rid="ref38">54</xref>
            ] for
oversampling the human examples, Petropoulos and Petropoulos [
            <xref ref-type="bibr" rid="ref25">41</xref>
            ] (0.641) with a Bi-LSTM between
the RoBERTa and the linear classification layer, Wu et al. [
            <xref ref-type="bibr" rid="ref26">42</xref>
            ] (0.608) with an extra transformer-layer
between the BERT and the linear classification layer, Zhu and Kong [
            <xref ref-type="bibr" rid="ref28">44</xref>
            ] (0.555) by using DeBERTa
instead of BERT, Sun et al. [
            <xref ref-type="bibr" rid="ref29">45</xref>
            ] (0.531) with a 2D-CNN-layer between a BERT and a linear layer, and
Huang et al. [
            <xref ref-type="bibr" rid="ref32">48</xref>
            ] (0.480) with a siamese DeBERTa with contrastive learning and domain adaptation.
          </p>
          <p>
            Two systems build on contextualized word embeddings from BERT but use a more involved model:
Guo et al. [
            <xref ref-type="bibr" rid="ref18">34</xref>
            ] (0.763) uses a Bi-LSTM followed by a transformer layer for classification. As input for the
LSTM, the authors use BERT embeddings concatenated with seven stylometric and linguistic features
(lexical diversity, average sentence length, average word length, the number of grammatical errors,
sentiment tendency, repetition rate, and stop word ratio). Valdez-Valenzuela and Gómez-Adorno [
            <xref ref-type="bibr" rid="ref20">36</xref>
            ]
(0.727) use a mixture of co-occurrence graph features embedded with a GNN, stylometric features, and
x
x
x
x
x
x
x
x
x
x
x
x
BERT document embeddings and augment the training data with additional human texts.
          </p>
          <p>
            Two systems use ensembles based on multiple fine-tuned BERT models: Qin et al. [
            <xref ref-type="bibr" rid="ref24">40</xref>
            ] (0.680) use a
voting ensemble of a basic BERT and a BERT with R-Drop regularization, trained on additional data
from Kaggle. Sheykhlan et al. [
            <xref ref-type="bibr" rid="ref33">49</xref>
            ] (0.460) use a voting ensemble with BERT, RoBERTa, and Electra.
          </p>
          <p>
            Finally, five systems use generative LLM as base for a classifier: Gritsai et al. [
            <xref ref-type="bibr" rid="ref16">32</xref>
            ] (0.796) use an
ensemble of multiple Mistral models, each fine-tuned via QLoRA on texts generated by diferent types
of LLMs. Ye et al. [
            <xref ref-type="bibr" rid="ref21">37</xref>
            ] (0.722) fine-tune a T5 model with language modeling head to predict the tokens
“positive” (machine text) or “negative” (human text) for a given document. If neither of these tokens is
the most likely, the system outputs “undecided”. Lin et al. [
            <xref ref-type="bibr" rid="ref27">43</xref>
            ] (0.565) also fine-tune a T5 to predict
“positive” or “negative” after a new, special token and assign the probability of whichever token is more
likely. Lei et al. [
            <xref ref-type="bibr" rid="ref30">46</xref>
            ] (0.504) fine-tune a ChatGLM model for authorship attribution, i.e. the model learns
to predict tokens that indicate either “Human” any one of the particular LLMs. The predicted classes
are then transformed back into a binary AI detection score. Wu and Guan [
            <xref ref-type="bibr" rid="ref34">50</xref>
            ] (0.450) use a language
model pre-trained for NLI under the assumption that LLM-generated texts show weaker coherence and
textual entailment between sentences.
          </p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.2. Perplexity-based Systems</title>
          <p>
            Miralles et al. [
            <xref ref-type="bibr" rid="ref13">29</xref>
            ] (0.806) use an XGBoost classifier with features (mean, stddev, etc.) of the distributions
of next-token probabilities across five current LLMs, in combination with established, stylometric
features.
          </p>
          <p>
            Huang and Grieve [
            <xref ref-type="bibr" rid="ref23">39</xref>
            ] (0.683) use the perplexity of “authorial language models” as features for
an SVM classifier. As authorial language, a GPT-2 is fine-tuned for each known LLM and one for all
humans. For both disputed text, the perplexity of all authorial LMs is measured and provided as feature
vector for the SVM. The SVM is then trained for verification.
          </p>
          <p>
            Liu and Kong [
            <xref ref-type="bibr" rid="ref31">47</xref>
            ] (0.497) use the perplexity of a GPT-2 model as discriminator under the assumption
that the text with the lower GPT-2 perplexity score is AI-generated. The perplexity is calculated as a
sum on a sliding 1,024-token window.
          </p>
        </sec>
        <sec id="sec-2-2-3">
          <title>2.2.3. Systems Using Terms and Perplexity</title>
          <p>
            Tavan and Najafi [ 23], the first-ranked system (0.924), uses an ensemble of two LLMs (Mistral and
Llama2) and the Binoculars baseline. The LLMs (with classification head) were fine-tuned on the training
data via LoRA. The final score is the average decision across all three models. Abburi et al. [
            <xref ref-type="bibr" rid="ref12">28</xref>
            ] (0.843)
use a combination of a RoBERTa-based AI detector, token-level probability features from multiple GPT-2
variants and E5 document embeddings to classify AI texts. Finally, Chen and Kong [
            <xref ref-type="bibr" rid="ref22">38</xref>
            ] (0.694) use a
voting ensemble of three models: 2× BERT and a GPT-2. The BERT models predict which of two given
texts is written by a human, the first or the second. Trained on diferent splits of the training data (with
inverted positions and labels). Both models are identical, but the input pairs are flipped to counteract
the truncation of long documents. The GPT-2 model is used to calculate the perplexity of each text (for
ca. 500 character chunks and summed for each text), where text with the higher perplexity is taken as
the human written one. The training data is augmented with the Kaggle DAIGT v2 Train [
            <xref ref-type="bibr" rid="ref39">55</xref>
            ] dataset.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <p>We used two dataset collections in the evaluation of the “Voight-Kampf” task. The first collection, in
accordance with the builder-breaker pattern, collects all datasets submitted by individual ELOQUENT
participants. This collection is described in detail in Section 3.1.</p>
      <p>As both labs were ran concurrently, we created a second “bootstrap” dataset collection PAN AI News
2021 in lieu of training data to get participants started. Participants were encouraged to also use
additional data from other sources; Part of this bootstrap collection was held back as a test collection.</p>
      <sec id="sec-3-1">
        <title>Genre and Style:</title>
        <p>The text is an informative piece providing a comprehensive overview of Malaysia’s geography, history,
government structure, economy, and cultural diversity. Its tone is neutral and factual, aiming to
educate the reader about various aspects of the country.</p>
        <p>Content:
• Malaysia is a federal constitutional monarchy in Southeast Asia, comprising thirteen states
and three federal territories.
• Geographically divided into Peninsular Malaysia and East Malaysia (Malaysian Borneo) by
the South China Sea.
• Shares borders with Thailand, Singapore, Vietnam, Indonesia, Brunei, and maritime borders
with the Philippines.
• Capital city: Kuala Lumpur; federal government seat: Putrajaya.
• Multi-ethnic and multi-cultural country with Islam as the state religion, but freedom of religion
is protected.
• Boasts a strong economy, historically driven by natural resources but expanding into sectors
like science, tourism, commerce, and medical tourism.</p>
        <p>The PAN AI News 2021 dataset consists of news articles written by humans or LLMs and its creation is
described in Section 3.2.</p>
        <sec id="sec-3-1-1">
          <title>3.1. Breaker Submissions to ELOQUENT</title>
          <p>ELOQUENT formulated the breaker challenge for this task, with the objective for participants to use
models and systems of their choice to fool classifiers into believing their output is authored by a human.
The organizers selected 29 human authored texts, five sample items for pre-experiment tuning and
testing purposes, and 24 items proper. Each text was of 300 to 600 words length and summaries of each
text were generated by the organizers using OpenAI’s ChatGPT service with the prompt:
Summarize the following text in five to six short bullet points and give an overall
description of the genre and tone of the text.</p>
          <p>Those machine-generated summaries were then shared with the participants, so their systems could
generate derivative short texts. A sample summary is given in Figure 2 and a list of all test item titles
are given in Table 2. We suggested the following prompt but the participants were free to formulate
their own prompts as they saw fit.</p>
          <p>Write a text of about 500 words which covers the following items:</p>
          <p>
            The task had 35 registered teams. By the deadline three teams participated, with five experimental
conditions submitted. The models used are Poro [
            <xref ref-type="bibr" rid="ref40">56</xref>
            ] and Mistral [
            <xref ref-type="bibr" rid="ref41">57</xref>
            ] submitted by team Reindeer [
            <xref ref-type="bibr" rid="ref42">58</xref>
            ],
GPT-SW3 [
            <xref ref-type="bibr" rid="ref43">59</xref>
            ], a RAG-enhanced system based on the Command-R model submitted by team “Verbanex”
from Universidad Tecnológica de Bolívar, and a GPT 3.5-based baseline produced by the organizers.
Poro is a decoder-only model with a parameter count of 34 billion and 54 layers, trained on the LUMI
supercomputer with 1 trillion tokens for Finnish, English, and code. In testing, Poro has been found to be
reasonably competent in several other languages as well, due to the multilinguality of the Finnish data
set. The Mistral model was used as a comparison since it is better instruction trained for conversational
data. Poro and Mistral are open source models, freely available for use in experimentation. GPT-SW3
is based on the GPT-3 architecture and trained on the Berzelius supercomputer with 300B tokens for
Swedish, Norwegian, Danish, Icelandic, English, and code. GPT-SW3 is available for research purposes.
The Command-R series of models, built for RAG with a longer input context than many other models,
are tested for quality in several languages including English and are available for research purposes.
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.2. The “PAN AI News 2021” Dataset Collection</title>
          <p>For creating the second dataset collection, we first scraped 1,359 articles of major 2021 U.S. news
headlines from Google News, then generated summaries of each article, and finally re-generated
news-alike articles from these using nine large language models.</p>
          <p>We scraped Google News using the GNews Python library2 and Newspaper3k3 to download the
articles and extract the plain texts. We chose the year 2021 specifically as it predates the release of
GPT-3.5 so that we could be reasonably certain the articles were actually human-authored.</p>
          <p>Using the plain texts as input, we instructed GPT-4 Turbo to generate a bulleted summary of each
article as shown in Figure 3. To be able to generate convincing and high-quality articles with a high
similarity to the human texts, we also extracted (1) the article’s type (nine classes), (2) the target audience
(three classes), (3) the authors political stance (three classes), (4) the articles dateline, and (5) the names
and functions of directly quoted spokespersons, if any. The result was to be returned as JSON with
a predefined schema (the output was mostly valid, though some syntax and schema errors had to be
corrected later by hand).</p>
          <p>Given the so-generated summaries, we prompted several instruction-tuned downstream LLMs to
assume the role of a journalist from the respective source medium in writing an article of the given
type about the extracted key points. The article should have the same stance, target audience, and start
2https://github.com/ranahaani/GNews
3https://github.com/codelucas/newspaper</p>
          <p>Summary You are a news article and press release summarizer. Given an article, you summarize the key
points in 10 bullet points.</p>
          <p>Type You also classify the article type ("breaking news", "press release", "government agency
statement", "financial news", "opinion piece", "fact check", "celebrity news", "general reporting",
"speech transcript").</p>
          <p>Dateline Extract the dateline from the beginning of the article if one exists (e.g. "WASHINGTON " or
"May 28 (Reuters)").</p>
          <p>Quotes If spokespersons are cited verbatim, list their names, functions, and titles (if any).
Audience Determine the article’s target audience ("general public", "professionals", "children").</p>
          <p>Stance Classify whether the article’s stance is "left-leaning", "right-leaning", or "neutral".</p>
          <p>Structure Answer in structured JSON format (without Markdown formatting) like so:
{
}
"key_points": ["key point 1", "key point 2", . . . ],
"spokespersons": ["person1 (title, function)", . . . ],
"article_type": "article type",
"dateline": "dateline",
"audience": "audience",
"stance": "stance"
with the same dateline. Direct quotations from the originally cited spokespersons were to be included
as well, though we did not prescribe what those spokespersons were alleged to have said.</p>
          <p>
            In particular, we used the following LLMs for generating the articles:
1. GPT-3.5 Turbo [
            <xref ref-type="bibr" rid="ref44">60</xref>
            ]
2. GPT-4 Turbo [61]
3. Gemini Pro [62] (with temperatures of 0.6 and 0.9)
4. PaLM2 Text-Bison [63]
5. Meta Llama2 7B / 13B / 70B Chat [64]
6. Mistral 7B Instruct v0.2 [17]
7. Mixtral 8x7B Instruct v0.1 [65]
8. BLOOMZ 7B1 [66]
9. Qwen-1.5 72B Chat [67] (8-bit-quantized).
          </p>
          <p>Unless stated otherwise, we used the API default settings for GPT, Gemini, and PaLM. The remaining
models were retrieved from Huggingface [68] to run on our own infrastructure. Llama2 13B was used
with two diferent settings for contrastive decoding [ 69] ( = 0.1 and  = 0.6). We also planned to
include Falcon 7B and 40B Instruct [16], but were unable to get sensible articles out of it.</p>
          <p>In addition to the LLMs above, we also used a GPT-2 model fine-tuned on the Open-Instruct
dataset [70] and two Alpaca models [71] based on Llama2 7B and 13B as lower-quality baselines
(with a shortened prompt to fit the smaller input sizes).</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Text Pre- and Post-Processing</title>
        <p>We manually reviewed all of the extracted plain text from the human-written articles that we scraped
from Google News and removed any remaining artifacts, such as page footers, navigation fragments, or
ifgure captions.</p>
        <p>Similarly, we manually reviewed all LLM responses and thoroughly removed all obvious artifacts that
might give away the LLM authorship too easily. We removed typical LLM chat phrases such as “Sure, I’d
be happy to help.”, “Sorry, I cannot. . . ”, “Here’s your article:”, “Here are 10 paragraphs:” or “In this article, I will. . . ”,
markers and placeholders such as “[your name]”, ”[email]” or “[end of article]”, but also more complex
structural artifacts. Typical structural artifacts included numbering of individual paragraphs, excessive
use of bulleted lists, or newlines after the dateline. A very peculiar artifact many LLMs exhibited was
to append a bulleted list of “quotations” from the spokespersons to the end of the article instead of
incorporating them into the article itself. Some LLMs also liked giving (very “approximate”) word or
paragraph counts at the end. With the exception of Llama2, the open-source LLMs tended to produce
more such artifacts than the closed-source LLMs.</p>
        <p>As a final post-processing step, we truncated the (generally longer) human texts in the bootstrap
dataset to approximately the same length as the average LLM text. We did this by fitting a log-normal
distribution to the LLM text lengths and truncated the human texts accordingly by drawing from this
distribution. To avoid cutting texts in the middle of sentences, we used the drawn text lengths only as a
starting point to search for the nearest paragraph ending within a window of at most 200 characters.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Bootstrap and Main Test Split</title>
        <p>We initially published 1,087 of the original 1,359 human-authored articles, together with the re-generated
counterparts from 13 of the 16 LLMs. This “bootstrap” (i.e., training) dataset was released so the
participants could calibrate their systems. The texts from each LLM were collected in a separate
newline-delimited JSON files together with one file for all human texts.</p>
        <p>The remaining 272 human-authored articles and their counterparts were kept back for testing. This
resulted in 3,984 test cases (pairs of human and generated texts), which together form the main
(sub-)collection of the test set. The test set is formatted also as newline-delimited JSON, but in a single
ifle. Each line contains one of the original human texts and its corresponding machine generations as
a random-order pair. Texts within a pair were cut to the same number of words within a window of
65 words, trying to preserve full sentences if possible.</p>
        <p>Notably, all generated text variants from Llama2 13B with contrastive decoding and the
highertemperature variant of Gemini Pro were excluded completely from the “bootstrap” data to test how
robust detectors are to unseen models. In contrast, all variants from Alpaca 13B were added only to the
bootstrap dataset.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Test Data Variations for Robustness Evaluation</title>
        <p>
          To further test the robustness of the submitted systems against certain text modifications, we generated
multiple, unknown variants of the original test cases:
1. For German Text, we amended the prompt to generate German instead of English texts. This
was already part of the Main Test set but not of the initially published bootstrap dataset.
2. For Unicode, we replaced 15 % of the characters in (a) the machine texts and (b) both the human
and machine texts with Unicode lookalike characters.
3. For Cross-topic, we shufled the test case pairings to break the topic coherence.
4. For Contrastive Decoding, we used contrastive decoding [72] instead of top- / top- sampling;
5. For Short Text, we cropped the texts to 35 words; and
6. For Kaggle Paraphrase, we used the prompt from a previous Kaggle competition on LLM
detection [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to generate more faithful and direct paraphrases of the original articles:
The following is a human-written article. Now, please rewrite this article in
your writing style, also optimize sentence structures and correct grammatical
errors. You can appropriately add or remove content associated with the
article, but should keep the general meaning unchanged. Just return the
modified article.
        </p>
        <p>It is followed by the original text instead of the bullet-point summary.</p>
        <p>In total, we created 65 test set variants from the 13 source LLMs and the eight conditions given above.
Table 3 shows a systematic overview of all variants.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>At test time, participants were given pairs of human and LLM texts and had to calculate a score between
0 and 1 indicating which text was more likely to be human. Scores below 0.5 indicate that the left text
is human, and scores above 0.5 indicate that the right text is human. A score of exactly 0.5 could be
given to signal a non-decision. We borrowed this scoring scheme from previous iterations of the PAN
authorship verification task.</p>
      <p>We evaluate the overall efectiveness of the systems on each dataset variant as the arithmetic mean of
the following five metrics, which have also been used for evaluation in previous PAN shared tasks on
authorship verification (all with comparable 0–1 scales):
• Roc-Auc: The area under the Receiver Operating Characteristic curve.
• Brier: The complement of the Brier score, which is in our case equivalent to the mean squared
loss.
• C@1: A modified accuracy score that assigns non-answers (score = 0.5) the average accuracy of
the remaining cases.
• F1: The harmonic mean of precision and recall.
• F0.5u: A modified F 0.5 measure (precision-weighted F measure) that treats non-answers (score =
0.5) as false negatives.</p>
      <p>All metrics were corrected by discounting half a standard deviation, estimated on each dataset
individually, from the system’s scores with  − 1 degrees of freedom. This penalizes unstable systems
with widely varying scores on the individual dataset variants, and promotes systems that are more
robust to text obfuscation or other text modifications (even if their mean performance may be slightly
worse than that of other systems). We decided to use the macro average across all datasets because, even
though the datasets have diferent numbers of examples, we consider all datasets equally important as
performance indicators.</p>
      <sec id="sec-4-1">
        <title>4.1. PAN Submission Ranking</title>
        <p>We determined the final rank of each system by its macro-average mean efectiveness across all  = 70
dataset variants (including the five ELOQUENT submissions). Table 4 lists all systems sorted by their
rank. Twelve of the systems beat the best baseline (Binoculars with Falcon 7B). Six more beat the
second-best baseline (Binoculars with Mistral 7B). Figure 4 visualizes the mean score distribution of the
systems as a boxplot. It can be seen that weaker systems not only have a lower mean score, but also a
much higher variance over the diferent dataset variants. The bottom half of the systems even have
worse-than-random scores on some variants, while performing quite well on others.</p>
        <p>
          In comparison, Table 5 shows the system rankings only on the more homogeneous and thus easier
main test collection (without obfuscations and length restrictions, but with two German variants). In
both rankings, Tavan and Najafi [ 23] ranks first, but the middle and lower ranks difer substantially,
leading to an overall moderate to strong mean score rank correlation coeficient (Kendall’s  = .661,
 ≪ .001). When measuring the main collection only, more systems outperform the lower baselines,
but not the Binoculars baseline, which is actually beaten by fewer systems. Larger rank diferences
can also be seen among the top ranks. A notable example are Lorenz et al. [25], who rank ninth place
on the main collection, but third place on all dataset variants. Conversely, Valdez-Valenzuela and
Gómez-Adorno [
          <xref ref-type="bibr" rid="ref20">36</xref>
          ], who rank second on the main collection, rank only 14th on all dataset variants
aDtse Di
lA System
.05
.04
.10
.01
.02
.0
hSort
eTxt
nUicode
M
nUicode
H /
        </p>
        <p>M
aPrp
rhase E</p>
        <p>T</p>
        <p>UEN
O</p>
        <p>ci
Q
ros</p>
        <p>otp
L C
aMin
notrasive
C
nUicode
M
hSort
eTxt
nUicode
H /
rhase
aPrp
M</p>
        <p>ci
C
otp L
ors- E</p>
        <p>Q
O</p>
        <p>UENT Main
notrasive
C</p>
        <p>3
TP-SW
G</p>
        <p>
          Detection dificulty of all individual PAN test collections and the mean of all ELOQUENT
submissions in comparison. Right: Detection dificulty of individual ELOQUENT submissions. The dificulty of a
dataset is the inverse mean detection score of either the 10 best-performing systems or all systems. German texts
are not shown separately, as they were part of the main collection. A more detailed comparison of individual
dataset variants is given in Figure 6.
(although it should be noted that their short-text performance had to be estimated due to run failures).
This shows that the top performance of many systems does not generalize well to unexpected or
out-of-domain test cases. The top three systems [23, 24, 25], on the other hand, defy this trend and
generalize quite well. Interestingly, all three use diferent approaches. Tavan and Najafi [
23] built an
ensemble with Binoculars and a fine-tuned Mistral LLM, Huang et al. [ 24] use a BERT classifier trained
with PU loss [
          <xref ref-type="bibr" rid="ref35">51</xref>
          ] on sentence trigrams, and Lorenz et al. [25] employ a TF-IDF-based SVM classifier
with surprisingly strong results.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. PAN Dataset Variant Discussion</title>
        <p>by the inverse mean efectiveness score of the detection systems. Figure 6 shows in more detail the
dificulty of all individual dataset variants that are part of the collections (in this case, based only on the
scores of the ten best-performing systems).</p>
        <p>
          Which of the other dataset variants was particularly dificult varied from system to system. Truncating
texts to 35 words unsurprisingly yielded the most dificult dataset variant and all systems struggled
with it. Three of systems even failed to run at all on these texts due to hard-coded assumptions in the
software about the minimum length of a text. The efectiveness scores of these systems on the missing
dataset variants was therefore estimated using the mean efectiveness score of all other systems. The
three systems this afects [
          <xref ref-type="bibr" rid="ref19 ref20 ref24">35, 40, 36</xref>
          ] are marked with an asterisk in Table 4 and Figure 4.
        </p>
        <p>Already part of the main test collection, but on its own very challenging for most systems was the
set of German texts. Systems that relied on multilingual LLMs (such as Tavan and Najafi [ 23]) were
less afected, but many other systems struggled. Interestingly, Lorenz et al. [ 25], achieved near-perfect
scores on the German variant despite using a term-based SVM classifier with a fixed-length vocabulary
trained on English texts. This suggests that the classifier fitted more on the typical characteristics of a
human text rather than what makes a typical LLM text.</p>
        <p>Of course, more dificult than short texts or German texts on their own were short German texts.
Similarly, breaking the topic coherence (“cross-topic”) on its own was not very efective in making the
dataset more dificult, but combining this condition with the German texts produced one of the most
challenging variants (see Table Figure 6).</p>
        <p>Unicode lookalike character replacements in the machine texts were another dificult obfuscation
even for the best systems, whereas making the same replacements in the human texts as well equalized
the efect to a certain degree. The resulting dataset variant in which both texts of a pair were obfuscated,
was slightly less dificult than if only the machine text was obfuscated, yet still significantly more
dificult than the unobfuscated original texts.</p>
        <p>Contrastive decoding emerged as the easiest variant of all, even easier than the main dataset. However,
this should be taken with a large grain of salt, since we used only Llama2 13B with two diferent settings
for the hyper parameter α for creating these texts and Llama2 itself turned out to be the easiest LLM to
detect (also see Figure 6).</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. ELOQUENT Submission Ranking</title>
        <p>The ELOQUENT leader board is listed in Table 6. Of the four ELOQUENT submissions, only one
submission (GPT-SW3) managed to beat the GPT-3.5 baseline in terms of dataset dificulty and all five
(including the baseline) rank in the range of the PAN main test collection (see Figure 5 and Figure 6).
None of the submissions proved more efective than any of the other PAN conditions, such as Unicode
obfuscations or shortening the text length. This sobering result only goes to show how dificult to hide
the fingerprints of current-generation LLMs still are without making drastic text modifications.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This year’s shared task was designed as a builder-breaker task in collaboration with the ELOQUENT
lab. PAN participants would build systems to detect LLM-generated texts and ELOQUENT participants
would submit datasets trying to break the detectors.</p>
      <p>In total, we received LLM detection systems from 30 diferent participants and evaluated them on
70 diferent dataset variants, five of which were contributed by ELOQUENT (including one baseline). Of
the systems submitted to PAN, twelve beat the best of the provided baselines, demonstrating a robust
performance across all dataset variants. The dataset variants were designed to test the generalization
capabilities of the systems by exposing them to unseen and possibly unexpected conditions. We
implemented this by replacing characters with Unicode lookalikes, shortening the texts to 35 words,
generating German instead of English texts, and other similar means. Many of these variants proved
quite challenging for the systems, though the best systems were able to handle most of them well.</p>
      <p>ELOQUENT received four submissions, one of which beat the provided baseline. Unfortunately, none
of the submissions proved efective enough to reduce the detection performance of the PAN systems.</p>
      <p>We conclude that current LLMs are still easy to detect (some more than others) and their stylistic
ifngerprints are hard to hide. On the other hand, none of the detection systems managed to classify
all test cases correctly, which means that despite all, there is a margin of error and we can expect this
margin to increase with newer and better LLMs.</p>
      <sec id="sec-5-1">
        <title>Acknowledgments</title>
        <p>The “Voight-Kampf” Generative AI Authorship Detection Task @ PAN 2024 has been funded as part of
the OpenWebSearch project by the European Commission (OpenWebSearch.eu, GA 101070014).
(Eds.), Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), volume
3497 of CEUR Workshop Proceedings, 2023, pp. 2476–2491. URL: https://ceur-ws.org/Vol-3497/
paper-199.pdf.
[12] A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, T.
Goldstein, Spotting LLMs with Binoculars: Zero-shot detection of machine-generated text, arXiv
[cs.CL] (2024). arXiv:2401.12070.
[13] J. Su, T. Y. Zhuo, D. Wang, P. Nakov, DetectLLM: Leveraging log rank information for zero-shot
detection of machine-generated text, arXiv [cs.CL] (2023). arXiv:2306.05540.
[14] E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, C. Finn, DetectGPT: Zero-shot machine-generated
text detection using probability curvature, International Conference on Machine Learning 202
(2023) 24950–24962. doi:10.48550/arXiv.2301.11305. arXiv:2301.11305.
[15] G. Bao, Y. Zhao, Z. Teng, L. Yang, Y. Zhang, Fast-DetectGPT: Eficient zero-shot
detection of machine-generated text via conditional probability curvature, arXiv [cs.CL] (2023).
arXiv:2310.05130.
[16] E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, D. Hesslow, J. Launay, Q.
Malartic, D. Mazzotta, B. Noune, B. Pannier, G. Penedo, The Falcon series of open language models,
arXiv [cs.CL] (2023). arXiv:2311.16867.
[17] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand,
G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril,
T. Wang, T. Lacroix, W. E. Sayed, Mistral 7B, arXiv [cs.CL] (2023). arXiv:2310.06825.
[18] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring
the limits of transfer learning with a unified text-to-text transformer, arXiv [cs.LG] (2019).
arXiv:1910.10683.
[19] D. Sculley, C. E. Brodley, Compression and machine learning: A new perspective on feature space
vectors, in: Data Compression Conference (DCC’06), IEEE, 2006, pp. 332–341. doi:10.1109/dcc.
2006.13.
[20] O. Halvani, C. Winter, L. Graner, On the usefulness of compression models for authorship
verification, in: Proceedings of the 12th International Conference on Availability, Reliability and
Security, volume Part F1305, ACM, New York, NY, USA, 2017. doi:10.1145/3098954.3104050.
[21] M. Koppel, J. Schler, Authorship verification as a one-class classification problem, in: Twenty-first
international conference on Machine learning - ICML ’04, ACM Press, New York, New York, USA,
2004, pp. 489–495. doi:10.1145/1015330.1015448.
[22] J. Bevendorf, B. Stein, M. Hagen, M. Potthast, Generalizing unmasking for short texts, in:
Proceedings of the 2019 Conference of the North, Association for Computational Linguistics,
Stroudsburg, PA, USA, 2019, pp. 654–659. doi:10.18653/v1/n19-1068.
[23] E. Tavan, M. Najafi, Marsan at PAN: BinocularLLM and Fusing Binoculars’ Insight with the
Proficiency of Large Language Models for Cutting-Edge Machine-Generated Text Detection, in:
G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024
Conference and Labs of the Evaluation Forum, CEUR-WS.org, 2024.
[24] J. Huang, Y. Chen, M. Luo, Y. Li, Generative AI Authorship Verification Of Tri-Sentence Analysis
Base On The Bert Model, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. de Herrera (Eds.),
Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum, CEUR-WS.org, 2024.
[25] L. Lorenz, F. Z. Aygüler, F. Schlatt, N. Mirzakhmedova, BaselineAvengers at PAN 2024:
OftenForgotten Baselines for LLM-Generated Text Detection, in: G. Faggioli, N. Ferro, P. Galuščáková,
A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 - Conference and Labs of the Evaluation
Forum, CEUR-WS.org, 2024.
[26] M. Guo, Z. Han, H. Chen, J. Peng, A Machine-Generated Text Detection Model Based on Text
Multi-Feature Fusion, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. de Herrera (Eds.), Working
Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum, CEUR-WS.org, 2024.
[27] Z. Lin, Z. Han, L. Kong, M. Chen, S. Zhang, J. Peng, K. Sun, A Verifying Generative Text Authorship
Model With Regularized Dropout, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. de Herrera
(Eds.), Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum, CEUR-WS.org,</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Sarvazyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Á.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Franco-Salvador, Supervised machine-generated text detectors: Family and scale matters, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and Interaction - 14th
          <source>International Conference of the CLEF Association</source>
          , Springer,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bafour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Crossley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Holbrook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Demkin</surname>
          </string-name>
          , Llm - detect
          <source>ai generated text</source>
          ,
          <year>2023</year>
          . URL: https://kaggle.com/competitions/llm-detect
          <article-title>-ai-generated-text.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Molla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>Overview of the 2023 ALTA shared task: Discriminate between human-written and machine-generated text</article-title>
          , in: S.
          <string-name>
            <surname>Muresan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Casey</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>David</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Nina</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Koji</surname>
          </string-name>
          , E. Erik, U. Stefan (Eds.),
          <source>Proceedings of the 21st Annual Workshop of the Australasian Language Technology Association, Association for Computational Linguistics</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>148</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarvazyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Á.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Franco-Salvador</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Overview of AuTexTification at IberLEF 2023:
          <article-title>Detection and attribution of machine-generated text in multiple domains</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          )
          <fpage>275</fpage>
          -
          <lpage>288</lpage>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2309. 11285. arXiv:
          <volume>2309</volume>
          .
          <fpage>11285</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mansurov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsvigun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. M.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mahmoud</surname>
          </string-name>
          , G. Puccetti, T. Arnold, SemEval
          <article-title>-2024 Task 8: Multidomain, Multimodel and Multilingual MachineGenerated Text Detection</article-title>
          ,
          <source>in: Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>2057</fpage>
          -
          <lpage>2079</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M. R.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>Overview of the PAN/CLEF 2015 evaluation lab</article-title>
          , in: J.
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Savoy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Pinel-Sauvagnat</surname>
            ,
            <given-names>G. J. F.</given-names>
          </string-name>
          <string-name>
            <surname>Jones</surname>
          </string-name>
          , E. SanJuan, L. Cappellato, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction - 6th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2015</year>
          , Toulouse, France, September 8-
          <issue>11</issue>
          ,
          <year>2015</year>
          , Proceedings, volume
          <volume>9283</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2015</year>
          , pp.
          <fpage>518</fpage>
          -
          <lpage>538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ghanem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giachanou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , I. Markov,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M. R.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Specht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2020:
          <article-title>Authorship verification, celebrity profiling, profiling fake news spreaders on twitter, and style change detection</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction - 11th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2020</year>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , Proceedings, volume
          <volume>12260</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2020</year>
          , pp.
          <fpage>372</fpage>
          -
          <lpage>383</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. L. D. la Peña</given-names>
            <surname>Sarracén</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          , I. Markov,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wolska</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2021:
          <article-title>Authorship verification, profiling hate speech spreaders on twitter, and style change detection</article-title>
          , in: K. S. Candan,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Larsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maistro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction - 12th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2021</year>
          ,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          ,
          <source>September 21-24</source>
          ,
          <year>2021</year>
          , Proceedings, volume
          <volume>12880</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2021</year>
          , pp.
          <fpage>419</fpage>
          -
          <lpage>431</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Manjavacas</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Markov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <source>Overview of the Cross-Domain Authorship Verification Task at PAN</source>
          <year>2021</year>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Maistro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          (Eds.),
          <source>Working Notes Papers of the CLEF 2021 Evaluation Labs</source>
          , volume
          <volume>2936</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2021</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2936</volume>
          /paper-147.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kestemont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kredens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pezik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Heini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <source>Overview of the Authorship Verification Task at PAN</source>
          <year>2022</year>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          , M. Potthast (Eds.),
          <source>CLEF 2022 Labs and Workshops</source>
          , Notebook Papers, volume
          <volume>3180</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2022</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3180</volume>
          /paper-184.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kredens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pezik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Heini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <source>Overview of the Authorship Verification Task at PAN</source>
          <year>2023</year>
          , in: M.
          <string-name>
            <surname>Aliannejadi</surname>
            , G. Faggioli,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>M. Vlachos</given-names>
          </string-name>
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Abburi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pudota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Veeramani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Bowen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          , Team Deloitte at PAN:
          <article-title>Generative AI Text Detection</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>P.</given-names>
            <surname>Miralles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Martín</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Camacho</surname>
          </string-name>
          , Ensembling Normalized Log Probabilities, in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Yadagiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kalita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ranjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bostan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Toppo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pakray</surname>
          </string-name>
          ,
          <article-title>Team cnlp-nits-pp at PAN: Leveraging BERT for Accurate Authorship Verification: A Novel Approach to Textual Attribution</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lv</surname>
          </string-name>
          , Y. Han,
          <string-name>
            <surname>L</surname>
          </string-name>
          . Kong,
          <article-title>Meta-Contrastive Learning for Generative AI Authorship Verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>G.</given-names>
            <surname>Gritsai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Boyeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Grabovoy</surname>
          </string-name>
          ,
          <article-title>Team ap-team at PAN: LLM Adapters for Various Datasets</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>H.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            .
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          , Y. Han,
          <article-title>Enhancing Human-Machine Authorship Discrimination in Generative AI Verification Task with BERT and Augmented Data</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>L.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          , L. Ma, J. Ruan,
          <article-title>BLGAV: Generative AI Author Verification Model Based on BERT and BiLSTM</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>B.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yan</surname>
          </string-name>
          , Y. Han,
          <article-title>Author authentication of generative AI based on BERT by regularization method</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>A.</given-names>
            <surname>Valdez-Valenzuela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gómez-Adorno</surname>
          </string-name>
          ,
          <article-title>Team iimasnlp at PAN: Leveraging Graph Neural Networks and Large Language Models for Generative AI Authorship Verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <article-title>Token Prediction as Implicit Classification for Generative AI Authorship Verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <article-title>Integrating Dual BERT Models and Causal Language Models for Enhanced Detection of Machine-Generated Texts</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>W.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grieve</surname>
          </string-name>
          ,
          <article-title>Authorial Language Models For AI Authorship Verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>R.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <article-title>A model fusion approach for generative AI authorship verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>P.</given-names>
            <surname>Petropoulos</surname>
          </string-name>
          , V. Petropoulos,
          <article-title>RoBERTa and Bi-LSTM for Human vs AI generated Text Detection</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Zhao,</surname>
          </string-name>
          <article-title>BertT: A Hybrid Neural Network Model for Generative AI Authorship Verification</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Voight-Kampf Generative AI Authorship Verification</surname>
          </string-name>
          <article-title>Based on T5</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , L. Kong,
          <source>AI Authorship Verification Based On Deberta Model</source>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          , L. Ma, BCAV:
          <string-name>
            <given-names>A Generative</given-names>
            <surname>AI</surname>
          </string-name>
          <article-title>Author Verification Model Based on the Integration of Bert and CNN</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          , G. Niu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Generative AI Authorship Verification based on ChatGLM</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          , L. Kong,
          <article-title>AI Text Detection Method Based on Perplexity Features with Strided Sliding Window</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>K.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Voight-Kampf Generative AI Authorship</surname>
          </string-name>
          <article-title>Verification based on Contrastive Learning and Domain Adaptation</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sheykhlan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abdoljabbar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mahmoudabad</surname>
          </string-name>
          ,
          <article-title>Team karami-kheiri at PAN: Enhancing MachineGenerated Text Detection with Ensemble Learning Based on Transformer Models</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>G.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <article-title>Team lm-detector at PAN: Can NLI be an Appropriate Approach to MachineGenerated Text Detection</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Multiscale positive-unlabeled detection of ai-generated texts</article-title>
          ,
          <source>CoRR abs/2305</source>
          .18149 (
          <year>2023</year>
          ). URL: https://doi.org/10.48550/arXiv. 2305.18149. doi:
          <volume>10</volume>
          .48550/ARXIV.2305.18149.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T. Liu, R-drop:
          <article-title>Regularized dropout for neural networks</article-title>
          , in: M.
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>Y. N.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Vaughan</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems</source>
          <year>2021</year>
          ,
          <article-title>NeurIPS 2021</article-title>
          , December 6-
          <issue>14</issue>
          ,
          <year>2021</year>
          , virtual,
          <year>2021</year>
          , pp.
          <fpage>10890</fpage>
          -
          <lpage>10905</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Achiam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <article-title>On first-order meta-learning algorithms</article-title>
          ,
          <source>arXiv [cs.LG]</source>
          (
          <year>2018</year>
          ). arXiv:
          <year>1803</year>
          .02999.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>S.</given-names>
            <surname>Thite</surname>
          </string-name>
          , Llm - detect
          <source>ai generated text dataset</source>
          ,
          <year>2023</year>
          . URL: https://www.kaggle.com/datasets/ sunilthite/llm-detect
          <article-title>-ai-generated-text-dataset.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kłeczek</surname>
          </string-name>
          ,
          <source>Daigt v2 train dataset</source>
          ,
          <year>2023</year>
          . URL: https://www.kaggle.com/datasets/thedrcat/ daigt-v2
          <string-name>
            <surname>-</surname>
          </string-name>
          train-dataset.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [56]
          <string-name>
            <given-names>R.</given-names>
            <surname>Luukkonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Burdge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Zosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Talman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Komulainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hatanpää</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sarlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <article-title>Poro 34b and the blessing of multilinguality</article-title>
          ,
          <source>arXiv preprint: 2404</source>
          .
          <year>01856</year>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [57]
          <string-name>
            <given-names>A. Q.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sablayrolles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bamford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Chaplot</surname>
          </string-name>
          , D. d. l. Casas,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bressand</surname>
          </string-name>
          , G. Lengyel,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Saulnier</surname>
          </string-name>
          , et al.,
          <source>Mistral 7b, arXiv preprint arXiv:2310.06825</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [58]
          <string-name>
            <given-names>V.</given-names>
            <surname>Neralla</surname>
          </string-name>
          , S. Bijl de Vroe,
          <source>Evaluating Poro-34B-Chat and Mistral-7B-Instruct-v0</source>
          .
          <article-title>1: LLM System Description for ELOQUENT at CLEF 2024</article-title>
          , in: G. Faggioli,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuščáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.), Working Notes of CLEF 2024 -
          <article-title>Conference and Labs of the Evaluation Forum, CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [59]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Cuba</given-names>
            <surname>Gyllensten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stollenwerk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Öhman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Isbister</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gogoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Carlsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Casademont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          , GPT-SW3:
          <article-title>An autoregressive language model for the Scandinavian languages</article-title>
          ,
          <source>in: Proceedings of the 2024 Joint International Conference on Computational Linguistics</source>
          ,
          <article-title>Language Resources and Evaluation (LREC-COLING</article-title>
          <year>2024</year>
          ), Torino, Italia,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [60]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Ramesh,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>