<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The Impact of Prompts on Zero-Shot Detection of AI-Generated Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>KaitoTaguchi</string-name>
          <email>gu@inf.kyushu-u.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yujie Gu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kouichi Sakurai</string-name>
          <email>sakurai@inf.kyushu-u.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kyushu University</institution>
          ,
          <addr-line>Fukuoka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Skydisc Inc.</institution>
          ,
          <addr-line>Fukuoka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years, there have been significant advancements in the development of Large Language Models (LLMs). While their potential for misuse, such as generating fake news and committing plagiarism, has posed significant concerns. To address this issue, detectors have been developed to evaluate whether a given text is human-generated or AI-generated. Among others, zero-shot detectors stand out as efective approaches that do not require additional training data and are often likelihood-based. In chat-based applications, users commonly input prompts and utilize the AI-generated texts. However, zero-shot detectors typically analyze these texts in isolation, neglecting the impact of the original prompts. It is conceivable that this approach may lead to a discrepancy in likelihood assessments between the text generation phase and the detection phase. So far, there remains an unverified gap concerning how the presence or absence of prompts impacts detection accuracy for zero-shot detectors. In this paper, we introduce an evaluative framework to empirically analyze the impact of prompts on the detection accuracy of AI-generated text. We assess various zero-shot detectors using both white-box detection, which leverages the prompt, and black-box detection, which operates without prompt information. Our experiments reveal the significant influence of prompts on detection accuracy. Remarkably, compared with black-box detection without prompts, the white-box methods using prompts demonstrate a significant increase in AUC across all zero-shot detectors tested, which calls for attention to the impact of prompts on zero-shot detectors. Code is availhatbtlep:s://github.com/kaito25atugich/Detec. tor</p>
      </abstract>
      <kwd-group>
        <kwd>zero-shot detector</kwd>
        <kwd>AI-generated text</kwd>
        <kwd>prompt</kwd>
        <kwd>LLM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>development of Large Language Models (LLMs) 1[, 2, 3],</p>
      <p>illustrated in Table1.</p>
      <p>Recent years have seen significant advancements in the detection stages. A summary of zero-shot detectors is
misuse.
and their practical applications have become widespread. In this paper, we assess to what extent this phenomenon
Meanwhile, their potential misuse have raised significant afects likelihood-based zero-shot detectors. First, we
proconcerns. In particular, the generation of fake news and pose two methods for detecting AI-generated text using
plagiarism using LLMs is a notable issue. Detectors thatzero-shot detectors: white-box detection, which
leverevaluate whether a given text is human-generated or AIa-ges the prompts used to generate the text, and black-box
generated serve as a defense mechanism against such detection, which detects AI-generated text without
relying on a prompt. Next, we conduct extensive experiments
ifed into three categories: a zero-shot detector leveragingisting zero-shot detectors in black-box detection.</p>
      <p>Detectors for AI-generated text can be broadly classia-nd demonstrate a decrease in detection accuracy for
exstatistical properties4,[5, 6, 7, 8, 9, 10, 11], a detector
employing supervised learning [12, 13, 14, 15], and a
detector utilizing watermarking16[, 17].</p>
      <p>Zero-shot detectors, such as DetectGP5T],[which do
not require additional training, are designed in many</p>
      <p>Our results show a significant diference in the
performance of zero-shot detectors for AI-generated text with
and without prompts, highlighting the need to consider
the impact of prompts on these detectors.</p>
      <p>These results further point out that likelihood-based
we usually input prompts and utilize the generated outd-etectors due to their higher sampling rate.
methods using likelihood-based scores. In other words, zero-shot detectors face challenges for practical use.
Adthe zero-shot detection is carried out by replicating thdeitionally, the experimental results demonstrate that fast
likelihood at the generation phase. When using LLMs, zero-shot detectors are more robust compared to other
put. However, at the detection phase, it is anticipated
that reproducing likelihood becomes challenging due
to the absence of the contextual information provided2. Related work
nEvelop-O
crafting of prompts with malicious intent to deliberately text and updating the content of the instructions to
rereduce detection accuracy. In contrast, the second cadtu-ce detection accuracy.
egory encompasses research that employs tasks with Kumarage et al. [21] proposed an attack named Soft
benign prompts, devoid of malicious intent. Prompt, which generates a vector using reinforcement
learning to induce misclassification by detectors. This
2.1. Malicious prompts Soft Prompt vector is then used as input for DetectGPT
and RoBERTa-based detectors1[2], demonstrating a
deFirst, we delve into studies that specifically concentratecrease in detection accuracy21[].
on the deliberate creation of malicious prompts.</p>
      <p>In [19], Koike et al. proposed OUTFOX, utilizing in- 2.2. Benign prompts
context learning with the problem
statemen,thumangenerated text , and AI-generated text . By construct- We review cases involving tasks with benign prompts.
ing prompts such as “  ∈  → ℎ  ∈  is the correct label Liu et al. conducted experiments using the CheckGPT
by humans, and   ∈  →   ∈  is the correct label model, an approach based on supervised learning. Their
by AI,” they aim to generate text for a given problemfindings indicate that when using diferent prompts,
alstatement in such a way that the generated text aligtnhsough all surpass 90%, there is an experimental
demonwith human-authored content. This approach makes thestration of approximately a 7% decrease in detection
acdetection of artificially generated content challenging. curacy [15].</p>
      <p>Shi et al. conducted an attack on OpenAI’s Detec- Dou et al. [14] performed experiments envisioning
tor [22] by employing an Instructional Prompt, confirm- the utilization of LLMs by students. In their study, they
ing a decrease in detection accuracy18[]. The Instruc- demonstrated a decrease in DetectGPT’s detection
accutional Prompt involves adding a reference te xt and racy when prompts were employed.
an instructional tex t with characteristics that reduce Hans et al. [8] pointed out the dificulty in
reproducthe detection accuracy to the original in pu,tthereby ing likelihoods depending on the presence or absence of
undermining the detection accuracy. prompts, using unique prompts like “Write about a
capy</p>
      <p>In [20], Lu et al. proposed SICO, a method that low-bara astronomer.” In response to the capybara problem,
ers detection accuracy by instructing the model withitnhey proposed Binoculars.
prompts to mimic the writing style of human-authored We assume performing benign tasks such as
summarization. Therefore, unlike malicious prompt attacks,3.2.2. Entropy
there is no need to deliberately choose prompts that
would lower accuracy using the detector when
constructing prompts, nor is there a requirement to collect pairs
of data for in-context learning.</p>
      <p>On the other hand, Dou et al.1[4] experimentally ing the average across each context.</p>
      <p>Entropy is a method that utilizes the entropy of the
vocabulary for detection. The formula is shown in(3). Entropy
is calculated using the likelihood of the vocabulary,
takdemonstrated unintended decreases in detection
accuracy. However, they did not delve into why the accuracy
decreases or make references to other likelihood-based
zero-shot detectors. Additionally, Hans et a8l.][did not
provide specific verification regarding the impact of a
detector knowing or not knowing the prompt on
detection accuracy. Therefore, the resilience of Binoculars
to changes in likelihood due to prompts has not been
adequately assessed. The supervised learning based
approach [15] is excluded from our experiments in this
context.</p>
      <p>−1
 − 1 =2 =1
Entropy=
∑
∑   
(| &lt; ) log</p>
      <p>(| &lt; ). (3)
3.2.3. Rank
Rank is a method that utilizes the order of likelihood
magnitude of tokens in the vocabulary when sorted. The
formula is presented in(4). Rank is the average position
of tokens constituting a given text. The function
is a function that sorts the given array in descending
order, and 
is a function that, given an array and an
prompts unintentionally leads to a decrease in accuracy
tasks such as summarization, the presence or absence of the given array.
when using likelihood-based zero-shot detectors.</p>
      <p>In this study, we demonstrate that even in ordinaryelement as input, returns the index of the element within
DetectGPT=
 =̃
  ̃ =
1
 =1

1
 − 1 =1</p>
      <p>∑ log</p>
      <p>( ̃ )
∑(log</p>
      <p>( ̃ ) − )̃ 2
log    () −  ̃
  ̃
(5)

=1
 ( 1,  2, … ,   ) = ∏  ( 
| 1, … ,  −1 )</p>
      <p>(1)</p>
      <sec id="sec-2-1">
        <title>3.2.1. Log-Likelihood</title>
        <p>The log-likelihood is a method that utilizes the likelihood
of tokens composing a text for detection. The formula is
presented in(2). The log-likelihood is the average of the
log-likelihoods of tokens constituting a given text.
where
1</p>
        <p>− 1 =2
Log-likelihood =
∑ log   
(  | &lt; ).</p>
        <p>(2)</p>
        <sec id="sec-2-1-1">
          <title>3.2. Existing zero-shot detectors</title>
          <p>We provide a brief introduction to existing zero-shotof the obtained replacement texts is then compute(5d).
detectors, summarized in Table 1. Here ,  refers to the represents the score, calculating the diference between
language model utilized for detection. The vocabular y
is composed of  tokens. The input tex tis composed
the log-likelihood of the original text and the average
log-likelihood of the acquired replacement texts. It is
token sequence from 1 to −1 is denoted as &lt; .
of  tokens, represented a s = { 1,  2, … ,   }, and the permissible to standardize by dividing by the standard
deviation of the log-likelihood of the replacement texts.</p>
          <p>If the score is above the threshold, it is deemed to be</p>
          <p>AI-generated text.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Preliminary</title>
      <sec id="sec-3-1">
        <title>3.1. Language model</title>
        <p>A model that captures the probability of generating wordstext generation, whereas humans create text
indepenor sentences is referred to as a language model. Le t</p>
        <p>dently of likelihood. DetectGPT focuses on this
pherepresent the vocabulary. The language model for a word nomenon and posits a hypothesis that by rewriting
ceris defined by the following ( 1).
sequence of length , denoted as 1,  2, … ,   where   ∈  ,
tain words, the likelihood of the text decreases for
AIgenerated content and can either increase or decrease for
rank =</p>
        <p>−1
 − 1 =2

∑ ((
log</p>
        <p>(  | &lt; )),   ). (4)</p>
        <sec id="sec-3-1-1">
          <title>3.2.4. DetectGPT</title>
          <p>The language model aims to maximize likelihood during
human-generated content5][.</p>
          <p>The overview of DetectGPT is presented in Figure
1. The replacement process is achieved by utilizing a
mask model   , such as T5 [24], on some of the words
contained in the given te xt. This operation is repeated
for a total o f iterations, and the average log-likelihood
Replace words
calculate the likelihood



・
・
・
mask
model
 
・
・
・
 ǁ1
and  ̃ ∼   (  ) represent the mean, sample variance, and</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.2.5. FastDetectGPT</title>
          <p>On the other hand, NPR, like DetectGPT, performs the
substitution of words in the te xtimes. It takes the ratio
of the average log-rank of the obtained substituted texts
to the log-rank of the original text. This is defined in(7).</p>
          <p>In [6], Bao et al. highlighted challenges in DetectGPT’s
use of diferent models for substitution and score
calculation, as well as the cost-related aspect of requiring
model access for each substitution iteration. In response,
FastDetectGPT is a modified detector that reduces access3.2.7. Binoculars
  =
1 ∑=1 log   ( ̃ )
log   ()
likelihood, whereas human-generated text is not.”
“AI-generated text is likely to be around the maximumillustrated in Figure3.
to the model, addressing the cost issue while enablingHans et al. proposed Binoculars, a detection method
substitutions. Although the methodology involves setu-tilizing two closely related language models,
Falconting hypotheses similar to DetectGPT, there is no funda7-b [26] and Falcon-7b-instruct, by employing a metric
mental change. It still operates on the assumption thactalled cross-perplexity 8[]. The overall framework is
Let the first model be denoted a s 1 (such as
Falcon</p>
          <p>We present the overall architecture of FastDetectGP7Tb), and the second model as 2 (like Falcon-7b-instruct).
in Figure 2. In FastDetectGPT, the substitution process iIsn this case, using 1, we calculate the log perplexity as
(7)
a mask model. Similar to the detection model, it utilizes
replaced with an alternative method that does not rely osnhown in (8).
an autoregressive model, and   and   
The substitution for the-th word involves randomly
extracting a word from the next-word list, considering</p>
          <p>can be the same.
the substituted te x,̃tand by conducting sampling
durwords, performing this substitution times results in product.
ing word selection, the replacement process generat es
substitution texts in a single access.
log  -</p>
          <p>The subsequent score calculation is omitted as it
follows the same procedure as DetectGPT.
3.2.6. LLR &amp; NPR
represents the rank when usin g   .</p>
          <p>LLR (Likelihood Log-Rank ratio) and NPR (Normalized
perturbed logrank) are classical log-rank enhancement
techniques proposed by Su et al.7[]. Both methods have
simple configurations. LLR literally takes the ratio of
log-likelihood to log-rank, as expressed in(6). Here,  
 = −


∑=1 log</p>
          <p>(  | &lt; )
∑=1 log   (  | &lt; )
the context up to th( e− 1) -th word in the input text, Next, using 1 and  2, we calculate the cross-perplexity,
and replacing the word with the chosen one. In otheras shown in (9). Here, the symbol⋅ represents the dot

1
 =1
log  
 1() = −
∑ log( 1(  | &lt; ))</p>
          <p>(8)
 1, 2() =
− 1
 
∑
 =1 =1</p>
          <p>∑  1(| &lt; ) ⋅ log( 2(| &lt; )) (9)
The score in Binoculars is determined by1(0).</p>
          <p>1, 2() =</p>
          <p>log  
log  - 
 1()
 1, 2()</p>
          <p>(10)</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Proposal</title>
      <p>detection method.</p>
      <p>In this study, we propose a detection flow to investigate
the impact of prompts on likelihood. Before
present(6) ing the experimental setup, we introduce an additional</p>
      <p>Replace words calculate the likelihood</p>
      <p>LLM</p>
      <p>LLM
5. Experiment
・  ǁ2
・
・
 ǁ1
 2
model. In this research, aiming for cost reduction, we
emWord replacements in NPR are performed using a masked 5.1. Configuration
ploy FastNPR, a method that replaces word replacementsTo begin, we utilize the GPT2-XL2[3] as the detection
with sampling, akin to FastDetectGPT.</p>
      <sec id="sec-4-1">
        <title>4.2. Detection methods</title>
        <p>model, excluding Binoculars. Due to GPU constraints,
Binoculars employs the pre-trained and instruct-tuned
Phi1.5 [27] instead of Falcon.</p>
        <p>For DetectGPT and NPR, we generate five replacement
We explain the detection methodology. Letrepresent sentences for 10% of the entire text, while the Fast series
the text to be detected, and isf an AI-generated text, let generates 10,000 replacement sentences. T5-Large24[]
 denote the prompt used for its generation. Detectioins used for word replacement in DetectGPT and NPR,
can be categorized into two patterns: Black-box detectiownhile the Fast series employs the GPT2-XL, the same
and White-box detection. An overview is presented indetection model. Also, we use the XSum datase2t8][.
detector having knowledge of prompt information. For5.2. Result</p>
        <p>following sentences, please? text”.
a consistent decrease of 0.1 or more across all methods,
tokens, making replacement not always feasible,
espethe use of a masked language model to replace plausible on AUC(NPR)
of replacement sentences. DetectGPT and NPR requireEfect of Substitution Rate(SR) and Sample Size(SS) Variation
prompt</p>
        <p>Use for calculating score
(1+1 equals 2.| What is 1+1? )
Particularly in recent years, there is a trend toward
practical applications, emphasizing high true positive rates at
low false positive rates, suggesting that at least an AUC in
the late 0.9s would be necessary3[0, 8]. Furthermore, the
lack of improvement in detection accuracy with
DetectGPT and NPR may be attributed to the limited number
of substitutable tokens.
placed, and the sample size, representing the numberTable 4
cially for higher replacement percentages. Therefore, we
primarily vary the replacement ratio in the Fast series to
conduct the investigation.</p>
        <p>The results for DetectGPT are presented in Tab3le,
and the results for NPR are shown in Table4. From these
results, it is evident that increasing the replacement ratio
and sample size helps mitigate the decrease in detection
accuracy. This observation is similar to Chakraborty et
al.’s assertion that increasing the sample size can enable
detection if the distribution slightly difers29[].</p>
        <p>However, in our validation, the improvement in
accuracy plateaus at around 10 samples, reaching a maximum
AUC of approximately 0.8, which is not considered high.</p>
        <p>Method
FastDetectGPT
FastDetectGPT
FastDetectGPT
FastDetectGPT
FastDetectGPT
FastDetectGPT
FastDetectGPT
DetectGPT
DetectGPT
DetectGPT
DetectGPT
DetectGPT</p>
        <p>SR
10%
20%
100%
10%
20%
100%
100%
10%
20%
30%
10%
30%</p>
        <p>SS
5
5
5
10
10
10
5
5
5
10
10
10000</p>
        <p>AUC
0.640
0.697
0.779
0.704
0.739
0.821
0.819
0.453
0.522
0.490
0.446</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Limitation and future work</title>
      <sec id="sec-5-1">
        <title>6.5. Relationship with watermarking</title>
        <p>6.1. Hypotheses for zero-shot detectors Watermarking techniques utilize statistical methods for
verification [ 16]. Since these methods are based on
likeWhile our investigation has focused solely on prompts, lihood during both generation and verification, a
failsimilar phenomena could potentially be observed with ure to reproduce likelihood during the verification stage
other factors. For instance, variations in Temperature ormay lead to a decrease in accuracy. On the other hand,
Penalty Repetition between the generation and detectioronbust watermarking techniques against paraphrase
atstages might introduce diferences in the selected tokens,tacks have emerged 1[7]. These methods may exhibit
making detection challenging based on likelihood. Gen- robustness against prompts as well.
eralizing these observations, we hypothesize that any
act that fails to replicate the likelihood during language6.6. Towards resilient zero-shot detectors
generation could undermine the detection accuracy of
zero-shot detectors relying on likelihood from next-wordCurrently, many methods perform likelihood-based
deprediction. tection. Combining these methods with other
sophisticated techniques may lead to more robust detection.
6.2. Tasks One such approach is Intrinsic Dimension 1[1].
Intrinsic Dimension refers to the minimum dimension needed
While our investigation has focused on summary textto represent a given text. Tulchinskii et al. propose a
generation, there are several other potential tasks to conde-tector based on Persistent Homology to estimate the
sider, such as paraphrase generation, story generation, Intrinsic Dimension and use it as a score. However, this
and translation text generation. It is plausible that dme-ethod requires a constant length of text and was not
tection accuracy could also decrease in these commonapplicable in our experiment. It would be interesting to
tasks. Since these tasks may be utilized without malie-xplore the application of this method in experiments
cious intent, it is crucial to conduct similar evaluationisnvolving longer texts.
for them. Approaches utilizing representations obtained with
masked language models, including Intrinsic Dimension,
6.3. Number of parameters calculate likelihood in a diferent way from the detectors
used in our experiment, which are based on
autoregresIn this study, each detection method utilized a languagesive language models. Combining these elements may
model of approximately 1 billion parameters. It would be lead to the development of a more robust zero-shot
deof interest to investigate whether increased robustnestsector.
can be observed when experimenting with larger
language models. Conversely, there are experimental
studies that have demonstrated the ability of smaller language7. Conclusion
models to achieve a higher likelihood for AI-generated
texts across a broader range of language models3[1]. In this paper, we experimentally demonstrated a
signifiConsidering these findings, conducting experiments with cant gap in the detection of AI-generated text with and
without prompts for likelihood-based zero-shot
detecsmaller language models and verifying if there are
diferences in robustness could also provide valuable insights. tors. These findings call for attention to the impact of
prompts on enhancing zero-shot detectors in practical
applications.</p>
      </sec>
      <sec id="sec-5-2">
        <title>6.4. Relationship with supervised learning detectors</title>
        <p>Even when using supervised learning, it has been noted
that generated text from prompt-based tasks may exhibit[1] OpenAI. (2023). GPT-4 Technical Report, arXiv.
decreased detection accuracy15[]. However, there is a [2] Microsoft. Microsoft Copilot, Retrieved October 31,
possibility that these models could be more robust com- 2023, from https://adoption.microsoft.com/ja-jp/
pared to zero-shot detectors. For instance, RADA1R3][ copilot/.
achieved an AUC of 0.939 in the task used in this experi- [3] Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac,
ment. In comparison, the RoBERTa-large detecto1r2[] J. B., Yu, J., ... &amp; Ahn, J. (2023). Gemini: A
had an AUC of 0.767. This suggests that robust detectors family of highly capable multimodal models.
against paraphrase attacks might demonstrate similarly arXiv:2312.11805.
robust results in other tasks. [4] Gehrmann, S., Strobelt, H., &amp; Rush, A. (2019). GLTR:
Statistical detection and visualization of generated
text. In Proceedings of the 57th Annual Meeting &amp; Hsieh, C. J. (2023). Red teaming language model
of the Association for Computational Linguistics: detectors with language models. arXiv:2305.19713.</p>
        <p>System Demonstrations (pp. 111–116). [19] Koike, R., Kaneko, M., &amp; Okazaki, N. (2023).
Out[5] Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., fox: LLM-generated essay detection through
in&amp; Finn, C. (2023). DetectGPT: Zero-shot machine- context learning with adversarially generated
exgenerated text detection using probability curva- amples. arXiv:2307.11729.</p>
        <p>ture. In ICML 2023. [20] Lu, N., Liu, S., He, R., &amp; Tang, K. (2023). Large
lan[6] Bao, G., Zhao, Y., Teng, Z., Yang, L., &amp; Zhang, Y. guage models can be guided to evade AI-generated
(2023). Fast-DetectGPT: Eficient zero-shot detec- text detection. arXiv:2305.10847.
tion of machine-generated text via conditional prob[-21] Kumarage, T., Sheth, P., Morafah, R., Garland, J., &amp;
ability curvature. arXiv:2310.05130. Liu, H. (2023). How reliable are AI-generated-text
[7] Su, J., Zhuo, T. Y., Wang, D., &amp; Nakov, P. (2023). detectors? An assessment framework using evasive
DetectLLM: Leveraging log rank information for soft prompts. arXiv:2310.05095.
zero-shot detection of machine-generated text[2. 2] OpenAI. (2023). New AI classifier for indicating
AIarXiv:2306.05540. written text, Retrieved November 30, 2023.
[8] Hans, A., Schwarzschild, A., Cherepanova, V., [23] Radford, A., Wu, J., Child, R., Luan, D., Amodei,
Kazemi, H., Saha, A., Goldblum, M., ... &amp; Gold- D., &amp; Sutskever, I. (2019). Language models are
unstein, T. (2024). Spotting LLMs with Binoculars: supervised multitask learners. OpenAI blog, 1(8),
Zero-shot detection of machine-generated text. 9.</p>
        <p>arXiv:2401.12070. [24] Rafel, C., Shazeer, N., Roberts, A., Lee, K., Narang,
[9] Liu, S., Liu, X., Wang, Y., Cheng, Z., Li, C., Zhang, S., Matena, M., ... &amp; Liu, P. J. (2020). Exploring the
Z., ... &amp; Shen, C. (2024). Does DetectGPT fully uti- limits of transfer learning with a unified
text-tolize perturbation? Selective perturbation on model- text transformer. The Journal of Machine Learning
based contrastive learning detector would be better. Research, 21(1), 5485-5551.</p>
        <p>arXiv:2402.00263. [25] Touvron, H., Martin, L., Stone, K., Albert, P.,
Alma[10] Sasse, K., Barham, S., Kayi, E. S., &amp; Staley, E. W. hairi, A., Babaei, Y., ... &amp; Scialom, T. (2023). Llama
(2024). To burst or not to burst: Generating and 2: Open foundation and fine-tuned chat models.
quantifying improbable text. arXiv:2401.15476. arXiv:2307.09288.
[11] Tulchinskii, E., Kuznetsov, K., Kushnareva, L., Cher- [26] Almazrouei, E., Alobeidli, H., Alshamsi, A.,
Capniavskii, D., Barannikov, S., Piontkovskaya, I., ... pelli, A., Cojocaru, R., Debbah, M., ... &amp; Penedo, G.
&amp; Burnaev, E. (2023). Intrinsic dimension estima- (2023). The falcon series of open language models.
tion for robust detection of AI-generated texts. arXiv:2311.16867.</p>
        <p>arXiv:2306.04723. [27] Li, Y., Bubeck, S., Eldan, R., Del Giorno, A.,
Gu[12] Solaiman, I., Brundage, M., Clark, J., Askell, A., nasekar, S., &amp; Lee, Y. T. (2023). Textbooks are all you
Herbert-Voss, A., Wu, J., ... &amp; Wang, J. (2019). Re- need ii: phi-1.5 technical report. arXiv:2309.05463.
lease strategies and the social impacts of language[28] Narayan, S., Cohen, S. B., &amp; Lapata, M. (2018). Don’t
models. arXiv:1908.09203. give me the details, Just the summary! Topic-aware
[13] Hu, X., Chen, P. Y., &amp; Ho, T. Y. (2023). RADAR: convolutional neural networks for extreme
summaRobust AI-text detection via adversarial learning. rization. In Proceedings of the 2018 Conference on
arXiv:2307.03838. Empirical Methods in Natural Language Processing
[14] Dou, Z., Guo, Y., Chang, C. C., Nguyen, H. H., &amp; (pp. 1797–1807).</p>
        <p>Echizen, I. (2024). Enhancing robustness of LLM- [29] Chakraborty, S., Bedi, A. S., Zhu, S., An,
synthetic text detectors for academic writing: A B., Manocha, D., &amp; Huang, F. (2023). On
comprehensive analysis. arXiv:2401.08046. the possibilities of AI-generated text detection.
[15] Liu, Z., Yao, Z., Li, F., &amp; Luo, B. (2023). Check me if arXiv:2304.04736.</p>
        <p>you can: Detecting ChatGPT-generated academic[30] Krishna, K., Song, Y., Karpinska, M., Wieting, J.,
writing using CheckGPT. arXiv:2306.05524. &amp; Iyyer, M. (2023). Paraphrasing evades detectors
[16] Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, of ai-generated text, but retrieval is an efective
I. &amp; Goldstein, T. (2023). A watermark for large defense. arXiv:2303.13408.</p>
        <p>language models. ICML 2023. [31] Mireshghallah, F., Mattern, J., Gao, S., Shokri, R., &amp;
[17] Ren, J., Xu, H., Liu, Y., Cui, Y., Wang, S., Yin, D., Berg-Kirkpatrick, T. (2023). Smaller language
mod&amp; Tang, J. (2023). A robust semantics-based water- els are better black-box machine-generated text
demark for large language model against paraphras- tectors. arXiv:2305.09859.</p>
        <p>ing. arXiv:2311.08721.
[18] Shi, Z., Wang, Y., Yin, F., Chen, X., Chang, K. W.,</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>