<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>F. of Information Project of Privacy In-
ternational, Freedom of information
around the world</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Using ChatGPT for the FOIA Exemption 5 Deliberative Process Privilege</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jason R. Baron</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nathaniel W. Rollings</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Douglas W. Oard</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Maryland</institution>
          ,
          <addr-line>College Park, Maryland</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <volume>2006</volume>
      <abstract>
        <p>Government transparency frameworks such as the Freedom of Information Act (FOIA) in the United States must balance the public's right to know with a number of other considerations. This paper focuses on one such issue, assessment of whether the deliberative process privilege applies in specific cases under FOIA Exemption 5. Providing automated support to the reviewers charged with making such determinations could help to improve responsiveness while controlling review costs. This paper applies ChatGPT-3.5 to explore three ways in which the emerging family of Large Language Models (LLM) might help reviewers with this task: (1) suggesting which passages should and should not be withheld, (2) explaining the basis for those suggestions to the reviewer, and (3) helping the reviewer explain the basis for their decisions to the requestor. The results show that suggestions by ChatGPT-3.5 are not more accurate than previously reported supervised text classification results, that legal analyses in explanations provided by ChatGPT-3.5 are somewhat superficial but generally not unreasonable, that hallucinations are rare, and that explanations provided by ChatGPT may be viewed as useful to a requestor given explanations typically provided with an initial FOIA response.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Freedom of Information Act</kwd>
        <kwd>Deliberative process privilege</kwd>
        <kwd>ChatGPT</kwd>
        <kwd>Sensitivity review</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>was fully annotated in a fashion that establishes “ground
truth” with respect to the factual/deliberative distinction.</p>
      <p>
        Recent research has demonstrated that well-known ma- We can therefore measure the overall accuracy in
Chatchine learning classifiers can achieve at least a modestly GPT’s determinations in labeling passages subject to the
high level of success in being able to discern “deliberative” deliberative process privilege.
material in government documents that fall within the Additionally, given ChatGPT’s ability to provide
exscope of a public access exemption in the U.S. Freedom of planatory narratives in responding to prompts, we have
Information Act [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].1 The research task involved training undertaken as part of this exercise to observe how
difclassifiers to segregate deliberative material (constituting ferences in prompts afect ChatGPT’s explanations and
opinions, recommendations, options, and policy-related “bottom line” conclusions with respect to
deliberativediscussions) from factual material contained in a given ness. We posed variant prompts requesting ChatGPT
document, in line with what FOIA law expects human re- make a determination either on the basis of (i) a simple
viewers to do at federal agencies in response to applicable request for a determination; (ii) a request for a
determiFOIA requests. nation that also asks that specific case law be cited; (iii)
      </p>
      <p>
        To the authors’ knowledge, there has not yet been a a request that a correct determination that is specified as
research efort in applying large language model (LLM) part of the prompt be substantiated; (iv) a request that an
software, for example in the form of ChatGPT, to the task incorrect determination provided as part of the prompt
of segregating factual from deliberative material in the be substantiated; and (v) a request that additional
subcontext of FOIA. This paper represents a preliminary ex- ject matter that is actually irrelevant to making a FOIA
ploration of how ChatGPT performs on selected example determination be considered. In all cases, we wished
passages from documents drawn from the Clinton White to observe how well ChatGPT justified determinations
House document set used in Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. One goal with citations to “real” case law as well as the quality
of doing so is that each paragraph in the document set of these case law citations, as subjectively determined
by the first author of this paper, a legal expert in FOIA
law. The ultimate object of the exercise was to determine
how helpful ChatGPT might be to human reviewers in
cases where a large number of documents determined to
be responsive need to be further reviewed for possible
withholding from public access.
      </p>
      <p>
        A more complete description of applicable FOIA law is
contained in Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. “The fundamental principle
animating FOIA is public access to government
docuIn: Proceedings of the Third International Workshop on Artificial
Intelligence and Intelligent Assistance for Legal Professionals in the Digital
Workplace (LegalAIIA 2023), held in conjunction with ICAIL 2023, June
19, 2023, Braga, Portugal
$ jrbaron@umd.edu (J. R. Baron); nrolling@umd.edu
(N. W. Rollings); oard@umd.edu (D. W. Oard)
      </p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
1Title 5, U.S. Code, Section 552. See
https://www.law.cornell.edu/uscode/text/5/552.
ments.”2 Government records are presumptively open
and available for access, subject to nine general
exemptions.3 Exemption 5 allows for withholding of
“interagency or intra-agency memorandums or letters which
would not be available by law to a party other than an
agency in litigation with the agency.” As a threshold
matter, to be considered “inter- or intra-agency” in nature, a
document4 must not have been sent to or received from
an outside source; only internal communications within
the Executive branch are covered by the exemption.</p>
      <p>In accordance with relevant case law, Exemption 5
allows for (but does not require) agencies to withhold
records in whole or in part that are covered by the
“deliberative process privilege.” To further satisfy the test for
deliberative process privilege, a document must be
“predecisional” in nature, i.e., drafted for internal discussion
prior to a policy decision made by a senior decisionmaker.</p>
      <p>Finally, exempt material must also be “deliberative” in
nature, as opposed to simply a recitation of facts. The
“deliberative process privilege . . . protects ‘documents
relfecting advisory opinions, recommendations and
deliberations comprising part of a process by which government
decisions and policies are formulated.’"5</p>
      <p>Pursuant to the FOIA, “any reasonably segregable
portion of a record” is releasable “after deletion of the
portions which are exempt.”6 Courts have routinely found
factual material in documents to be outside the scope of
Exemption 5. “Purely factual material usually cannot be
withheld under Exemption 5 unless it reflects an exercise
of discretion and judgment calls.”7</p>
      <p>This paper makes the following contributions:
more often than can be achieved using supervised
machine learning techniques.
• We study the legal quality of ChatGPT’s
explanations in justifying its determinations, concluding
that its performance meets minimum standards
roughly equivalent to the quality of legal
analysis and explanations contained in determination
letters at the initial agency stage of responding
to FOIA requests.
• We explore how prompt variations influence</p>
      <p>
        ChatGPT responses, including where erroneous
or irrelevant information is embedded in the
prompt, finding that ChatGPT has significant
limitations in its explanatory powers.
• We suggest future lines of research, including
more directly comparing the eficacy of large
language model systems against an existing set of
classifiers previously used in similar research.
• We release our results as supplemental data to the
fully annotated test collection of Clinton White
House documents previously provided in Baron,
et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Work on automatic detection of sensitive content in texts
has been performed in many contexts, including
privilege review in e-discovery [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ], privacy protection
in search [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ], declassification of materials whose
distribution had been limited for national security
reasons [
        <xref ref-type="bibr" rid="ref8">8, 9</xref>
        ], and redaction of exempt material in response
• We show that ChatGPT is often able to make a to requests under government transparency regimes such
correct “legal” determination as to whether given as FOIA. Among these, we are aware of three research
passages in a document are within the scope of groups who have focused specifically on FOIA or
FOIAthe deliberative process privilege, although no like government transparency applications. Graham
McDonald and others at the University of Glasgow in the
U.K. have published an extensive line of work focused
2Valencia-Lucena v. U.S. Coast Guard, 180 F.3d 321, 325 (D.C. Cir. on review for two exemptions in the U.K. Freedom of
3159U99.S)..C. 552(b)((1)-(9). In addition, there exist three additional Information regime: international relations and personal
narrow exceptions to access involving types of law enforcement material [10, 11, 12, 13, 14]. The first author and others
records. None of the exemptions and exceptions other than one at the University of Maryland in the U.S. have focused
aspect of Exemption 5 are relevant to this research. on the deliberative process privilege under Exemption
4We refer here to “documents” and “records” interchangeably. Of 5 of the U.S. FOIA regime. Finally, a team led by Karl
nmoetnet,se,”-wmiatihl ocormwmithuonuictaatcioconmspaarenycionngsaidttearcehdmsetanntsd.-alone “docu- Branting at The MITRE Corporation in the U.S. has
de5Waterman v.IRS, 2023 WL 2125253 (D.C. Cir. 2023) (quoting NLRB scribed, but not yet published, their work on automatic
v. Sears, Roebuck &amp; Co., 421 U.S. 132, 150 (1975)). detection of content that could be subject to withholding
65 U.S.C. 552(b). under several FOIA exemptions [15].
7A(Dn.Cci.eCnitr.C2o0i1n1C).oHlloewcteovresrv,.coUu.rSt.sDheapv’et aolfsoStsattaet,e6d4t1haFt.3ddis5t0in4c,t5io1n3 All published research of which we are aware in which
as between facts and opinions “must not be applied mechanically.” machine learning techniques have been applied to the
Mapother v. Dep’t of Justice, 3 F.3d 1533, 1537 (D.C. Cir.2011). See, task of segregating sensitive content in documents have
e.g., Reporters Committee for Freedom of the Press v. FBI, 3 F.4th employed either rule-based or supervised
descrimina350, 361, 365-66 (D.C. Cir. 2021) (agency comments “on the accuracy tive classifiers, using statistical techniques such as linear
bofecpauurseelythfaisctfuaaclt-scthateecmkienngtsexiner[cai]sedr.a.ft. rdeipdonrottwcearlel fnoortjduedlgibmeeranttivoer regression or support vector machines (e.g., [11]),
suthe candid exchange of ideas”) (internal quotes omitted). pervised neural classifiers using models such as BERT
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>
        (e.g., [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), or supervised sequence detection models
using, for example, Begin, Inside, Outside (BIO) classifiers
(e.g., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). The general framework for human review processing of
      </p>
      <p>
        Recently, generative Large Language Model (LLM) FOIA requests was set out in Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Here, we
techniques have risen to prominence as an alternative are interested in whether ChatGPT can identify
deliberato supervised descriminative classifiers. Most prominent tive passages, and how well it explains a decision. The
among these at present are a family of LLMs from Ope- Clinton White House collection of materials
comprisnAI known as Generative Pre-Trained Transformer (GPT) ing the test collection consists of documents previously
models. These models adopt a fundamentally diferent selected in Baron, et al. from the files of two high-level
ofapproach to the task, one more focused on explanation ifcials who worked in the Clinton White House. Among
and nuance than is typical of supervised discriminative other positions, Elena Kagan held the title of Deputy
classifier designs. Essentially, these GPT models are ques- Assistant to the President for Domestic Policy and was
tion answering systems that can perform three tasks: deputy director of the Domestic Policy Council.8
Cynthia Rice was a Special Assistant to the President for
• Interpret a question that is posed to the model Domestic Policy.9 The Domestic Policy Council was (and
(either with or without any prior dialog context), still is) responsible for coordinating the policy-making
• Find existing information that can be used to con- process and making recommendations to the President
struct an answer, and with respect to Administration policies.10 The subject
• Generate an answer to the question that is appro- matters covered in the Clinton White House collection
priate to the context in which the question was range across a wide variety of matters of domestic policy
asked. (see Table 1 in Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
      </p>
      <p>
        There is now a burgeoning research community ex- As analyzed here, the set of Clinton White House
ploring how best to use LLMs, including GPT models records analyzed in Baron, et al. were originally divided
in particular, for a wide spectrum of tasks [16]. Since into five “batches” (K1, K2, K3, R4, and K5). (Documents
GPT models are controlled by issuing questions (which, n= 38; Paragraphs n=2213.)11 Four of the five batches in
because they can also be declarative, are generally re- Baron, et al. were previously reviewed by the first author
ferred to as “prompts”), how best to craft those prompts to of this paper, who annotated each paragraph as either
elicit a desired response has received considerable atten- factual or deliberative. In one batch (K2), the author was
tion [17, 18, 19]. This is typically referred to as "prompt joined by a second legal subject matter expert. Where
engineering.” Another active line of work involves as- the two lawyers initially disagreed, they later came to
sessing the degree to which LLMs, and GPT models in a consensus position on all paragraph annotations (see
particular, generate correct and useful responses [20]. Table 2 in Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
      </p>
      <p>Here, one serious concern is that LLMs are prone to For the present exercise, we used the ChatGPT
Vera problem known as “hallucination,” which describes a sion 3.5 API, which for brevity we refer to as ChatGPT.12
situation in which a model that doesn’t know an answer ChatGPT is an autoregressive model, predicting the next
simply makes one up [21]. What makes hallucinations output token (e.g., word) based on (a) the user-provided
particularly problematic is that LLMs are relatively good prompt and (b) all tokens it has already produced for this
at generating believable text, whether that text is correct input in this session. It is also capable of tracking and
conor not. For example, if one asks ChatGPT-3.5 which of sidering (b) the prior prompts and responses provided
the astronauts who walked on the Moon was the tallest, within a session. That conversational tracking ability
the model will ofer an answer, and will helpfully include within a session is not, however, used for the
experithe height of that astronaut. The height it provides would ments in this paper. Like many LLMs, ChatGPT operates
indeed be reasonable for the general population, but it on a stochastic model of language developed during its
is often well over the actual height limit for Apollo-era training, and therefore the same input may produce
difastronauts. Only a domain expert would know that, how- ferent outputs. However, ChatGPT ofers a “temperature”
ever, so the risk is that whoever asked the question would setting, which we set to zero to maximize consistency in
have no way of knowing that in this case ChatGPT was
essentially making things up. In our work, we look at
8https://clinton.presidentiallibraries.us/collections/show/34.
9https://clinton.presidentiallibraries.us/collections/show/60.
both the accuracy of ChatGPT’s recommendations re- 10https://en.wikipedia.org/wiki/United_States_Domestic_Policy_
garding the deliberative process privilege and the quality Council.
of its explanations. While these questions have each been 11For this paper, we excluded a sixth batch in Baron, et al., consisting
studied in other contexts [22], and are being increasingly of a high number of documents not within the scope of Exemption
discussed in various legal contexts, [23, 24, 25], this is 125C.hatGPT has a training cutof date of September
the first application of ChatGPT for review of content for 2021, so it cannot access more recent information.
FOIA exemptions of which we are aware. https://platform.openai.com/docs/models/gpt-3-5
its results for our experiments.</p>
      <p>ChatGPT is not the only LLM capable of making a
decision and producing an explanation when provided
user input. However, its convenient hosting by OpenAI
and its impressive performance on many tasks make it a
compelling starting point for exploring the capabilities
of modern LLMs for determining and explaining FOIA
exemptions.</p>
      <p>The prompt plays a critical role in ChatGPT’s output
because it is considered for every word ChatGPT
produces. As a result, slight changes in the prompt may Figure 1: Relationships among prompts. Each prompt adds
produce substantial changes in the output. Even sim- a new component to consider for its impact on ChatGPT’s
ply adding “Let’s think through this step by step” to the performance at both classification and explanation tasks.
end of a provided logic problem can result in improved
performance [18]. We therefore investigate a variety of
prompts. We conduct five runs across all non-trivial para- Would the following be protected under
graphs in the entire collection of documents (n=1719). FOIA exemption 5? Explain your reasoning
Trivial paragraphs (n=494), marked by the annotators as and cite any case law that supports your
T0, are those which are readily apparent to any reader conclusion.
as being nonexempt. These paragraphs frequently in- SOCIAL SECURITY Office of the
clude signature blocks, headers, etc. Each run across the Commissioner April 25, 1997 MEMORANDUM
documents uses a diferent prompt and considers each TO: Bruce Reed Assistant to the President
paragraph individually: for Domestic Policy SUBJECT: Proposed
Legislation Regarding Nazi War criminals
be The actual paragraph being evaluated would be
ap</p>
      <p>pended to the end of this prompt, separated by two line
following be breaks. Note that while we are adding this additional
exemption 5? information to each paragraph, it does not change the
underlying determination. Nothing in the metadata
provided is exempt, and the entire phrase would still be
marked as exempt if anything in the target paragraph is
covered under FOIA Exemption 5.
• Prompt 0: Would the following</p>
      <p>protected under FOIA exemption 5?
• Prompt 1: Would the
protected under FOIA</p>
      <p>
        Explain your reasoning.
• Prompt 2: Would the following be
protected under FOIA exemption 5?
Explain your reasoning and cite
any case law that supports your
conclusion. 4. Results
• Prompt 3: Would the following be
protected under FOIA exemption 5? Our results are presented in two parts. First, we explore
Explain your reasoning. (with document how often ChatGPT makes the right recommendation
metadata appended) based on a comparison as between ChatGPT’s legal
deter• Prompt 4: Would the following be minations regarding whether a given passage is “factual”
protected under FOIA exemption 5? or “deliberative” in nature, scored using ground truth
Explain your reasoning and cite annotations from Baron, et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Second, we examine
any case law that supports your some selected examples of ChatGPT narrative responses
conclusion. (with document metadata ap- to the prompts in order to illustrate the quality of
narpended) ratives from the standpoint of an expert legal observer.
      </p>
      <p>Although the examples are typical of narratives across the</p>
      <p>Prompt 0 is a generic request for ChatGPT’s determi- entirety of the collection, they are by no means
exhausnation as to whether a given passage is factual or deliber- tive in the types of variations in responses encountered
ative under Exemption 5. Each of the later prompts still throughout the entirety of the test collection.13
seeks this same determination, but extends it in some
way. The relationships between prompts are shown in
Figure 1.</p>
      <p>The metadata included in Prompts 3 and 4 may include
a variety of information, the nature of which depends
on the document in question. Below is an example of a 13The raw output for both sections may be found at https://github.
request using a memorandum’s header as this metadata. com/nater82/ChatGPT_FOIA_Exemption5_Data</p>
      <sec id="sec-3-1">
        <title>Definitely Exempt</title>
        <p>0.186
0.047
0.161
0.135
0.049</p>
      </sec>
      <sec id="sec-3-2">
        <title>Possibly Exempt</title>
        <p>0.010
0.056
0.035
0.053
0.076</p>
      </sec>
      <sec id="sec-3-3">
        <title>Unsure</title>
        <p>0.657
0.550
0.564
0.448
0.461</p>
      </sec>
      <sec id="sec-3-4">
        <title>Possibly Nonexempt</title>
        <p>0.115
0.325
0.152
0.250
0.389</p>
      </sec>
      <sec id="sec-3-5">
        <title>Definitely Nonexempt</title>
        <p>
          0.032
0.023
0.088
0.115
0.025
4.1. Measures of Accuracy The category distribution for each prompt type is
shown in Table 1 of the 1,719 D0 (“decided as
nonWe provided the five prompts shown in Figure 1 to Chat- exempt”) or D1 (”decided as exempt”) passages, and in
GPT for every paragraph in the dataset. While manually Table 2 for the 494 T0 (“trivial to classify”) passages.14
examining the explanations of all 1719 paragraphs for As can be seen, ChatGPT is often unsure of what to do
each of these prompts would be infeasible, we were able with a short T0 passage, whereas it is more confident
to conduct an overall evaluation of the classification por- on the (typically much longer) D0 and D1 paragraphs
tion of this task. Although each response can be unique, as can be seen in Tables 1 and 2. In the remainder of
the initial sentences are often similar, especially in the this section, we report results only for D0 and D1
paraifrst dozen or so words. ChatGPT frequently provided graphs. We note here that although T0 paragraphs might
its overall determination near the beginning of this first be “trivial to classify” for a human reviewer, our design
sentence, so we were able to use exact string match to in which ChatGPT sees those passages without any of
cluster its commonly used initial expressions. However, their surrounding context does seem to make them far
it did not always provide a definitive answer. It would from trivial for ChatGPT.
sometimes make statements such as “The text provided With all of ChatGPT’s responses classified, we then
calwould likely be exempt under FOIA exemption 5.” While culated the measures shown in Table 3 using the ground
this claim clearly leans towards the document being ex- truth determinations provided by Baron, et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. We
empt, it is not as strong a statement as a result such as used two diferent scoring approaches to analyze the
re“Yes, the information would be protected under FOIA ex- sults of these runs. First, we only considered the cases
emption 5.” To account for these diferences, the second where ChatGPT’s responses were definitely exempt to
author of this paper manually classified responses as ei- be “exempt,” and the cases in which its response was
ther definitely exempt, possibly exempt, unsure, possibly definitely not exempt to be “nonexempt,” with the
possinonexempt, or definitely nonexempt. The unsure cate- bly exempt, unsure, and possibly nonexempt categories
gory included all statements where ChatGPT refused to always marked as wrong. We call this the “Hard” scoring
make any commitment either direction, instead making condition. Our second scoring approach treated both
claims along the lines of “It is unclear whether the above definitely exempt and possibly exempt as exempt, unsure
record would be protected under FOIA exemption 5.” While as always wrong, and possibly nonexempt and definitely
the majority of these determinations could be made on nonexempt as nonexempt. We call this the “Soft” scoring
large clusters of identical verbiage as a result of Chat- condition.
        </p>
        <p>
          GPT’s usually consistent phrasing of its determinations
in the initial sentence, slightly over 200 examples had to
be individually classified. These cases typically involved
the inclusion of unusual terms or determinations buried
in later sentences of the response.
14Baron, et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] refer to all of D0, D1, and T0 as ”paragraphs,”
but the T0 cases are more commonly individual lines or short
passages containing only what we might think of as metadata (e.g.,
date, sender, ...). For clarity, when discussing T0 in particular, we
therefore refer to those items as passages rather than as paragraphs.
        </p>
        <p>
          The most important observation we can draw from cision for every type of prompt. This suggests that
igTable 3 is that ChatGPT-3.5 is not particularly impres- noring “possibly” recommendations may not be the best
sive as a classifier for this task. The best 1 achieved approach. Rather, some value from the positive
recomby any discriminative classifier reported by Baron, et mendations might be obtained if they could be assigned
al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] (in Table 10, column 4), when evaluated on the a lower confidence value when using ChatGPT as one
same collection using cross-validation, was 0.704 (with system among many in a classifier ensemble.
Precision 0.720 and Recall 0.689). Our best 1 in Table 3, While we cannot determine the inner workings of
by contrast, is just 0.584 (with Precision 0.528 and Recall ChatGPT, both due to the size and general architecture
0.654). As Table 4 shows, We could increase this to an 1 of LLMs and because its code and parameters are only
of 0.610 (Precision 0.572, Recall 0.654) by the simple ex- visible to OpenAI, these results provide us some insights
pedient of treating the unsure cases as the majority class into potential reasons for varying performance between
(which in terms of ground truth is nonexempt) rather prompts. The additional document metadata in Prompts
than the conservative approach we have taken in Table 3 3 and 4 can indicate whether a paragraph is a
commuof always marking unsure as wrong. But even that result nication within or between government agencies, a fact
is still just 87% of the 1 that a discriminative classifier not always evident within a given paragraph. As a result,
achieved.15 it may help encourage ChatGPT to take more definite
        </p>
        <p>Indeed, ChatGPT’s best 1 on this collection could positions. As we can see in Table 3, this improved both
be matched by just guessing that everything is exempt, Precision and Recall.
since that’s the ground truth answer for 45.0% of the cases The case law requests in Prompts 2 and 4 are also
inter(1=0.620, Precision=0.450, Recall=1.000), while discrimi- esting, because ChatGPT often presents its determination
native classifiers easily beat that simple baseline. So if all before citing a court case. This is potentially
consequenthat is needed is a classifier, it is not clear that ChatGPT tial, because the GPT family of models only considers the
alone would be the best choice. Of course, discriminative prompt and the words it has previously predicted when
classifiers lack the nuance of ChatGPT recommendations, determining the next word to output. Comparing results
and they ofer no explanations for their decisions. So the with Soft scoring for Prompts 1 and 2 in Table 3, we see
best we can say from these 1 statistics is that it is per- that requesting a reference to case law helps on average
haps ChatGPT’s talents for nuance and explanation that if no metadata is provided. However, comparing Prompts
might be its most interesting capabilities. 3 and 4, we see that additionally requesting a reference</p>
        <p>We can also see from Table 3 that, as expected, Soft to case law has a much smaller impact once metadata
scoring has better Recall for every type of prompt. Per- has been provided.
haps more surprisingly, Soft scoring also has better
Pre15We note, however, that Baron et al also includes results showing
that a domain shift in the training data can reduce the 1 on this
collection to 0.525 (Precision 0.410 Recall 0.730), although we note
that the domain shift experiments in that paper also made use of
less training data.</p>
        <sec id="sec-3-5-1">
          <title>4.2. Quality of Explanations</title>
          <p>Beyond the aggregate measures of accuracy reported
above, we also sought to characterize the overall
quality of ChatGPT’s responses. What follows is necessarily
impressionistic, based on a much smaller set of
examples pulled from the overall collection. In this section,
the main features of ChatGPT responses we analyze are
ChatGPT’s determination of whether specific content is
deliberative or factual in nature, whether there are
misstatements or inconsistencies in how Exemption 5 case
law is set out, and how useful we expect its explanations
would likely be. Unlike the analysis in the prior section,
in this section we sometimes include more than one
paragraph in a prompt, and we also look at some cases in
which the prompt contains a T0 passage rather than D0
or D1 paragraph(s).
4.2.1. ChatGPT Correct in Finding Content Exempt
Our first example consists of an excerpt from an e-mail
sent to Elena Kagan from a White House colleague
concerning a meeting on a controversial proposal by the
Social Security Administration to terminate benefits to
Nazi suspects when they were “denaturalized” (i.e., had
their citizenship removed), prior to when they might be
deported.16 Opponents of the proposal (including Neil
Sher, former head of the Ofice of Special Investigations
at the Department of Justice, and Doug Bloomfield,
representing the World Jewish Congress), argued that
judges would be less likely to order deportation if the
government was not only “trying to strip them of their
citizenship,” but “also impoverishing them while they
are trying to defend their rights during the judicial
process.”17 The proposal was not enacted by Congress at
the time, but a “No Social Security for Nazis Act” was
enacted many years later, in 2014.18 The email contains
the following paragraph:
16https://buchanan.house.gov/sites/buchanan.house.gov/files/AP%</p>
          <p>20Report%20Nazis.pdf
17Id.
18Pub.L. 113-270.</p>
          <p>And make no mistake, Messrs. Sher and
Bloomfield were absolutely clear that they
and their organizations would protest
vociforously [sic]. They would protest in the
media, and they would protest in Congress. They
would make sure that it is known that the
President has introduced legislation that would
compromise the prosecutions of Nazis.</p>
          <p>With Prompt 2, which included a request for legal
citations, ChatGPT responded as follows:</p>
          <p>The passage would likely be protected under FOIA exemption
5, which covers privileged inter-agency or intra-agency
communications. The passage appears to be a communication between
individuals discussing potential protests and political action
related to a government decision or action. This could be considered
part of the internal deliberative process of an agency, which is
protected under exemption 5.</p>
          <p>One relevant case is NLRB v. Sears, Roebuck &amp; Co. (421 U.S.
132 (1975)), in which the Supreme Court held that exemption 5
applies to documents that are “predecisional” and “deliberative”
in nature. The Court explained that this exemption is intended
to protect the decision-making process of government agencies
by ensuring that oficials can freely exchange ideas without fear
of public scrutiny.</p>
          <p>In the case of the passage given, the communication is
predecisional in that it pertains to potential actions related to a decision
that has not yet been made. Additionally, the communication
appears to be deliberative in that it discusses potential responses
and strategies related to that decision. Therefore, it is likely that
this passage would be protected under exemption 5.</p>
          <p>ChatGPT’s response is both accurate and coherent. It is
correct in guessing that the document was as a threshold
matter an internal meeting in the Executive branch. The
meeting is clearly pre-decisional and contained
deliberations of senior oficials on a matter of policy. Specifically
with respect to legal analysis, ChatGPT chose to cite to an
early, leading opinion by the Supreme Court discussing
the Exemption 5 test. NLRB v. Sears Roebuck is the first
cited Supreme Court case in the definitive guide to the
FOIA issued by the Department of Justice, in its section
on the deliberative process privilege.19</p>
          <p>Interestingly, rerunning Prompt 2, using identical
language in the prompt, resulted in ChatGPT spontaneously
added a paragraph discussing Exemption 7(E),
authorizing withholding documents pertaining to certain types
of law enforcement information. ChatGPT said:</p>
          <p>Finally, it is worth noting that even if the statement were not
protected under exemption 5, it could potentially be protected
under exemption 7(E), which protects law enforcement
information that would disclose investigative techniques and procedures.</p>
          <p>The statement suggests that Messrs. Sher and Bloomfield are
involved in some sort of investigation or prosecution of Nazis,
and therefore could potentially be considered law enforcement
information. However, without more context it is dificult to
determine whether exemption 7(E) would apply in this case.</p>
          <p>This supplemental narrative provides an erroneous
interpretation of FOIA law. The document is not part of
a law enforcement investigatory file; nor does it contain
any discussion of the details of any ongoing investigation
of Nazi War Criminals under threat of deportation.
ChatGPT here may have been led astray by certain keywords
appearing in the passage, most prominently the word
“prosecution.”</p>
          <p>Finally, changing the prompt to ask ChatGPT to
“cite any DoD documents that support your
conclusion” resulted in ChatGPT citing to
“Department of Defense (DoD) Directive 5400.7-R, which
provides guidance on the implementation of FOIA.”20 While
this Directive exists, it is a short document that does
not contain any discussion of Exemption 5, and does
not contain the language that ChatGPT quotes from in
the document. This is an instance of a hallucination on
ChatGPT’s part.
other academic areas taught in the school.</p>
          <p>Springbrook has a very diverse student body
(35% African American, 20% Asian, 16%
Hispanic, 28% white), and it enrolls a total of 2,200
students.</p>
          <p>ChatGPT’s response to Prompt 2:</p>
          <p>This information does not appear to be protected under FOIA
exemption 5, as it does not relate to any agency decision-making
process, inter- or intra-agency communication, or deliberative
process. Rather, it is a general description of a high school’s
demographics and academic programs.</p>
          <p>Exemption 5 only applies to materials that are “inter-agency
or intra-agency memorandums or letters which would not be
available by law to a party other than an agency in litigation with
the agency.” 5 U.S.C. §552(b)(5). The purpose of this exemption
is to protect the decision-making process of the government by
allowing oficials to freely exchange ideas and opinions without
fear of premature disclosure.</p>
          <p>In the absence of any connection to an agency decision-making
process or communication, exemption 5 is not applicable. There
is no relevant case law on this specific issue as it falls outside the
scope of exemption 5.</p>
          <p>ChatGPT’s response is correct that the paragraph is
simply “a general description of a high school’s
demographics and academic programs,” and as such, is factual
and therefore non-exempt. However, the last sentence
stating that there is no relevant case law is erroneous,
and inconsistent with the majority of responses in which
ChatGPT has cited case law stating that factual material
is outside the scope of Exemption 5. ChatGPT’s failure
to do so here is an unexplained anomaly.
4.2.3. ChatGPT Incorrectly Determines Content To</p>
          <p>Be Non-Exempt
This short paragraph constitutes one of three “strategy
4.2.2. ChatGPT Correct in Finding Content options” discussed in an email from 1997 circulating</p>
          <p>Non-Exempt among members of the Domestic Policy Council,
regardThe following paragraph is from the “Background” sec- ing whether the Administration should draft a bill on
tion of a briefing document dated March 14, 1998, de- an environmental policy known as “Superfund.” The
scribing a meeting at Springbrook High School in Illinois paragraph below posits the option of the Clinton
Adminthat President Clinton will have the next day with leaders istration merely assisting legislative stafs in coming up
from the government, business, education, and scientific with bill language; two other options involve the
Admincommunities. He will discuss recent findings from an istration taking the lead in drafting either a narrow or
international math and science study that showed U.S. broad bill.
12th graders lagging below the international average in
these subjects. The President will announce at that
meeting new on-line math and science assistance to parents,
teachers, and students.</p>
          <p>Springbrook High School is a good school
with above-average test scores. Although the
school has a solid math and science program,
it is not particularly stronger than any of the
19https://www.justice.gov/oip/page/file/1573681/download#page=15
20https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/
dodd/540007p.pdf</p>
          <p>A) Prepare Administration principles, but
allow agency assistance to both majority and
minority staf to improve/prepare their bills.
(There is a republican bill, there is as yet no
dem alternative.) It would need to be
absolutely clear that assistance does not guarantee
Administration endorsement. EPA is likely to
support this view.</p>
          <p>ChatGPT’s response to Prompt 0 was:</p>
          <p>The content of the message itself does not appear to contain
information that would fall under FOIA exemption 5. It does
not appear to involve discussions or recommendations made The proposed legislation in Option A would not be protected
in the process of formulating or developing policies, nor does under FOIA Exemption 5 because it involves the formulation of
it involve pre-decisional or deliberative discussions. Rather, it policy rather than the deliberative process. The Supreme Court
discusses potential strategy options and considerations related has held that Exemption 5 protects only documents that are
to a legislative issue. However, as always, the determination both predecisional and deliberative, meaning they are part of the
of whether a specific document is subject to FOIA exemption 5 decision-making process and contain opinions, recommendations,
ultimately depends on the specific context and content of the or advice. However, the proposed legislation in Option A is a final
document, and it is possible that additional information could decision rather than a recommendation or advice. Therefore, it
change the analysis. would not be protected under Exemption 5. On the other hand, the</p>
          <p>ChatGPT has made two errors in its narrative response. proposed legislation in Option B would likely be protected under
First, the passage is discussing taking a policy position FOIA Exemption 5 because it involves the deliberative process.
that the Environmental Protection Agency is “likely to The proposed legislation is a recommendation or advice to limit
support.” While the passage does not expressly make a the Administration’s proposed legislation to compensation of
recommendation, it nevertheless easily fits within the miners with lung cancer for whom it is more probable than not
that the cancer was caused by exposure to radiation in uranium
case law supporting withholding under the deliberative mines. This recommendation is part of the decision-making
process privilege. Second, ChatGPT makes an erroneous process and contains opinions and advice. Therefore, it would
statement of law where it goes on to say: “Rather, [the be protected under Exemption 5. Case law that supports this
passage] discusses potential strategy options and con- conclusion includes NLRB v. Sears, Roebuck &amp; Co. (1975), which
siderations related to a legislative issue.” Discussions of established that Exemption 5 protects documents that are both
“potential strategy options and considerations” are at the predecisional and deliberative. Additionally, in Coastal States
core of the deliberative process. It is dificult to recon- Gas Corp. v. Department of Energy (1981), the Supreme Court
cile ChatGPT’s justification here with other instances of held that Exemption 5 protects documents that are part of the
making correct determinations. decision-making process and contain opinions, recommendations,
or advice.</p>
          <p>Here, ChatGPT has only gets the correct answer for
4.2.4. ChatGPT Response After Being Provided an</p>
          <p>Option B, and not A. For reasons not entirely clear, it
conExternally Supplied Determination
cludes that the text in Option A would be nonexempt
“beThis example consists of two paragraphs taken from a cause it involves the formulation of policy rather than the
memorandum shared between two members of the Do- deliberative process.” This evidences a misunderstanding
mestic Policy Council titled “Proposed Changes to RECA - of the scope of the deliberative process privilege, which
Decision Needed by January 7,” pertaining to the Radioac- presumptively involves “formulations” of policy as
equivtive Exposure and Compensation Act of 1990. As enacted, alent to deliberations. ChatGPT also states that “Option
RECA provided for uranium miners working on nuclear A is a final decision rather than a recommendation or
weapons programs who had been exposed to specified advice.” This is factually incorrect.
high dosages of radiation and who contracted lung can- However, when Prompt 2 was instead modified to
incer to be compensated up to $100,000. The memorandum form ChatGPT that the two paragraphs were deliberative,
discussed policy options for expanding compensation ChatGPT responded:
to miners where it cannot be determined how high the This text is protected under FOIA Exemption 5, which exempts
original dosage of radiation would have been. from disclosure “inter-agency or intra-agency memorandums or
The paragraphs given to ChatGPT were: letters that would not be available by law to a party other than an
agency in litigation with the agency.” The text contains
delibera</p>
          <p>Option A: Propose legislation to provide tive material, which is defined as “predecisional and deliberative
at least partial compensation ($50,000) to all material, which is both predecisional and deliberative in nature,
miners for whom the probability is between and which reflects the give-and-take of the consultative process.”
20% and 50% that their lung cancer was caused The text presents two options for legislation regarding
compensaby exposure to radiation in uranium mines. tion for miners with lung cancer caused by exposure to radiation
This approach would also include proposing in uranium mines. The options are presented as proposals and
the measures outlined in Section 1, above, for are part of a consultative process to determine the best course of
miners with greater than 50% probability. action. The text reflects the internal deliberations of the agency</p>
          <p>Option B: Limit the Administration’s pro- and is therefore protected under Exemption 5. Case law that
posed legislation to compensation of miners supports this conclusion includes NLRB v. Sears, Roebuck &amp; Co.
with lung cancer for whom it is more prob- (1975), which established that Exemption 5 protects “documents
able than not that the cancer was caused by reflecting advisory opinions, recommendations and deliberations
exposure to radiation in uranium mines (i.e. comprising part of a process by which governmental decisions
only propose measures outlined in Section 1 and policies are formulated.” Additionally, in Coastal States
above). Gas Corp. v. Department of Energy (1981), the Supreme Court
First, in response to Prompt 2 (not supplying an an- held that Exemption 5 protects “documents that would reveal
swer), ChatGPT said the mental processes of decisionmakers.” The text in question
contains advisory opinions, recommendations, and deliberations, The statement would likely be protected under FOIA exemption
and would reveal the mental processes of decisionmakers, making 5, which pertains to privileged inter- or intra-agency
communiit eligible for protection under Exemption 5. cations. The statement appears to be a communication between</p>
          <p>Asking ChatGPT in the prompt to justify that the para- individuals in the government discussing potential reactions to a
graphs are deliberative results in a correct, coherent, ex- proposed policy, and therefore could be considered privileged.
planation as to why both Options A and B reflect internal There is no direct Ulysses S. Grant quote that applies to this
deliberations. In this example, however, ChatGPT’s more specific situation, as the Freedom of Information Act was not
“directed” response doesn’t amount to a noticeably sig- enacted until 1966, long after Grant’s presidency. However, Grant
nificant enhancement in the quality of the explanation. cdoidunetmabpihliatsyizien tghoeviemrnpmoretnatn.
cIenohfitsrfirasntsipnaaruegnucyraalnadddpruebssli,chaecstated, “The country having just emerged from a great rebellion,
4.2.5. ChatGPT Response After Being Provided many questions will come before it for settlement in the next
Erroneous Determination four years which preceding Administrations have never had to
deal with. In meeting these it is desirable that they should be
Continuing with the prior example, for these paragraphs approached calmly, without prejudice, hate, or sectional pride,
we instead provided to ChatGPT a determination that was remembering that the greatest good to the greatest number is the
erroneous, i.e., that Options A and B were factual, asking object to be attained.” This suggests a commitment to openness
it to justify that determination with legal citations. In and fairness in government decision-making, which could inform
response to this modified version of Prompt 2, ChatGPT interpretations of FOIA exemptions.
stated: ChatGPT’s second paragraph contains an interesting</p>
          <p>The text is not protected under FOIA Exemption 5 because it discussion. The first sentence of the response is correct,
consists of factual material related to proposed legislation and as insertion of the name of 19th century U.S. President
does not involve any deliberative process or policy-making dis- was intended to be absurd on its face with no applicability
cussions. As such, it would not be subject to exemption under to FOIA. Or was it? The remaining discussion contains
the deliberative process privilege. Case law supporting this con- an ironic determination regarding Grant “emphasiz[ing]
(ca2lpu0p1sli1ioe)n,sitinnocwmluhaditceehsriMtahlisel ntSheuraptvr.aemrDeeebpCoatorhtumprtreenhdteelodcfisttihhoaentaNElaxavenymd, p5dt6ei2loinbUe.5rSa.ot5ni6vly2e itnhegiomvepronrmtaenncte,”ofgitvraennstphaerehnicsytoarnicdaplucbolnicseancscuousnthtaabtiltihtye
in nature. Additionally, in Judicial Watch, Inc. v. U.S. Depart- Grant Administration was one of the most corrupt in
ment of Justice, 365 F.3d 1108 (D.C. Cir. 2004), the D.C. Circuit American history. [26] Despite the erroneous premise,
held that factual material is not protected under Exemption 5, the last sentence in the paragraph provides an unexpected
even if it is contained within a document that also includes de- connection, in the essentially correct statement that "a
liberative material. commitment to openness and fairness in government</p>
          <p>Here, ChatGPT appears to be simply parroting back decision-making" does "inform interpretations of FOIA
the erroneous prompt, without making any kind of in- exemptions." Indeed, recent examples of this include
guiddependent decision as to whether the prompt accurately ance by Attorneys General in the Obama and Biden
Addescribes the passages. The justification and case law ministrations urging agencies to err on the side of
disclocited simply support its conclusory assumption. sure where FOIA exemptions (including pertaining to the</p>
          <p>We did, however, see a few examples where ChatGPT deliberative process privilege) are discretionary [27] [28].
made an independent determination. In one instance, In making the connection between Grant’s inaugural
adwhen incorrectly told that a passage was deliberative, dress and FOIA we don’t mean to claim that ChatGPT
ChatGPT simply stated that the passage was factual, go- actually “thought” this way, but rather that what
Chating on to cite to authorities excluding factual material GPT wrote causes us to think this way. ChatGPT has
from the scope of the privilege without supplying any thus shown its ability to serve as what in other contexts
further analysis. has been referred to as a “provocation” [29, 30].
4.2.6. ChatGPT Response When Prompt Contains</p>
          <p>Irrelevant Language
For our example in section 4.2.1 regarding social security
payments to former Nazis undergoing deportation
proceedings, we added irrelevant material in the prompt.</p>
          <p>Here is the prompt variation ChatGPT evaluated for
the passage in 4.2.1, supra: Would the following be
protected under FOIA exemption 5? Explain your
reasoning and cite any Ulysses S. Grant to
support your conclusion.</p>
          <p>ChatGPT responded stating the following:
4.2.7. Additional Examples
Here we comment on three additional aspects of
ChatGPT’s responses.</p>
          <p>1. Responding to just an isolated signature line with
name and title given, ChatGPT responded with:
“In this case, the statement is simply a signature
block and closing of a letter, which does not
contain any deliberative process information.
Therefore, it would not be protected under Exemption
5." Here ChatGPT was correct.
2. Responding to a single sentence passage stating:
“Let me know what you think. Thanks Elena.”
With Prompt 2, ChatGPT’s response was: As an
AI language model, I do not have personal
opinions or beliefs. However, I can provide an answer
to your question based on my programmed
knowledge." ChatGPT went on to say it was impossible
to determine whether the passage would be
protected under Exemption 5. This is in error. Taken
in isolation, most FOIA reviewers would find that
a request for an opinion would be covered by</p>
          <p>Exemption 5.
3. In many of its responses ChatGPT stated the
important statutory requirement previously noted
in section 1, namely, that segregability is an
important criterion for evaluating particular pas- Example B.
sages in documents with respect to whether they
may or may not be exempt. For example, in one
case ChatGPT added to its response:
It is worth noting that the fact that a document
contains information that is exempt from disclosure under
FOIA exemption 5 does not necessarily mean that the
entire document is exempt. Rather, the exemption
applies only to the specific information that is covered by
the privilege.</p>
          <p>In other cases, ChatGPT went so far as to focus on
individual sentences, or even clauses within
sentences, when remarking on segregability.
However, in many cases ChatGPT made no explicit
references to segregation. So this is interesting
more for what it indicates about what ChatGPT
could do, rather than what it currently typically
does do.</p>
          <p>[W]hile a policy and schedule have been
agreed upon by the Commission, our
Capstone program is not yet fully implemented.</p>
          <p>Documents which are pre-decisional in
nature and related to the eventual
implementation of Capstone have been withheld under
B(5). All documents and discussions related to
the leadup to Capstone implementation have
been withheld under B(5). . . Exemption 5
protects from disclosure inter- or intra-agency
memoranda or letters that would not be
available by law to a party other than an agency
in litigation with the agency, including
documents covered by the attorney work-product,
deliberative process, and attorney-client
privileges. See 5 U.S.C. § 552(b)(5).22</p>
          <p>FOIA Exemption 5 protects from disclosure
those inter- or intra-agency documents that
are normally privileged in the civil discovery
context. The three most frequently invoked
privileges are the deliberative process
privilege, the attorney work-product privilege, and
the attorney-client privilege. After carefully
reviewing the responsive documents, I
determined that portions of the responsive
documents qualify for protection under [the]
Deliberative Process Privilege.</p>
          <p>The deliberative process privilege protects
the integrity of the deliberative or
decisionmaking processes within the agency by
exempting from mandatory disclosure opinions,
conclusions, and recommendations included
within inter-agency or intra-agency
memoranda or letters. The release of this internal
information would discourage the expression
of candid opinions and inhibit the free and
frank exchange of information among agency
personnel.23</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>4.3. Comparison To Real FOIA Responses</title>
          <p>The quality of the explanations contained in agency
responses to actual FOIA requests varies greatly. In
connection with other research, Baron filed FOIA requests
to numerous agencies asking for documents on the
“Capstone” approach to managing and preserving federal
email records [31].21 Responses to these actual FOIA
requests, contained in what are known as “determination
letters,” vary greatly in the quality of their narrative
justifications with respect to full or partial withholdings
of materials subject to the deliberative process privilege.
Relevant passages (each containing the entirety of the
explanation) are set out in the following examples.
Example A.
21An agency choosing to adopt the Capstone approach commits
to preserving all email records of designated senior oficials as
permanent records to be transferred to the National Archives
and Records Administration [31]. Some 250 components
of government have put into place Capstone policies (see
https://www.archives.gov/records-mgmt/rcs/schedules/capstoneforms).</p>
          <p>Example C.</p>
          <p>Regarding FOIA Exemption 5, draft
documents and internal memoranda are being
withheld pursuant to Exemption 5, 5 U.S.C.
§ 552(b)(5). Exemption 5 allows agencies to
withhold “inter-agency or intra- agency
memorandums or letters which would not be
available by law to a party other than an agency in
litigation with the agency,” and covers records
that would “normally be privileged in the civil
discovery context.” NLRB v. Sears, Roebuck
&amp; Co., 421 U.S. 132, 149 (1975); Tax Analysts
22Excerpt from Determination Letter of the Federal Election
Commission, dated July 27, 2021, responding to FOIA Request 2021-078
(dated July 5, 2021) (on file with authors).
23Excerpt from Determination Letter of the US Department of
Homeland Security, dated April 17, 2023, responding to FOIA Request
2021-HQFO-01478 (dated August 23, 2021) (on file with authors).
v. IRS, 117 F.3d 607, 616 (D.C. Cir. 1997). The
deliberative process and the attorney
workproduct privileges are two of the primary
privileges incorporated into Exemption 5. The
deliberative process privilege protects the
internal decision-making processes of government
agencies to safeguard the quality of agency
decisions. Competitive Enter. Inst. v. OSTP, 161
F. Supp.3d 120, 128 (D.D.C. 2016). The basis
for this privilege is to protect and encourage
the creative debate and candid discussion of
alternatives. Jordan v. U.S. Dep’t. of Justice,
591 F.2d 753, 772 (D.C. Cir.1978). Two
fundamental requirements must be satisfied before
an agency may properly withhold a record
pursuant to the deliberative process privilege.</p>
          <p>First, the record must be predecisional, i.e.,
prepared in order to assist an agency
decisionmaker in arriving at the decision.
Renegotiation Bd. v. Grumman Aircraft Eng’g Corp.,
421 U.S. 168, 184 (1975); Judicial Watch, Inc.
v. FDA, 449 F.3d 141, 151 (D.C. Cir. 2006).</p>
          <p>Second, the record must be deliberative, i.e.,
“it must form a part of the agency’s
deliberative process in that it makes recommendations
or expresses opinions on legal or policy
matters.” Judicial Watch, Inc. v. FDA, 449 F.3d
at 151 (quoting Coastal States Gas Corp. v.</p>
          <p>U.S. Dep’t of Energy, 617 F.2d 854, 866 (D.C.</p>
          <p>Cir. 1980)). To satisfy these requirements, the
agency need not “identify a specific decision
in connection with which a memorandum is
prepared. Agencies are . . . engaged in a
continuing process of examining their policies;
this process will generate memoranda
containing recommendations which do not ripen
into agency decisions; and the lower courts
should be wary of interfering with this
process.” Sears, Roebuck &amp; Co., 421 U.S. at 151
n.18 (1975). Moreover, the protected status of
a predecisional record is not altered by the
subsequent issuance of a decision, see, e.g.,
Fed. Open Mkt. Comm. v. Merrill, 443 U.S.
340, 360 (1979); Elec. Privacy Info. Ctr. v. DHS,
384 F. Supp. 2d 100, 112-13 (D.D.C. 2005) or by
the agency opting not to make a decision. See
Judicial Watch, Inc. v. Clinton, 880 F. Supp. 1,
13 (D.D.C. 1995), af’d, 76 F.3d 1232 (D.C. Cir.
1996) (citing Russell v. U.S. Dep’t of the Air
Force, 682 F.2d 1045 (D.C. Cir. 1982)). Here,
the responsive records being withheld meet
the requirements for Exemption 5 protection
under the deliberative process privilege. They
are internal and predecisional. They reflect
the views of Agency employees concerning
the implementation of Capstone. Since they
contain internal discussions, these case
handling records clearly reflect the deliberative
and consultative process of the Agency that
Exemption 5 protects from disclosure. Sears,</p>
          <p>Roebuck and Co., 421 U.S. at 150-52.24</p>
          <p>Example A constitutes a response that does provide a
rationale for withholding tailored to the request, but fails
to provide an adequate justification as a matter of law
where it merely cites to the language of the statute,
without further explanation. The response makes a reference
to the “pre-decisional” nature of documents withheld,
but not as to whether they are also “deliberative,” and
nowhere discusses case law opining on the scope of the
deliberative process privilege. In contrast, the
explanation in Example C provides an extended discussion of the
rationale behind the deliberative process privilege, with
numerous citations from FOIA case law, including
recent cases. Example C also contains a second paragraph
tying the prior discussion of case law to the specifics
of the FOIA request. Example B falls somewhere in the
middle: the agency provides a generic justification for
withholding documents under the deliberative process
privilege, but fails to cite case law or tie its discussion to
the specifics of the FOIA request. Measured subjectively,
the ChatGPT narratives generated as part of the research
(and without the benefit of “knowing” what the incoming
FOIA request was about), most closely approximate the
language contained in real-world Example B; except that
Prompt 2 asking for legal citations resulted in a greater
degree of legal justification in the ChatGPT response
than the agency chose to provide to the requestor here.</p>
          <p>Legal citations, if tailored to the specifics of a request,
can be important in that they provide a requestor with
precedent to be taken into account in making her
assessment as to whether resources should be expended
in filing an administrative appeal from this
determination, with the option of seeking judicial review after the
agency has made its final decision. However, even in the
case of Example C, it is likely at the initial determination
stage that the more nuanced language consists merely of
“boilerplate” in response to all requests involving a
deliberative process privilege determination. If this is in fact
the case, the explanatory value of even this example is
diminished, which in turn narrows the perceived gap as
between Example C and ChatGPT’s equally “boilerplate”
responses.</p>
        </sec>
        <sec id="sec-3-5-3">
          <title>4.4. General Observations</title>
          <p>Without a comparable baseline in terms of studies of the
accuracy of human review of FOIA Exemption 5
determinations, it is dificult to make an objective determination
of the overall accuracy rate of ChatGPT responses. Past
assumptions regarding the accuracy of human review
on matters of legal determinations have been shown to
24Excerpt from Determination Letter of the National Labor Relations</p>
          <p>Board, dated May 17, 2022, responding to FOIA Request
NLRB2021-001052 (dated June 24, 2021) (on file with authors).</p>
          <p>Case
National Labor Relations Board v. Sears, Roebuck &amp; Co.</p>
          <p>Coastal States Gas Corp. v. U.S. Department of Energy
United States v. Weber Aircraft
EPA v. Mink
Tax Analysts v. IRS
Public Citizen Health Research Group v. FDA
Judicial Watch, Inc. v. U.S. Department of Justice
Public Citizen, Inc. v. Ofice of Management and Budget
Milner v. Department of the Navy
National Wildlife Federation v. United States Forest Service
Schiller v. NLRB
United States v. Nixon
Citizens for Responsibility and Ethics in Washington v. U.S. Department of Justice
Mead Data Central, Inc. v. U.S. Department of the Air Force
Public Citizen v. U.S. Department of Justice
Department of the Interior v. Klamath Water Users Protective Association
Jordan v. U.S. Department of Justice
Crooker v. Bureau of Alcohol, Tobacco and Firearms
Armstrong v. Executive Ofice of the President
In re Sealed Case
Upjohn Co. v. United States
Federal Court
Supreme Court</p>
          <p>DC Circuit
Supreme Court
Supreme Court</p>
          <p>DC Circuit
DC Circuit
DC Circuit
DC Circuit
Supreme Court
Ninth Circuit</p>
          <p>DC Circuit
Supreme Court</p>
          <p>DC Circuit</p>
          <p>DC Circuit
Supreme Court</p>
          <p>DC Circuit
DC Circuit
DC Circuit
DC Circuit</p>
          <p>DC Circuit
Supreme Court
be measurably in error [32]. The authors had no pre- enhanced over the Prompt 2 condition, where we asked
conceived views as to the level of accuracy ChatGPT for a justification with citations but without providing
would achieve in this initial research exercise. As mea- the true determination. Further investigation across a
sured across all batches, ChatGPT’s overall ability to broader sample of documents would be needed to
estabdetermine whether paragraphs contain withholdable ma- lish whether ChatGPT has the capability of enriching its
terial, on the order of 60% as measured by accuracy or by explanations when informed of the correct answer in the
1, would need to be improved before actual deployment prompt.
by an agency is realistically contemplated. In similar fashion, in most – but not all – instances</p>
          <p>As noted in connection with the comparisons as be- where we asserted an erroneous determination to be
tween ChatGPT narratives and real-world examples, correct (e.g., a statement that a passage is deliberative
ChatGPT’s narratives on the whole consist of what when it is factual in nature), ChatGPT also took the
errolawyers would consider “boilerplate” responses, which neous statement as true. In such cases, it typically tried
while adequately setting out the most fundamental facets to substantiate the incorrect determination using
boilof how courts and commentators characterize the deliber- erplate language, without providing additional analysis
ative process privilege, do not provide additional context that would help a reviewer to make the correct decision.
or insight into agency deliberations with respect to why This lends a cautionary note in terms of how we should
particular material was found exempt. This characteriza- approach relying on ChatGPT’s narratives.
tion could, however, easily be said to also apply in some In 5%-10% of D0 and D1 cases, ChatGPT declined to
measure to actual agency responses to FOIA requests. opine on whether the paragraph was or was not
covViewed in this light, it is dificult to find particular fault ered by the deliberative process privilege, stating that it
with the quality of ChatGPT’s narratives for this purpose, would need additional information (i.e., context) to reach
especially where ChatGPT was not made aware of the a determination. In general, the shorter the passage, the
language of the incoming FOIA requests that generated more likely that ChatGPT stated a need for additional
the documents that ChatGPT was asked to review. information. This accords with the human review
experi</p>
          <p>For the small sample of paragraphs (n=34) in which ence, where shorter documents may provide less context
we included a correct legal determination as part of the in which to make a decision on withholding.
prompt, ChatGPT perhaps unsurprisingly generally re- ChatGPT’s narratives focus on citations to early
promistated the given legal conclusion as a working assumption nent legal cases that are frequently cited in scholarly
when finding relevant case law. In doing so, the quality of publications on FOIA and in FOIA litigation. As Table 5
the narrative explanation appeared to be only marginally shows, cited decisions consist primarily of FOIA opinions</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion</title>
      <p>from the 1970’s through the early 2000’s, with the
overwhelming majority consisting of Supreme Court and D.C.</p>
      <p>Circuit opinions. Moreover, the Supreme Court’s deci- We clearly are closer to the beginning than to the end of
sion in National Labor Relations Board v. Sears, Roebuck our investigation of the use of ChatGPT to protect
sensi&amp; Co. accounts for over half (57%) of the total citations tive content, even in the narrow context of one part of
in the corpus. Only one cited case with five or more men- one exemption of one government transparency regime.
tions was decided in the last dozen years. While there are Thus, we see this not as time for making definitive
statetwelve U.S. circuit courts of appeal, in only one instance ments, but rather initial conclusions that will need to be
did ChatGPT cite to a case in other than the D.C. Circuit revisited as we learn more. We have learned two
impor(citing to a Ninth Circuit case). tant things. One is that although not yet at the level of</p>
      <p>ChatGPT’s “choices” in citing to case law are very a sophisticated FOIA reviewer, ChatGPT-3.5 is already
much in line with the real world: the overwhelming at a point where it can bring useful recommendations to
number of briefs filed in FOIA cases cite to the Supreme the table. The second is that ChatGPT-3.5 is already at
Court, and since the majority of FOIA cases are filed in least as adept at explaining its findings as at least some
the US District Court for the District of Columbia, cita- agencies choose to be in issuing their responses to FOIA
tions to FOIA cases decided by the D.C. Circuit are both requestors.
appropriate and commonplace. ChatGPT understandably From a legal standpoint, this preliminary exploration
missed citation to the most recent Supreme Court FOIA into the accuracy and lucidity of ChatGPT responses
5 Exemption case25 decided only months prior to the in making FOIA legal determinations illustrates
ChatSeptember 2021 cutof date for ChatGPT’s training. It is GPT’s promise for perhaps ultimately evolving as an aid
less clear, however, how to explain the general absence of to human review. Conceding that ChatGPT responses
citations to recent case law, given the hundreds of FOIA are far from perfect should not discount their ability to
decisions in the Supreme Court, federal appellate and assist in large volume productions where it may be
imporfederal district courts that have been handed down since tant to automate the “flagging” of deliberative material
the turn of the century. One plausible theory is that the throughout a given universe of otherwise responsive
docearly Supreme Court cases and D.C. Circuit cases act as uments to a given FOIA request. In this respect, ChatGPT
a “sink,” drawing citations from subsequent FOIA case could be viewed as functioning as the equivalent of a
“julaw, and that ChatGPT has simply learned from that. nior colleague” in any overall agency review process.26</p>
      <p>It should also be noted that there were no instances Any human-machine collaboration that enhances FOIA
of “hallucinations” in ChatGPT’s citing to legal case au- productions could help meet widespread criticism as to
thorities, where ChatGPT either provided a completely the quality of the FOIA review process as practiced at
erroneous holding associated with a particular case, or the federal level in the U.S., including but not limited to
“made up” imaginary case citations. On the other hand, inordinate delays experienced by requestors, often
couas a general rule the addition of case law citations did pled with inadequate explanations of why agencies have
not much enhance the quality of the supplied narrative, withheld documents under FOIA exemptions, including
in terms of ChatGPT expressly applying case holdings to Exemption 5 [33, 34].
the specifics of what documents purported to be about. Given the rapid progress in Large Language Model</p>
      <p>The ChatGPT narratives contained a fair number of (LLM) development, with ChatGPT-4 already available
inaccurate choices with respect to whether documents to some users, we see this as a good start. Moreover, we
as a threshold matter are covered under the inter- or note that we have obtained the aggregate results that
intra-agency test in Exemption 5. Notably, the inclusion we report in this paper using only the first five prompt
of document metadata in Prompts 3 and 4 appears to types that we tried. We might further improve ChatGPT’s
improve overall 1, suggesting that additional contex- accuracy by structuring the prompts in a way that more
tual information might provide additional gains. Among explicitly brings the system through the points that we
other things, the metadata may reveal names and email think should be considered, in the order we think they
addresses of individuals involved with the document in would best be considered. Moreover, we might consider
question that could act as flags that a communication has leveraging the ability to fine-tune a model that OpenAI
been sent or received by individuals outside the Executive ofers (for a price) to help ChatGPT to better learn which
branch. Further investigation of the impact of diferent case law is available to be cited, and which citations
forms of metadata may provide additional insights into would be most useful.
how ChatGPT is using this information. To date we have studied ChatGPT in isolation, but
applying a wider lens, its possible uses are certainly
extendible in a variety of ways. One obvious idea is to use
25U.S. Fish and Wildlife Service v. Sierra Club, 141 S. Ct. 777 (2021).
26Alexandra Samuel, “A Guide to Collaborating With ChatGPT for</p>
      <p>Work,” Wall Street Journal (April 11, 2023).</p>
      <p>ChatGPT as one classifier among an ensemble of
classiifers that each have diferent strengths and weaknesses.</p>
      <p>A larger leap would be to add prior decisions by other
FOIA reviewers into the workflow. So far in our work,
we have always given the same prompt (or, in the case of
Prompts 3 and 5, at least a prompt with the same
structure) to ChatGPT for every request. But as reviewers
see documents and make decisions, a real reviewer will
learn. For our ChatGPT experiments, however, every
request has been decided de novo, with no reference to
prior decisions. This is known as a “zero-shot” model.</p>
      <p>An alternative would be to craft a few-shot model, in
which we show ChatGPT what some good and bad
answers look like. That, however, then raises the question
of how best to select those examples. If that can be done
well, few-shot learning might help without the greater
expense of fine-tuning the model for this specific task.</p>
      <p>We might also investigate the use of follow-on prompts
to further expand upon ChatGPT’s initial response. This
exploration could be particularly useful when it cites to a
case, either asking for the relevant facts of the cited case,
or requesting it to find more recent cases citing back to
the case, all in an efort to determine if ChatGPT can
provide additional insights into the relevance of the cases it
cites. Posing an additional prompt when ChatGPT
determines that the passage would be exempt which requests
the specific language that led to this decision may also
improve the ability for a human reviewer to determine
the accuracy of its prediction. As we have seen notable
impact from even relatively small modifications to the
prompts in this paper, additional prompt engineering
may provide improvements in the quality of decisions
and explanations or at least give some extra insights into
ChatGPT’s behavior.</p>
      <p>While we have ofered our own opinions on the degree
of salience, cogency, and correctness of what ChatGPT
has to say, we (one legal subject matter expert in FOIA
litigation with extensive experience adjudicating appeals of
FOIA exemption decisions, and two computer scientists)
are certainly not representative of the target audience
for such an automated assistant. Some user studies,
perhaps with FOIA review staf at agencies, would clearly
be useful.</p>
      <p>We are closer to the beginning than to the end not
just because there is more to be done, but also because
the tools we are using to do this are themselves
evolving rapidly. Thus, by the time we have answers to our
questions, it seems reasonable to expect that there will
be many more new questions to be answered.</p>
      <p>What will remain constant is the importance of
making correct decisions as to which documents or portions
thereof should be accessible to ordinary citizens, a legal
concept that is not limited to the FOIA experience in the
U.S. A substantial number of international FOIA statutes
contain exclusions from public access for documentary
materials pertaining to the internal deliberations of
government oficials [ 35, 36].27 How much in the way of
deliberative material a given jurisdiction chooses to
release about the workings of government in response to
citizen requests is an important measure of the health
of its democracy in its commitment to transparency and
openness [37]. Automated processes that include
machine learning, now including what many call generative
AI, may yet make useful contributions towards that
important goal.
History Reveals About America’s Top Secrets., Pen- Proceedings of the 2023 CHI Conference on Human
guin Random House, 2023. Factors in Computing Systems, 2023, pp. 1–21.
[9] C. Martin, D. Lam, A. Liu, M. Tech, Classification As- [20] N. F. Liu, T. Zhang, P. Liang, Evaluating
verifiasistance Prototype System: Final Report, Technical bility in generative search engines, arXiv preprint
Report LR-SISL-07-17, Applied Research Laborato- arXiv:2304.09848 (2023).
ries: The University of Texas at Austin, Austin, TX, [21] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii,
2007. Y. J. Bang, A. Madotto, P. Fung, Survey of
hal[10] E. Frayling, C. Macdonald, G. McDonald, I. Ounis, lucination in natural language generation, ACM
Using entities in knowledge graph hierarchies to Computing Surveys 55 (2023) 1–38.
classify sensitive information, in: Experimental IR [22] Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su,
Meets Multilinguality, Multimodality, and Interac- B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung,
tion: 13th International Conference of the CLEF Q. V. Do, Y. Xu, P. Fung, A multitask,
multiAssociation, CLEF 2022, Bologna, Italy, September lingual, multimodal evaluation of ChatGPT on
5–8, 2022, Proceedings, Springer, 2022, pp. 125–132. reasoning, hallucination, and interactivity, 2023.
[11] G. McDonald, A framework for technology-assisted arXiv:2302.04023.</p>
      <p>sensitivity review: using sensitivity classification to [23] D. M. Katz, M. Bommarito, S. Gao, P. Arredondo,
prioritise documents for review, Ph.D. thesis, Uni- GPT-4 passes the bar exam, 2023. URL: http://dx.
versity of Glasgow, 2019. doi.org/10.2139/ssrn.4389233.
[12] G. Mcdonald, C. Macdonald, I. Ounis, How the [24] B. Ambrigo, New GPT-based chat app from
accuracy and confidence of sensitivity classification LawDroid is a lawyer’s ‘copilot’ for research,
afects digital sensitivity review, ACM Transactions drafting, brainstorming and more, 2023.
on Information Systems (TOIS) 39 (2020) 1–34.
https://www.lawnext.com/2023/01/new-gpt[13] H. Narvala, G. Mcdonald, I. Ounis, The role of latent
based-chat-app-from-lawdroid-is-a-lawyerssemantic categories and clustering in enhancing the
copilot-for-research-drafting-brainstorming-andeficiency of human sensitivity review, in: ACM SI- more.html.</p>
      <p>GIR Conference on Human Information Interaction [25] A. Lamparello, ChatGPT and legal writing, 2023.
and Retrieval, 2022, pp. 56–66. URL: https://lawprofessors.typepad.com/appellate_
[14] H. Narvala, G. McDonald, I. Ounis, Sensitivity re- advocacy/2023/03/chatgpt-and-legal-writing.
view of large collections by identifying and priori- html.
tising coherent documents groups, in: Proceed- [26] R. Chernow, Grant, Penguin Press, New York, 2017.
ings of the 31st ACM International Conference on [27] B. Obama, Memorandum on freedom
Information &amp; Knowledge Management, 2022, pp. of information act, 2006. URL: https:
4931–4935. //www.justice.gov/sites/default/files/oip/legacy/
[15] N. Romps, Searching for Solutions: MITRE tool 2014/07/23/presidential-foia.pdf .
simplifies freedom of information act requests, [28] D. of Justice, Freedom of information act
guidePress Release, 2023. https://www.mitre.org/news- lines, 2022. URL: https://www.justice.gov/ag/page/
insights/impact-story/mitre-tool-simplifies- ifle/1483516/download.</p>
      <p>freedom-information-act-requests. [29] J. Drucker, B. Nowviskie, Speculative computing:
[16] Y. Wang, S. Mishra, P. Alipoormolabashi, Y. Ko- Aesthetic provocations in humanities computing,
rdi, A. Mirzaei, A. Naik, A. Ashok, A. S. in: A Companion to the Digital Humanities,
BlackDhanasekaran, A. Arunkumar, D. Stap, et al., Super- well, 2004. https://companions.digitalhumanities.
naturalinstructions: Generalization via declarative org/DH/.
instructions on 1600+ NLP tasks, in: Proceedings [30] P. J. Denning, The profession of IT: Can generative
of the 2022 Conference on Empirical Methods in AI bots be trusted, Communications of the ACM
Natural Language Processing, 2022, pp. 5085–5109. 66 (2023).
[17] S. Wang, A brief summary of prompting in us- [31] N. Archives, R. Administration, White paper on the
ing GPT models (2023). https://doi.org/10.32388/ capstone approach and capstone GRS, 2015. URL:
IMZI2Q. https://www.archives.gov/files/records-mgmt/
[18] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, Y. Iwasawa, email-management/final-capstone-white-paper.</p>
      <p>Large language models are zero-shot reasoners, in: pdf .</p>
      <p>Advances in Neural Information Processing Sys- [32] B. Hedin, S. Tomlinson, J. R. Baron, D. W. Oard,
tems, 2022. Overview of the TREC 2009 Legal Track, in: TREC,
[19] J. Zamfirescu-Pereira, R. Y. Wong, B. Hartmann, 2009. URL: https://trec.nist.gov/pubs/trec18/papers/
Q. Yang, Why johnny can’t prompt: how non-AI LEGAL09.OVERVIEW.pdf .
experts try (and fail) to design LLM prompts, in: [33] A. Lamparello, ’there is a big problem’:
Sen</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Baron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Sayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <article-title>Providing more eficient access to government records: a use case involving application of machine learning to improve FOIA review for the deliberative process privilege</article-title>
          ,
          <source>ACM Journal on Computing and Cultural Heritage (JOCCH) 15</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Grossman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hedin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <article-title>Overview of the TREC 2010 legal track</article-title>
          .,
          <source>in: TREC</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Vinjumur</surname>
          </string-name>
          ,
          <article-title>Jointly minimizing the expected costs of review for responsiveness and privilege in e-discovery</article-title>
          ,
          <source>ACM Transactions on Information Systems (TOIS) 37</source>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Vinjumur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Paik</surname>
          </string-name>
          ,
          <article-title>Assessing the reliability and reusability of an e-discovery privilege test collection</article-title>
          ,
          <source>in: Proceedings of the 37th international ACM SIGIR conference on research &amp; development in information retrieval</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1047</fpage>
          -
          <lpage>1050</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Sayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <article-title>Jointly modeling relevance and sensitivity for search among sensitive content</article-title>
          ,
          <source>in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>615</fpage>
          -
          <lpage>624</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Sayed</surname>
          </string-name>
          , Search Among Sensitive Content,
          <source>Ph.D. thesis</source>
          , University of Maryland, College Park,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Sayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Mallekav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          ,
          <article-title>Comparing intrinsic and extrinsic evaluation of sensitivity classification</article-title>
          ,
          <source>in: Advances in Information Retrieval: 44th European Conference on IR Research</source>
          , ECIR
          <year>2022</year>
          , Stavanger, Norway,
          <source>April 10-14</source>
          ,
          <year>2022</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>215</fpage>
          -
          <lpage>222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Connelly</surname>
          </string-name>
          , The Declassification Engine:
          <article-title>What 27One example is especially pertinent to the 3rd</article-title>
          <source>Legal AIIA Workshop at ICAIL</source>
          <year>2023</year>
          <article-title>: in Portugal, the 1993 Law of Access to Administrative Documents (LADA) “allows any person to demand access to administrative documents held by state authorities, public institutions, and local authorities in any form</article-title>
          .” However, “
          <article-title>[a]ccess to documents in proceedings that are not decided or in the preparation of a decision can be delayed until the proceedings are complete or up to one year after they were prepared</article-title>
          .” [36]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>