<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>William Fleshman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benjamin Van Durme</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Johns Hopkins University</institution>
          ,
          <addr-line>3101 Wyman Park Dr, Baltimore, MD 21218</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Large language models (LLMs) are increasingly capable of completing knowledge intensive tasks by recalling information from a static pretraining corpus. Here we are concerned with LLMs in the context of evolving data requirements. For instance: batches of new data that are introduced periodically; subsets of data with user-based access controls; or requirements on dynamic removal of documents with guarantees that associated knowledge cannot be recalled. We wish to satisfy these requirements while at the same time ensuring a model does not forget old information when new data becomes available. To address these issues, we introduce AdapterSwap, a training and inference scheme that organizes knowledge from a data collection into a set of dynamically composed low-rank adapters. Our experiments demonstrate AdapterSwap's ability to support eficient continual learning, while also enabling organizations to have fine-grained control over data access and deletion.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;LLM</kwd>
        <kwd>adapter</kwd>
        <kwd>access-control</kwd>
        <kwd>data removal</kwd>
        <kwd>forgetting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>A motivating application of AdapterSwap is displayed in Figure 1. In this fictitious example,
AdapterSwap has been fit to hospital data with subsets of the data subject to various access-controls. Both a
cardiologist and member of the finance ofice submit the query ‘ How much does the average cardiac
visit cost?’ The retriever model selects the most relevant adapters to the query for which each user has
access. The cardiologist’s unique access to appointment notes and patient records enables the model to
access specific payments from cardiac patients and respond accordingly. In contrast, the finance ofice’s
access to payroll and supply expenditures results in a response from the hospital’s perspective without
leaking the private patient information to the unauthorized employee.</p>
      <p>The rest of this paper is structured as follows. In Section 2 we discuss challenges which motivate
our research. In Section 3 we provide context with prior works from which we build upon. Section 4
details AdapterSwap, our proposed approach. We demonstrate and quantify benefits of AdapterSwap
through careful experimentation in Section 5. We discuss additional related works in Section 6. Finally,
we conclude and suggest future research in Section 7.</p>
      <p>Specifically, in this work we:
• Develop an eficient approach for continuous knowledge acquisition through the training and
dynamic composition of multiple LoRA adapters;
• Demonstrate our method’s ability to guarantee data access-control and handle data removal in
an eficient manner;
• Quantify our performance via a document completion task across a diverse set of LLMs: Falcon-7B,</p>
      <p>Gemma-7B, Llama-2-7B, and Mistral-7B; and
• Show that our approach mitigates forgetting better than iterative fine-tuning and retraining.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Motivation</title>
      <p>Our work is primarily motivated by three issues faced when fine-tuning and deploying LLMs in
realworld organizations. Namely, how to utilize data with access-controls, how to remove knowledge from
a model retroactively, and how to update knowledge over time as new data becomes available.</p>
      <sec id="sec-2-1">
        <title>2.1. Data Access-control</title>
        <p>
          It is common for organizations to apply access-control to their sensitive data [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For example, only
certain employees at a hospital have access to patient records to protect privacy. Similarly, employees at
a law firm are concerned with attorney-client privilege. Access-control can also be an important aspect
of business models such as a news aggregator giving access to users and their personalized chatbots
based on their paid subscriptions [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We would like a language model trained on these data to inherit
and guarantee these access restrictions.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data Protection and Removal</title>
        <p>
          Organizations can unexpectedly lose the rights to maintain certain data. For example, data protection
policies such as the General Data Protection Regulation (GDPR) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] allow users to recall their data
from corporations. The Stack, a popular dataset comprised of source code repositories, allows users to
opt-out of having their code in future versions of the dataset but ofers no solution for models trained
on previous versions [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Similarly, the removal of training data later found to be copyright protected
or unlicensed might be mandated through legal action, a growing concern for LM producers [
          <xref ref-type="bibr" rid="ref5 ref9">9, 5</xref>
          ].
Existing models provide no mechanism to remove all knowledge from individual training examples, so
to comply with these mandates would require the entire model to be retrained. Therefore, we would
like a more eficient approach to guarantee the removal of data from models.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Catastrophic Forgetting</title>
        <p>
          Organizations often have access to continuous or evolving streams of data. Catastrophic forgetting is an
issue that arises when a machine learning model forgets previously seen information as it learns from
new data [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. LLMs have been shown to sufer from forgetting during the fine-tuning process [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
We would like a method that addresses this issue and guarantees the ability to recall old information as
new knowledge is continuously acquired.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Background</title>
      <sec id="sec-3-1">
        <title>3.1. Parameter Eficient Fine-Tuning</title>
        <p>As language models have become more specialized, growing in size and capabilities, fine-tuning an
entire model has become unreasonable on commodity hardware. To address these challenges, several
methods have been developed for performing parameter eficient fine-tuning .</p>
        <p>
          References [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] present techniques for training a sparse subset of parameters in a multi-task
setting. Alternatively, prompt tuning and prefix tuning concatenate learned task-specific embeddings to
the sequence of inputs or activations being processed by a model [
          <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
          ].
        </p>
        <p>
          In the wider context of transfer learning, adapter layers provide a straight forward mechanism to
eficiently and efectively generalize a base model to a target task or domain by fine-tuning a new set
of parameters (an adapter) on the target data [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. Low-rank adapters (LoRA) have emerged as a
parameter eficient approach to fine-tuning large language models with reasonable amounts of compute
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. In this work, we leverage LoRAs to continuously update and control a language model’s knowledge.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model Averaging and Segmentation</title>
        <p>
          Several approaches have been suggested for combining model weights or outputs with demonstrated
increases in performance or eficiency in certain scenarios. Reference [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] proposed Model Soups which
average the weights of multiple models trained on the same data with diferent hyper-parameters.
While they show that the soups increase performance and robustness of language models, they do not
address combinations of models trained on separate data.
        </p>
        <p>
          AdapterFusion was introduced by [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] as an approach to multi-task learning that segments
taskspecific knowledge into separate adapters that are then combined via an attention mechanism. While
segmenting tasks is similar to segmenting data based on access controls, the attention mechanism adds
additional complexity and the model dependency prevents the eficient removal of data.
        </p>
        <p>
          Reference [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] extended ideas from AdapterFusion with their approach AdapterDrop. AdapterDrop
prunes adapters for increased eficiency but still lacks the ability to address eficient data removal or
continual training.
        </p>
        <p>
          AdaMix was proposed by [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and uses a mixture-of-experts approach to combining adapters at each
layer for the purpose of parameter sharing but not for specific knowledge segmentation.
        </p>
        <p>
          Hierarchical Adapters by [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and AdapterSoup by [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] are the most similar to our approach as they
train individual adapters on segmented domains in the training data, but their focus is on combining the
adapters to perform well with out-of-domain queries, while our focus is on in-domain access-control
and eficient knowledge deletion.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. AdapterSwap</title>
      <sec id="sec-4-1">
        <title>4.1. Data Segmentation and Adapter Training</title>
        <p>The first stage of our approach is choosing a data partitioning scheme. Data is segmented into separate
access-control categories and further sharded based on desired per-shard compute requirements2. A
pretrained LLM is then used as a base model for fine-tuning a separate LoRA adapter per data partition.
As the information from each partition is isolated to a single adapter, the adapter inherits the
accesscontrol categories of the data. This partitioning and fine-tuning stage can be done once for a static
dataset, or on a continual basis as data arrives.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Retrieval Model</title>
        <p>
          Similar to AdapterSoup, we use a Gaussian Mixture Model (GMM) to retrieve the subset of adapters
relevant to a given query during inference. To fit the GMM we use a pretrained SBERT [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] model
to embed randomly held-out samples from each partition into vectors of dimension 7683. We further
reduce the dimension of these vectors by applying linear discriminant analysis (LDA). While previous
approaches such as [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] and [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] use principal components analysis (PCA), in this work we compare
PCA to LDA. LDA has the theoretical benefit of maximizing the linear separability of the clusters in
the lower dimensional space by using their labels in a supervised manner. We found that using LDA
over PCA significantly improved our downstream retrieval accuracy which we discuss in Section 5.3.
Finally, the GMM is fit on the lower dimensional vectors with the number of components equal to the
number of adapters. The eficiency of LDA and GMMs allows for cheap retraining of the retriever if
new adapters are added over time.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Inference</title>
        <p>
          During inference, queries are embedded using the same approach, and the GMM is used to rank potential
adapters. The user’s access control categories are applied so that restricted adapters are prevented from
being selected. We explore several retrieval modes where either the top-1, top-2, or top-3 adapters with
the highest GMM density are selected and combined with the base model. We also attempted averaging
all non-restricted adapters with equal weight or by weighting according to GMM density, but the results
for those scenarios were poor and are therefore omitted. In practice, the choice of retrieval mode could
vary by context. For example, our experiments show that top-1 results are likely best conditioned on
knowledge that the information being retrieved was isolated to a single adapter. However, combinations
of multiple adapters have been shown to perform better for out of domain queries [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Data Removal</title>
        <p>
          Finally, if circumstances arise that require permanently removing data from our training set, only the
weights associated with the LoRA adapter trained on the removed data need to be retrained; less than
0.1% of the base model’s parameters in our experiments. This contrasts with emerging approaches such
as [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] for eficiently fine-tuning over the entire base model; which ofers no savings as all fine-tuned
parameters would require discarding if data is removed.
2See Section 5.2 for a discussion on this topic.
3Specifically, we use the all-mpnet-base-v2 model from https://huggingface.co/sentence-transformers/all-mpnet-base-v2.
        </p>
        <p>In Section 5 we demonstrate AdapterSwap using several models under both access-control and
data removal scenarios. We also compare AdapterSwap’s ability to prevent forgetting with alternative
methods for continuous learning. An overview of AdapterSwap training is illustrated in Figure 2.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <p>We demonstrate the efectiveness by AdapterSwap through multiple experiments. First, we describe the
datasets and models used for our experiments in Section 5.1. In Section 5.2 we establish the training
times and baseline performance when fitting multiple adapters over a sharded dataset.</p>
      <p>Next, in Section 5.3 we quantify our ability to retrieve and compose mixtures of adapters during
inference. We demonstrate AdapterSwap’s ability to conform to data access-controls in Section 5.4
and efectiveness when removing data in Section 5.5. Finally, we compare AdapterSwap’s resilience to
forgetting against two alternative approaches for continuous learning in Section 5.6</p>
      <p>For all of our experiments, we evaluate AdapterSwap by segmenting documents into equal halves
and measuring the perplexity of the model while force decoding the second half given the first. This
document completion task is suited for determining if a model has trained on and remembered particular
samples from the datasets. We also report training times in GPU Hours using a single 80GB A100 GPU.</p>
      <sec id="sec-5-1">
        <title>5.1. Data and Models</title>
        <p>
          We use two datasets for our experimentation. First, we use the subset of C4 [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] utilized by [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] which
contains 21 website domains where unique pages from the domain represent separate documents. We
treat each domain as a separate LoRA training group for our experiments involving access-control and
data purging. The list of training domains and their corresponding document counts are shown in
Table 1.
        </p>
        <p>
          We also use an English subset of the WMT News Crawl Dataset [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. Passages from articles published
in the year 2020 were extracted and deduplicated following [27]. These passages were then segmented
into LoRA training groups based on their month of publication. The chronological nature of this dataset
makes it suitable for measuring a model’s ability to recall previous training data as subsequent months
are trained.
        </p>
        <p>We leverage a diverse set of LMs as base models to ensure our approach generalizes. We replicate
experiments across Falcon-7B [28], Gemma-7B [29], Llama-2-7B [30] and Mistral-7B-v0.1 [31].</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Shard Size, Time, and Performance Trade-ofs</title>
        <p>For each model, we trained a separate adapter on all training groups. Because our groups naturally
difer in size we are able to measure average training time and performance for diferent sharding
strategies. Figure 3 displays diferences in training time and document completion performance as a
function of shard size. Partitioning the dataset into smaller groups results in the need for more adapters
but enables faster training of each. The individual training time is an important characteristic when
rs60
u
oH40
PU20
G
s 60
r
tep 40
ad 20
A
# 0</p>
        <p>
          We utilize the HuggingFace [32] library to access the pretrained models and fit LoRA adapters
using Parameter Eficient Fine-Tuning [ 33]. All adapters were trained on a single 80GB A100 GPU to
standardize timing comparisons. In practice, AdapterSwap can be trained in parallel across several
devices. Each adapter was trained for 10 epochs with a batch size of 20. Rank 32 LoRA adapters were
applied to all attention layers for the News Crawl data and rank 64 adapters on all linear layers for
C4. All adapters were initialized with the same random seed, which was identified by [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] as being
necessary for adapter mixing. A detailed list of hyperparameters are included in appendix.
y
itx 4
e
l
p
re 2
P
rs900
u
o
lH800
a
t
o
T700
faced with the need to retrain adapters if data is later purged. We observe that smaller partition sizes
tend to result in better perplexities, likely due to adapters having a higher ratio of parameters to training
tokens. Overall, the total GPU hours is roughly equivalent across partitioning schemes, with a slight
overhead resulting from adding each additional adapter.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Retrieval</title>
        <p>
          An optimal retriever should return the adapters which have knowledge related to a given query. In our
case we know each sample was seen by only a single adapter, which we refer to as the oracle adapter,
but multiple adapters might be necessary in general information retrieval scenarios. Table 2 displays
the document completion performance across all models when using the oracle adapter, as well as an
average of the top-1, top-2, and top-3 adapters as ranked by the retriever model. We also show the
top-1 adapter with PCA as the projection method to compare with previous works [
          <xref ref-type="bibr" rid="ref21 ref23">21, 23</xref>
          ]. For all
models, using the top-1 adapter with LDA resulted in the best completions and retrieved the oracle
adapter with accuracy varying from 69% to 81%. With the exception of the Gemma-7B model, results
for the top-2 and top-3 schemes are reasonable and [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] show that higher mixtures work well when the
query is out of domain. The accuracy for retrieving the oracle adapter in the top-3 ranged from 93% to
95%, suggesting a potential inference scheme like [34]’s, where output is selected from the mixture and
individual adapters via a separate content-selector.
AdapterSwap can be directly applied to scenarios where data is organized into access-control categories.
We simulate this with the C4 dataset by assigning each domain to a diferent category. For each
document completion, we use our retrieval model to return the top-1 adapter under two scenarios: with
access to all adapters and with access to all but the adapter trained on the restricted domain.
        </p>
        <p>The results are summarized in Table 3. As expected, the best completion was achieved using the
adapter with access to the relevant data, and performance dropped significantly when the retriever did
not have access to the domain. This demonstrates AdapterSwap’s ability to enforce access-control at
the adapter level, providing the best results to users while simultaneously preventing unauthorized
access to restricted training data.
↓No Access</p>
        <p>↓With Access</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.5. Data Removal</title>
        <p>We also measure our ability to eficiently purge documents from our dataset after training. When a
piece of data is removed from the corpus, only the adapter fine-tuned on that data requires retraining.
We show this by attempting to complete documents for each adapter before and after removing them
and retraining the adapter.</p>
        <p>Table 4 summarizes the results for this experiment. We see a significant performance drop when
trying to complete documents that have been purged from their adapter as the model is now guaranteed
to have lost access to that data. Referring back to Figure 3, retraining a single adapter is up to 80x more
eficient than if you had to fine-tune over the entire dataset. Purging with AdapterSwap provides the
guarantee that removed documents will not contribute to later inferences without the huge cost of full
retraining.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.6. Catastrophic Forgetting</title>
        <p>Finally, we compare AdapterSwap’s ability to recall past information with two alternative strategies. A
naive approach to handling new streams of data is to iteratively fine-tune a LM as new data becomes
available. This workflow can cause older information to be ‘overwritten’ within the architecture’s
parameters. Alternatively, as new data is added, the model can be fine-tuned from scratch on all available
data at once. While this mitigates some ‘forgetting’ it requires longer training times as models are
discarded and retrained. We use News Crawl to evaluate both methods and compare to AdapterSwap.
We iteratively measure performance on the first month of our dataset as we add subsequent months of
data to our models. For this experiment we only use Falcon-7B as the baseline due to the increased
computational demand required by the alternative approaches.</p>
        <p>Figure 4 displays the performance of AdapterSwap compared to chronological fine-tuning and
full retraining. Chronological fine-tuning quickly degrades as more data is presented to the model.
Retraining performs better, but sufers from the fixed capacity of a single adapter. AdapterSwap
maintains a static performance when recalling data from the first month as that adapter remains
unchanged over time.</p>
        <p>y 14
t
i
x
lrep 12
e
P
10</p>
        <p>FT
RT</p>
        <p>AS
0</p>
        <p>5 10</p>
        <p>Months of Training</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Additional Related Work</title>
      <sec id="sec-6-1">
        <title>6.1. Knowledge Editing</title>
        <p>Recent work in knowledge editing of LLMs has considered approaches which either add additional
parameters to a model, or directly edit existing parameters to update information [35]. Directly updating
the existing parameters is attractive as it does not require any additional parameters, and updates can be
applied whenever new knowledge is available [36, 37]. Knowledge vectors, in combination with hidden
representations of specific entities, have also been proposed as a tool to update or remove knowledge
[38]. In all cases, these approaches lack the ability to guarantee that any specific training example can
be removed entirely from the model.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Retrieval Augmented Generation</title>
        <p>
          Retrieval-Augmented Generation (RAG) has become a popular approach to incorporating new data
into a pre-trained LLM without having to rely on retraining or fine-tuning [
          <xref ref-type="bibr" rid="ref27">39, 40</xref>
          ]. While it ofers
some advantages, there remain challenges that can make deploying an efective RAG-based solution
dificult. For example, RAG is limited by how much retrieved context can be employed based on the
underlying LLM’s context window size. This contrasts with AdapterSwap where each adapter in our
experiments represented more than 50 million tokens on average. Recent work has shown that all
evidence in the context window is not treated equally, with models favoring evidence at the start and
end of each window [
          <xref ref-type="bibr" rid="ref28">41</xref>
          ]. Further, RAG is fully dependent on the ability of a retriever model to locate
all relevant documents without introducing too much noise into the context [
          <xref ref-type="bibr" rid="ref27 ref29 ref30">42, 40, 43</xref>
          ]. In addition, as
transformers are quadratic in the number of tokens considered, then requiring forced decoding over
retrieved content can add significant latency at inference time. While our solution avoids these issues,
we note that AdapterSwap does not preclude the use of RAG, and a hybrid approach could be useful in
some circumstances.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Federated Learning</title>
        <p>
          Federated learning [
          <xref ref-type="bibr" rid="ref31">44</xref>
          ] also deals with learning from siloed data, typically aggregating gradients on
local data before averaging into a global model. However, federated learning is intended to produce a
single centralized model without mixing data silos and thus does not provide any mechanism for access
control or deletion. Some work has combined federated learning and PEFT methods [
          <xref ref-type="bibr" rid="ref32 ref33 ref34">45, 46, 47</xref>
          ], but
these do not address adapter mixing, deletion, or data silos as distinct knowledge sources. Furthermore,
the privacy benefits of federated learning are unclear when applied to LLMs with large capacity for
memorization [
          <xref ref-type="bibr" rid="ref35 ref36 ref37 ref38">48, 49, 50, 51</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this paper, we introduced AdapterSwap, a parameter eficient approach to continuous learning with
access-control and data removal guarantees. We fine-tuned adapters using four modern pretrained
language models on separate domains. We showed that knowledge from specific domains can be
masked via access-control by preventing a retriever from accessing the controlled adapter at inference.
The multiple-adapter scheme also enables eficient knowledge removal via data deletion and adapter
retraining. The non-parametric behavior of AdapterSwap enables knowledge from the past to be
retained, and we showed that AdapterSwap outperforms both chronological fine-tuning and retraining.</p>
      <p>AdapterSwap enables a rich set of future research opportunities. We would like to directly improve
the approach by exploring better retrieval and adapter mixing methods. Additionally, AdapterSwap
could be further scaled to specific down-stream tasks such as question answering and directly compared
or combined with alternative data management schemes such as a RAG framework.
in: P. Koehn, L. Barrault, O. Bojar, F. Bougares, R. Chatterjee, M. R. Costa-jussà, C. Federmann,
M. Fishel, A. Fraser, M. Freitag, Y. Graham, R. Grundkiewicz, P. Guzman, B. Haddow, M. Huck,
A. Jimeno Yepes, T. Kocmi, A. Martins, M. Morishita, C. Monz, M. Nagata, T. Nakazawa, M. Negri,
A. Névéol, M. Neves, M. Popel, M. Turchi, M. Zampieri (Eds.), Proceedings of the Seventh
Conference on Machine Translation (WMT), Association for Computational Linguistics, Abu Dhabi,
United Arab Emirates (Hybrid), 2022, pp. 1–45. URL: https://aclanthology.org/2022.wmt-1.1.
[27] A. Liška, T. Kočiský, E. Gribovskaya, T. Terzi, E. Sezener, D. Agrawal, C. de Masson d’Autume,
T. Scholtes, M. Zaheer, S. Young, E. G.-M. S. Austin, P. Blunsom, A. Lazaridou, Streamingqa: A
benchmark for adaptation to new knowledge over time in question answering models, arXiv
preprint arXiv:2205.11388 (2022).
[28] G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei,
J. Launay, The refinedweb dataset for falcon llm: Outperforming curated corpora with web data,
and web data only, 2023. arXiv:2306.01116.
[29] J. Banks, T. Warkentin, Gemma: Introducing new state-of-the-art open models, Google (2024).</p>
      <p>URL: https://blog.google/technology/developers/gemma-open-models/.
[30] H. Touvron, L. Martin, K. R. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra,
P. Bhargava, S. Bhosale, D. M. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu,
J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. S. Hartshorn, S. Hosseini,
R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. M. Kloumann, A. V. Korenev, P. S. Koura,
M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra,
I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M.
Smith, R. Subramanian, X. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov,
Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, T. Scialom,
Llama 2: Open foundation and fine-tuned chat models, ArXiv abs/2307.09288 (2023). URL: https:
//api.semanticscholar.org/CorpusID:259950998.
[31] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand,
G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril,
T. Wang, T. Lacroix, W. E. Sayed, Mistral 7b, 2023. arXiv:2310.06825.
[32] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M.
Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger,
M. Drame, Q. Lhoest, A. M. Rush, Huggingface’s transformers: State-of-the-art natural language
processing, 2020. arXiv:1910.03771.
[33] S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, S. Paul, B. Bossan, Peft: State-of-the-art
parametereficient fine-tuning methods, https://github.com/huggingface/peft, 2022.
[34] S. Feng, W. Shi, Y. Bai, V. Balachandran, T. He, Y. Tsvetkov, Cook: Empowering general-purpose
language models with modular and collaborative knowledge, arXiv preprint arXiv:2305.09955
(2023).
[35] Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, N. Zhang, Editing large language
models: Problems, methods, and opportunities, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings
of the 2023 Conference on Empirical Methods in Natural Language Processing, Association for
Computational Linguistics, Singapore, 2023, pp. 10222–10240. URL: https://aclanthology.org/2023.
emnlp-main.632. doi:10.18653/v1/2023.emnlp-main.632.
[36] N. De Cao, W. Aziz, I. Titov, Editing factual knowledge in language models, in: M.-F. Moens,
X. Huang, L. Specia, S. W.-t. Yih (Eds.), Proceedings of the 2021 Conference on Empirical Methods in
Natural Language Processing, Association for Computational Linguistics, Online and Punta Cana,
Dominican Republic, 2021, pp. 6491–6506. URL: https://aclanthology.org/2021.emnlp-main.522.
doi:10.18653/v1/2021.emnlp-main.522.
[37] K. Meng, D. Bau, A. Andonian, Y. Belinkov, Locating and editing factual associations in GPT,</p>
      <p>Advances in Neural Information Processing Systems 36 (2022). ArXiv:2202.05262.
[38] E. Hernandez, B. Z. Li, J. Andreas, Inspecting and editing knowledge representations in language
models, 2023. arXiv:2304.00740.
[39] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. tau Yih,</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bapna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Firat</surname>
          </string-name>
          ,
          <article-title>Simple, scalable adaptation for neural machine translation</article-title>
          , in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>1538</fpage>
          -
          <lpage>1548</lpage>
          . URL: https://aclanthology.org/D19-1165. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1165.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giurgiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jastrzebski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Morrone</surname>
          </string-name>
          , Q. de Laroussilhe,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gesmundo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Attariyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          ,
          <article-title>Parameter-eficient transfer learning for nlp</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1902</year>
          .00751.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Allen-Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Lora:
          <article-title>Low-rank adaptation of large language models</article-title>
          ,
          <source>CoRR abs/2106</source>
          .09685 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2106.09685. arXiv:
          <volume>2106</volume>
          .
          <fpage>09685</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>European</given-names>
            <surname>Parliament</surname>
          </string-name>
          ,
          <article-title>Council of the European Union, Regulation (EU) 2016/679 of the European Parliament</article-title>
          and of the Council,
          <year>2016</year>
          . URL: https://data.europa.eu/eli/reg/2016/679/oj.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <article-title>Digital media outlets sue openai for copyright infringement</article-title>
          , The New York Times (
          <year>2024</year>
          ). URL: https://www.nytimes.com/
          <year>2024</year>
          /02/28/technology/openai
          <article-title>-copyright-suit-media</article-title>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V. C.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferraiolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          , et al.,
          <article-title>Assessment of access control systems, US Department of Commerce, National Institute of Standards and Technology</article-title>
          . . . ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Tars</surname>
          </string-name>
          ,
          <source>News over a chatbot</source>
          ,
          <year>2024</year>
          . URL: https://hellotars.com/chatbot-templates/
          <article-title>media-publication /r1FvBF/news-over-a-chatbot.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kocetkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Ben</given-names>
            <surname>Allal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Muñoz Ferrandis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jernite</surname>
          </string-name>
          , M. Mitchell, S. Hughes,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          , L. von
          <string-name>
            <surname>Werra</surname>
          </string-name>
          , H. de Vries,
          <article-title>The stack: 3 tb of permissively licensed source code</article-title>
          ,
          <source>Preprint</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>M. M. Grynbaum</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Mac</surname>
          </string-name>
          ,
          <article-title>The times sues openai and microsoft over a.i. use of copyrighted work</article-title>
          , The New York Times (
          <year>2023</year>
          ). URL: https://www.nytimes.com/
          <year>2024</year>
          /02/28/technology/openai
          <article-title>-copyr ight-suit-media</article-title>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. McCloskey</surname>
            ,
            <given-names>N. J.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          ,
          <article-title>Catastrophic interference in connectionist networks: The sequential learning problem</article-title>
          ,
          <source>Psychology of Learning and Motivation</source>
          <volume>24</volume>
          (
          <year>1989</year>
          )
          <fpage>109</fpage>
          -
          <lpage>165</lpage>
          . URL: https://www. sciencedirect.com/science/article/pii/S0079742108605368. doi:https://doi.org/10.1016/S007 9-
          <issue>7421</issue>
          (
          <issue>08</issue>
          )
          <fpage>60536</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Zhang,</surname>
          </string-name>
          <article-title>An empirical study of catastrophic forgetting in large language models during continual fine-tuning</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2308</volume>
          .
          <fpage>08747</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rush</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Parameter-eficient transfer learning with dif pruning</article-title>
          , in: C.
          <string-name>
            <surname>Zong</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>4884</fpage>
          -
          <lpage>4896</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>378</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>378</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Y.-L. Sung</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Nair</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          <string-name>
            <surname>Rafel</surname>
          </string-name>
          ,
          <article-title>Training neural networks with fixed sparse masks</article-title>
          , in: M.
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Vaughan</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>34</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2021</year>
          , pp.
          <fpage>24193</fpage>
          -
          <lpage>24205</lpage>
          . URL: https://proc eedings.neurips.cc/paper_files/paper/2021/file/cb2653f548f8709598e8b5156738cc51-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Constant</surname>
          </string-name>
          ,
          <article-title>The power of scale for parameter-eficient prompt tuning</article-title>
          , in: M.
          <article-title>-</article-title>
          <string-name>
            <surname>F. Moens</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , S. W.-t. Yih (Eds.),
          <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>3045</fpage>
          -
          <lpage>3059</lpage>
          . URL: https://aclanthology.org /
          <year>2021</year>
          .emnlp-main.
          <volume>243</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlp-main.
          <volume>243</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Prefix-tuning: Optimizing continuous prompts for generation</article-title>
          , in: C.
          <string-name>
            <surname>Zong</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>4582</fpage>
          -
          <lpage>4597</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>353</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>353</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wortsman</surname>
          </string-name>
          , G. Ilharco,
          <string-name>
            <given-names>S. Y.</given-names>
            <surname>Gadre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Roelofs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gontijo-Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Morcos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Namkoong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Carmon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kornblith</surname>
          </string-name>
          , L. Schmidt,
          <article-title>Model soups: averaging weights of multiple finetuned models improves accuracy without increasing inference time</article-title>
          , in: K. Chaudhuri,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jegelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Szepesvari</surname>
          </string-name>
          , G. Niu, S. Sabato (Eds.),
          <source>Proceedings of the 39th International Conference on Machine Learning</source>
          , volume
          <volume>162</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>23965</fpage>
          -
          <lpage>23998</lpage>
          . URL: https://proceedings.mlr.press/v162/wortsman22a.html.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pfeifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kamath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rücklé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>AdapterFusion: Non-destructive task composition for transfer learning</article-title>
          , in: P. Merlo,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tiedemann</surname>
          </string-name>
          , R. Tsarfaty (Eds.),
          <source>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics:</source>
          Main Volume,
          <article-title>Association for Computational Linguistics</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>487</fpage>
          -
          <lpage>503</lpage>
          . URL: https: //aclanthology.org/
          <year>2021</year>
          .eacl-main.
          <volume>39</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .eacl-main.
          <volume>39</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rücklé</surname>
          </string-name>
          , G. Geigle,
          <string-name>
            <given-names>M.</given-names>
            <surname>Glockner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pfeifer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>AdapterDrop: On the eficiency of adapters in transformers</article-title>
          , in: M.
          <article-title>-</article-title>
          <string-name>
            <surname>F. Moens</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , S. W.-t. Yih (Eds.),
          <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Online and
          <string-name>
            <given-names>Punta</given-names>
            <surname>Cana</surname>
          </string-name>
          , Dominican Republic,
          <year>2021</year>
          , pp.
          <fpage>7930</fpage>
          -
          <lpage>7946</lpage>
          . URL: https://aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>626</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .emnlp -main.626.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Awadallah</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Gao,</surname>
          </string-name>
          <article-title>AdaMix: Mixtureof-adaptations for parameter-eficient model tuning</article-title>
          , in: Y.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kozareva</surname>
          </string-name>
          , Y. Zhang (Eds.),
          <source>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Abu Dhabi, United Arab Emirates,
          <year>2022</year>
          , pp.
          <fpage>5744</fpage>
          -
          <lpage>5760</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .emnlp-main.
          <volume>388</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .emnlp-main.
          <volume>388</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chronopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dodge</surname>
          </string-name>
          ,
          <article-title>Eficient hierarchical domain adaptation for pretrained language models</article-title>
          , in: M.
          <string-name>
            <surname>Carpuat</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.-C. de Marnefe</surname>
            ,
            <given-names>I. V.</given-names>
          </string-name>
          <string-name>
            <surname>Meza Ruiz</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the</source>
          <year>2022</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics</article-title>
          , Seattle, United States,
          <year>2022</year>
          , pp.
          <fpage>1336</fpage>
          -
          <lpage>1351</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .naacl-main.
          <volume>96</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .n aacl-main.
          <volume>96</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Chronopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fraser</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Dodge,</surname>
          </string-name>
          <article-title>AdapterSoup: Weight averaging to improve generalization of pretrained language models</article-title>
          , in: A.
          <string-name>
            <surname>Vlachos</surname>
          </string-name>
          , I. Augenstein (Eds.),
          <source>Findings of the Association for Computational Linguistics: EACL</source>
          <year>2023</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Dubrovnik, Croatia,
          <year>2023</year>
          , pp.
          <fpage>2054</fpage>
          -
          <lpage>2063</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .findings-eacl.
          <volume>153</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .findings-eacl.
          <volume>153</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          , Sentence-BERT:
          <article-title>Sentence embeddings using Siamese BERT-networks</article-title>
          , in: K. Inui,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          Wan (Eds.),
          <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>3982</fpage>
          -
          <lpage>3992</lpage>
          . URL: https://aclanthology.org/D19-1410. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -1410.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R.</given-names>
            <surname>Aharoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          ,
          <article-title>Unsupervised domain clusters in pretrained language models</article-title>
          , in: D.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Schluter</surname>
          </string-name>
          , J. Tetreault (Eds.),
          <article-title>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>7747</fpage>
          -
          <lpage>7763</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>692</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .
          <article-title>a cl-main</article-title>
          .
          <volume>692</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Diao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. Han,
          <string-name>
            <surname>T</surname>
          </string-name>
          . Zhang, Lisa:
          <article-title>Layerwise importance sampling for memory-eficient large language model fine-tuning</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2403</volume>
          .
          <fpage>17919</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          . URL: http://jmlr.org/papers/v21/
          <fpage>20</fpage>
          -
          <lpage>074</lpage>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kocmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bawden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bojar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dvorkovich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Federmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fishel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gowda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Graham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Grundkiewicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Haddow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Knowles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Koehn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Monz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Morishita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nagata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nakazawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Novák</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Popel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Popović</surname>
          </string-name>
          ,
          <article-title>Findings of the 2022 conference on machine translation (WMT22), T</article-title>
          . Rocktäschel,
          <string-name>
            <given-names>S.</given-names>
            <surname>Riedel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiela</surname>
          </string-name>
          ,
          <article-title>Retrieval-augmented generation for knowledge-intensive nlp tasks</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <year>2005</year>
          .11401.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Retrievalaugmented generation for large language models: A survey</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2312</volume>
          .
          <fpage>10997</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>N. F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hewitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Paranjape</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bevilacqua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Lost in the middle: How language models use long contexts</article-title>
          ,
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>12</volume>
          (
          <year>2023</year>
          )
          <fpage>157</fpage>
          -
          <lpage>173</lpage>
          . URL: https://api.semanticscholar.org/CorpusID:259360665.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>S.</given-names>
            <surname>Barnett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kurniawan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thudumu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Brannelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdelrazek</surname>
          </string-name>
          ,
          <article-title>Seven failure points when engineering a retrieval augmented generation system</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2401</volume>
          .
          <fpage>05856</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Reyes</surname>
          </string-name>
          ,
          <article-title>Navigating retrieval augmented generation (rag) challenges and opportunities</article-title>
          ,
          <source>Flybridge</source>
          (
          <year>2024</year>
          ). URL: https://www.flybridge.com/ideas/navigating
          <article-title>-retrieval-augmented-generat ion-rag-challenges-and-opportunities.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>B.</given-names>
            <surname>McMahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hampson</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. A. y Arcas</surname>
          </string-name>
          ,
          <article-title>Communication-eficient learning of deep networks from decentralized data</article-title>
          ,
          <source>in: Artificial intelligence and statistics</source>
          , PMLR,
          <year>2017</year>
          , pp.
          <fpage>1273</fpage>
          -
          <lpage>1282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-L.</given-names>
            <surname>Mok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Client-customized adaptation for parameter-eficient federated learning</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd-Graber</surname>
          </string-name>
          , N. Okazaki (Eds.),
          <source>Findings of the Association for Computational Linguistics: ACL</source>
          <year>2023</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>1159</fpage>
          -
          <lpage>1172</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .findings-acl.
          <volume>75</volume>
          . doi:
          <volume>10</volume>
          .18653/v 1/
          <year>2023</year>
          .findings-acl.
          <volume>75</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>S.</given-names>
            <surname>Babakniya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Elkordy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Ezzeldin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-B. Song</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>El-Khamy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Avestimehr</surname>
          </string-name>
          , Slora:
          <article-title>Federated parameter eficient fine-tuning of language models</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2308</volume>
          .
          <fpage>06522</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Xu,</surname>
          </string-name>
          <article-title>FedPETuning: When federated learning meets the parameter-eficient tuning methods of pre-trained language models</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd-Graber</surname>
          </string-name>
          , N. Okazaki (Eds.),
          <source>Findings of the Association for Computational Linguistics: ACL</source>
          <year>2023</year>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>9963</fpage>
          -
          <lpage>9977</lpage>
          . URL: https://aclanthology.org/
          <year>2023</year>
          .findings-acl.
          <volume>632</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .findings-acl.
          <volume>632</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Recovering private text in federated learning of language models</article-title>
          , in: A. H.
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Belgrave</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          Cho (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          ,
          <year>2022</year>
          . URL: https://openreview.net/forum?id=
          <fpage>dqgzfhHd2</fpage>
          -.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>W.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ajith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Blevins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <article-title>Detecting pretraining data from large language models</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2310</volume>
          .
          <fpage>16789</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>K.</given-names>
            <surname>Tirumala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Markosyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aghajanyan</surname>
          </string-name>
          ,
          <article-title>Memorization without overfitting: Analyzing the training dynamics of large language models</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2205</volume>
          .
          <fpage>10770</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>N.</given-names>
            <surname>Carlini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ippolito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jagielski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tramer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Quantifying memorization across neural language models</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2202</volume>
          .
          <fpage>07646</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <article-title>• We trained each adapter for our experiments on a single 80GB A100 GPU for 10 epochs with a batch size of 4 and 5 gradient accumulation steps. We utilized the AdamW optimizer with default settings</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <article-title>• For adapters trained on C4 domains we used rank 64 LoRAs with  = 128 applied to all linear layers</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <article-title>• For the News Crawl experiment we used LoRA adapters with rank 32 and  = 64 applied to just the attention layers</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <article-title>• For all experiments we initialized adapters using a random seed of 42. We confirm [ 21]'s finding that using the same initialization is critical if mixing adapters</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>