<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Efectiveness of Portable Models versus Human Expertise under Continuous Active Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jeremy Pickens</string-name>
          <email>jpickens@opentext.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>hTomas C. Gricks III, Esq.</string-name>
          <email>tgricks@opentext.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>OpenText</institution>
          ,
          <addr-line>Denver</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        INTRODUCTION
eDiscovery is the process of identifying, preserving, collecting,
reviewing, and producing to requesting parties electronically stored
information that is potentially relevant to a civil litigation or
regulatory inquiry. Of these activities, the review component is by
far the most expensive and time consuming [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Modern, efective
approaches to document review run the gamut from pure
humandriven processes such as boolean keyword search followed by linear
review, to predominantly AI-driven approaches using various forms
of machine learning. A review process that involves a significant,
though not exclusive, supervised machine learning component is
typically referred to as technology assisted review (TAR).
      </p>
      <p>
        One of the most eficient approaches to TAR in recent years
involves a combined human-machine (IA, or intelligence
amplification) approach known as Continuous Active Learning (CAL) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
As with any TAR review, a CAL review will benefit in some
measure by overcoming the cold start problem: The machine typically
cannot begin making predictions until it has been fed some
number of training documents, aka seeds. In an early CAL approach,
initial sets of training documents were selected via human efort,
e.g., manual keyword searching. This approach to selecting seed
documents relies on human knowledge and intuition.
      </p>
      <p>Recently in the legal technology sector, another seeding
approach that does not rely on human assessment of the review
collection but is based on artificial intelligence (AI) methods and derived
from documents outside the collection has been gaining momentum.
For this technique, which is often referred to as “portable models”,
and known in the wider machine learning community as transfer
learning, initial seed documents are selected not via human input,
but by predictions from a machine learning model trained using
documents from prior maters or related datasets. Portable models
take a pure AI approach and eschew human knowledge in the cold
start seeding process.</p>
      <p>Notwithstanding the benefits asserted by the proponents of
portable models as a seed-generation technique, we are aware of
no formal or even informal studies addressing the overall impact of
portable model seeding on the eficiency of a TAR review relative to
2</p>
    </sec>
    <sec id="sec-2">
      <title>MOTIVATION</title>
      <p>Separate and apart from the inherent value of an assessment of
the impact of portable models on TAR, there are two principles
atendant to the creation of portable models that serve as a further
motivation for this study: (1) the increased regulatory pressure to
maintain personal privacy; and (2) the growing need for stringent
cyber security measures. Consideration of both principals is
generally recognized as an essential step in the development and utility
of modern AI applications, given their breadth and proliferation.</p>
      <p>
        Recent years have seen an increased scrutiny from EU and United
States regulatory agencies. Data collection and reuse is under heavy
examination as regulators seek to minimize data collection and
maximize privacy and security. Portable models are a form of data
reuse; the models would not exist were it not for the original data.
As such, there are rights and obligations around the use of the
data that goes in to training portable models, and a strong need for
clearer assessments of risk when porting models. As Bacon et al
noted [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]:
hTe use of machine learning (“ML”) models to process
proprietary data is becoming increasingly common
as companies recognize the potential benefits that
ML can provide. Many IT vendors ofer ML services
that can generate valuable insights derived from their
customer’s proprietary data and know-how. For
companies that have not yet established their own ML
expertise in-house, these services can ofer significant
business advantages. However, there may be cases
where one party owns the ML model, another party
has the business expertise, and a third party owns the
data. In such cases, significant intellectual property
(“IP”) and data protection and security risks may arise.
      </p>
      <p>Naturally, most companies that invest in building an
ML model are looking for a return on their investment.</p>
      <p>From a financial perspective, such companies focus on
using the IP laws and related IP contract terms, such
as IP assignments and license grants, to maximize
their control over the ML model and associated input
and results. Data protection laws can run counter to
these objectives by imposing an array of requirements
and restrictions on the processing of various types of
data, particularly to the extent they include personal
information. The interplay between these competing
considerations can lead to interesting results,
especially when a number of diferent parties have a stake
in the outcome.</p>
      <p>
        hTe second, perhaps more important challenge with respect to
portable models is the possibility of data leakage. In recent years,
computer security and machine learning researchers have increased
the sophistication of membership inference atacks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. These
attacks are a way of probing black box, non-transparent models to
“discover or reconstruct the examples used to train the machine
learning model” [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The basic process is that:
      </p>
      <sec id="sec-2-1">
        <title>An atacker creates random records for a target ma</title>
        <p>chine learning model served on a [portable model]
service. The atacker feeds each record into the model.</p>
        <p>
          Based on the confidence score the model returns, the
atacker tunes the record’s features and reruns it by
the model. The process continues until the model
reaches a very high confidence score. At this point,
the record is identical or very similar to one of the
examples used to train the model. After gathering
enough high confidence records, the atacker uses the
dataset to train a set of “shadow models” to predict
whether a data record was part of the target model’s
training data. This creates an ensemble of models that
can train a membership inference atack model. The
final model can then predict whether a data record was
included in the training dataset of the target machine
learning model. The researchers found that this atack
was successful on many diferent machine learning
services and architectures. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
        </p>
        <p>
          Carlini et al [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ] further elaborate on the potential for portable
models to reveal private or sensitive information:
        </p>
        <p>One such risk is the potential for models to leak details
from the data on which they’re trained. While this
may be a concern for all large language models,
additional issues may arise if a model trained on private
data were to be made publicly available. Because these
datasets can be large (hundreds of gigabytes) and pull
from a range of sources, they can sometimes contain
sensitive data, including personally identifiable
information (PII): names, phone numbers, addresses, etc.,
even if trained on public data. This raises the
possibility that a model trained using such data could reflect
some of these private details in its output.</p>
        <p>
          Tramer et al [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] note that entire models may even be stolen
via such techniques, even when the adversary only has black box
(observations of outputs only, rather than internal workings)
access to the model: “The tension between model confidentiality and
public access motivates our investigation of model extraction
attacks. In such atacks, an adversary with black-box access, but no
prior knowledge of an ML model’s parameters or training data,
aims to duplicate the functionality of (i.e., “steal”) the model…We
show simple, eficient atacks that extract target ML models with
near-perfect fidelity for popular model classes including logistic
regression, neural networks, and decision trees.”
        </p>
        <p>Given the potential dangers associated with modern AI
applications such as portable models, we therefore ask: Do portable models
provided a cognizable sustained advantage over human augmented
IA processes suficient to warrant their use in the face of privacy
and cybersecurity concerns? If not, perhaps the safer and more
appropriate approach is to continue using traditional human-driven
techniques.
3</p>
        <p>
          RELATED WORK
hTe key foundation in our investigation is the observation that the
current state-of-the-art document review TAR process is based on
continuous active learning (CAL) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Given seed documents, the
basic CAL process induces a supervised machine learning model
which then predicts the most likely responsive, unreviewed
documents. After some (relatively small) number of those top-ranked
predictions are reviewed and coded, another model is induced and
the next most likely documents are queued for review. The process
continues until a high recall target is hit.
        </p>
        <p>Review workflows that are based on CAL have what might be
called a “just in time” approach to prediction. Rather than
atempting to induce a perfect model up front, CAL workflows dynamically
adjust as the review continues. Often this means that early
disadvantages, and even early advantages, wash out in the process. For
example [citation anonymized for review] found that four searchers
each working independently to find seed documents found diferent
and diferent numbers of seeds. But after separately using each seed
set to initialize a CAL review, approximately the same number of
documents needed to be reviewed to achieve high recall. This study
asks similar questions in the context of portable models—whether
there is a significant improvement in review eficiency when using
portable models relative to traditional, non-AI techniques.</p>
        <p>
          Another common portable model theme is the claim that the
more historical data they are trained on, the beter their predictions
will be. While that may be true in some instances, it may not be in
others. What constitutes privileged documents in one mater might
have a diferent set of characteristics as privileged documents in
others mater. What constitutes evidence of fraud, or sexual
harassment in one mater might be diferent than in other maters. No
amount of “big data” gathered from dozens (hundreds? thousands?)
of prior maters and composed into a monolithic portable model
may be relevant to the current problem if the paterns in the current
problem don’t match the historical ones. Therefore a question that
every eDiscovery practitioner should be asking herself is where
the best source of evidence for seeding the current task lies. As [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
notes: “The real goal should not be big data but to ask ourselves,
for a given problem, what is the right data and how much of it is
needed. For some problems this would imply big data, but for the
majority of the problems much less data is necessary.”
4
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>RESEARCH QUESTIONS</title>
      <p>We engage three primary research questions. The first question
level-sets the value of the pure AI (portable model) approach. The
second two questions compare the portable model approach to a
human-initiated process.</p>
      <p>RQ1 Does a CAL review seeded by a portable model
outperform (at high recall) one seeded randomly
RQ2 Do portable models initially find more relevant
documents than does human efort
RQ3 Does a CAL review seeded by a portable model
outperform (at high recall) one seeded by human efort</p>
      <p>When atempting to consider these questions in general, issues
naturally arise: What portable models are we talking about? Trained
on what data? And how close was that data to the target
distribution? And what humans seeded the comparison approach? And
what was their prior knowledge of the subject mater?
hTese questions mater, and while we cannot answer them for
every possible training set and human searcher, we have structured
the experiments in such a way as to give the most possible “benefit
of the doubt” to the portable model, and the least possible benefit
to the human searcher. Thus if there are significant advantages of
portable models over human efort, these should be most readily
apparent when portable models are given the most afordances and
humans the least.</p>
      <p>hTe primary manner in which portable models are given an
advantage is that we train them on a set of documents that is drawn
from the exact same distribution as the target collection to which
they will be applied. In practice, portable models are never given
this advantage. Prior cases in eDiscovery are not always exactly the
same. Diferent collections, even from the same corporate entity,
exhibit diferent distributions, especially as employees and business
activities change and evolve over time. Naturally, the more diferent
the source distribution, the less efective portable models will be
when applied to a new target collection. However, by holding the
distribution the same, this gives us an upper bound on portable
model efectiveness and establishes a strong baseline against which
the human efort can be compared.</p>
      <p>
        At the same time, the human efort is minimized. As will be
described in more detail below, a small team of human searchers
worked for a collective total of approximately half an hour per
topic. None of the humans were experts in any of the topics, nor
did anyone have recent prior knowledge on the topics, as the events
in this Jeb Bush TREC collection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] took place a decade or more
prior to when the searchers worked and most of the issues were
local to Florida and did not make national news. In practice, humans
are rarely given this disadvantage. They often work for more than
thirty minutes on a problem and can have broad domain expertise
that comes from having worked on similar cases in the past.
      </p>
      <p>hTus, our experiments consist of a comparison between portable
models trained in the best possible light vs human efort that is
kept at a minimum. We do this because the core concept of portable
models is that they will be suficiently broad in scope so as to be
able to identify relevant documents in a collection that contains
documents of a similar content and context to those on which they
were trained. (“Relevance” here refers to the notion of “what is
desired”, be it some sort of topical similarity such as age
discrimination or fraud cases, or something like privilege.) That distributional
similarity is not always guaranteed, and in fact it can be dificult a
priori to know whether you a portable model has been trained on
data similar enough to be useful. By using documents intentionally
drawn from the exact same distribution, we are able to show an
upper bound on portable model efectiveness. In practice, portable</p>
      <sec id="sec-3-1">
        <title>Topic</title>
      </sec>
      <sec id="sec-3-2">
        <title>Collection Stats</title>
        <p>Total Rel Richness Queries</p>
      </sec>
      <sec id="sec-3-3">
        <title>Manual Seeding Stats</title>
        <p>
          Minutes Total Docs
We test these research questions using the TREC 2016 total recall
track document collection, topics, and relevance judgments [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. This
dataset contains 34 topics each with a varying number of relevant
documents. Nonetheless, the richness of the majority of topics is
under 1%, i.e. relatively low richness topics where portable models
allege to be most efective. Table 1 contains statistics on each topic.
hTe first column is the topic ID from 401 to 434, sorted in a manner
that will be described in Section 6.3. The next two columns contain
the number of total relevant documents and the richness for each
topic. There are 290,099 total documents in the collection.
        </p>
        <p>Human efort, aka manual seeding, was done with a small team
of four searchers. For each topic, two of the searchers were
instructed to run a single query and code the first 25 documents that
resulted from that query. The other two searchers were given more
interactive leeway and were instructed to utilize as many searches
and whatever other analytic tools (clustering, timeline views, etc.)
as they wanted, with a goal of working for about 15-30 minutes and
stopping once they had tagged 25 documents. This was not strictly
controlled, and some reviewers worked a few minutes longer, some
a few minutes shorter. And some marked a few more than 25
documents, and some a few less, as is to be expected in normal, “in
the moment” flow of knowledge work. Table 1 contains the manual
efort statistics — with the total number of queries, total number of
minutes, and total unique documents tagged as either relevant or
non-relevant — for each topic. On average, the human reviewers
worked for 31.4 minutes, issued 9.9 queries, and coded 64.3
documents so the overall efort was done at a fairly high pace and was
relatively minimal in comparison to the size of the collection.
5.2</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiment Structure</title>
      <p>
        In order to compare portable models against both random and
human-seeded techniques in RQ1 through RQ3, there needs to be
a collection on which the portable model can be trained, separate
from the collection on which it and the comparative approaches
are deployed. We will refer to these two collections as “source” and
“target’, respectively. For the reasons enumerated in Section 4, we
carve out the portable model training source collection from the
same distribution as the target collection, and do so by selecting
documents at random. For a given topic, we:
(1) Shufle the collection randomly
(2) Split the collection into k groups
(3) For each group:
(a) Use that group as the portable model training “source”
collection S
(b) Train a model M using every document in S
(c) Use the remaining groups as the “target” collection T
(d) Select manual (human) seeds H by intersecting all found
docs (see Table 1) with T
(e) Select random seeds R from T until five positives
examples are found
(f) Selected portable seeds P from the top of the M-induced
ranking on T in an amount equal to jH j
(g) Use the appropriate seeds to run each experiment RQ1
through RQ3
(4) Average results across all k groups for the topic, but do not
average across topics
hTe specifics of step (3g) depends on the research question being
tested. For example, for RQ1, R and P are each (separately) used to
seed a continuous active learning (CAL) process. For RQ3, H and
P are used. Other than the diferent seedings, these CAL processes
are run exactly as in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] except that updates are done every 30
documents rather than every 1000 documents. And unlike some
approaches, the learning is not relevance feedback for a limited
number of steps. It is truely continuous in that it does not stop until
the desired recall level is achieved, which in these experiments are
set to 80%.
      </p>
      <p>While selecting a source collection for portable model training
from the same distribution as the target collection already ofers
great advantage to the predictive capabilities of a portable model,
i.e. puts it above where it would likely perform in more realistic
scenarios, we extend this advantage even further by giving the
model larger and larger source collections on which to train. We
compare three primary source/target partitions: 20/80, 50/50, and
80/20, with k=5, k=2, and k=5, respectively. (In the 80/20 case, steps
(3a) and (3c) are reversed, with the current group used as the target
collection and the other groups used as the source collection.) The
reason for the 20/80 partition is that the eDiscovery problem is
a recall-oriented task. The larger the review population, aka the
target collection, the more realistic the CAL process is likely to be.
However, the disadvantage is that only 20% of the TREC collection is
be used for training the portable model. The 80/20 partition reverses
the balance: 80% of the collection is used to train the portable model,
but only 20% of the collection is available to simulate the CAL
review, which can be problematic for especially sparse topics. The
50/50 partition splits the diference.</p>
      <p>Comment: An astute observer may find slight fault with the
structure of this experimental setup, in that there is a small amount of
knowledge overlap between the source and target partitions when
doing human seeding. Specifically, the human searchers originally
searched across the entire collection rather than across split
collections. It is possible that a document found by a human searcher
that ended up in a source partition has led the human to issue a
query that found more or beter documents that ended up in the
target partition. Thus even though the human-found documents
in only the target partition are used to seed a CAL process (Step
3d, above), the existence of some of those seeds could have been
influenced by knowledge of documents in the source partition. We
note this issue and make it explicit, but do not think that it afects
the overall conclusions of the experiment. One reason is that even if
humans had some knowledge of documents in the source partition
when finding the documents in the target partition, the portable
model M is given knowledge of every document, positive and
negative, in the source partition. Table 3 shows the raw number of
positive documents used for training in the source partition, and
it swamps the documents that the humans would have looked at
in their short search sessions. Thus, one can think of any overlap
during human seed selection as the background knowledge that
human would likely already be expected to possess when working
in a real scenario. E.g. an investigator working on detecting fraud or
sexual harassment likely has some implicit background knowledge
of fraud or sexual harassment.
6
6.1</p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS</title>
      <p>RQ1: Portable- vs Random-Seeded Recall
hTe results for our first question are found Table 2. Under the rubric
of symmetry, the results are expressed in terms of raw percentage
point (not percentage) diferences between the precision achieved
at 80% recall for the portable model P-seeded review versus a
random R-seeded review, and averaged across all partitions for each
topic. Positive values indicate beter portable model performance;
negative values the opposite. Nearly universally, P-seeding
outperforms random seeding; the p-value under a binomial test is &lt;
0.00001.</p>
      <p>hTis is a wholly expected result. Closer examination of the
simulated review orderings shows that in low richness domains most of
the precision loss comes not from the CAL iterations, but from the
larger number of documents needed to find enough positive ones
to start ranking. Note that random seeding on the target partition
also outperforms fully linear review by an average of 12.2, 10.6,
and 5.6 percentage points on the 20/80, 50/50, and 80/20 partitions,
respectively. So even random seeding of CAL is beter than no CAL
at all.
Topic
20/80 Partition
50/50 Partition
precision
80/20 Partition</p>
      <p>hTus in answer to the question: Does P-seeding produce an
eficacious result, the answer is yes. However, the more important
question is not whether portable models are useful; it is whether
they are useful relative to other reasonable, simpler, or less risky
alternatives. For that we turn to the remaining research questions.
6.2</p>
    </sec>
    <sec id="sec-6">
      <title>RQ2: Portable vs Human Seed Initial</title>
    </sec>
    <sec id="sec-7">
      <title>Relevance</title>
      <p>hTe results for our second questions are found in Table 3 under the
Target Portable and Target Manual columns for each partition group.
For example, in the 20/80 partition, where 80% of the collection is
used as the target collection and on average across all 5 folds, on
topic 403 human efort found 18.4 positively-coded seed documents
whereas the portable model M found 41.6 at the same level of efort
(i.e. at 60 documents, as per Table 1). On topic 421 under the 20/80
partition, humans found an average 10.4 documents and M found
3.2.</p>
      <p>hTe average number of positive training documents in the source
partition, i.e. the data on which M is trained, is shown. The number
of negative training examples is the remainder of the fold. Averages
across all 34 topics are shown at the botom of the table, as is a
binomial p-value.</p>
      <p>hTese results show that when 20% of the collection is used to
train M (20/80 parition), even though that data is literally from
the same distribution as the target fold, the various M are able
to find seed documents at a rate no beter than a small amount of
human efort. There is only a diference of 0.2 documents across
all topics, and while there is some variation between topics the
diferences are not statistically significant (p=0.303). As the training
partition increases, and 50% then 80% of the collection is used to
train each M, so too does the ability of the model to find more seed
documents. On the 50/50 partition M finds on average 3.5 more
documents than the human at the given efort level, and on the
80/20 partition it finds an average 1.5 more documents. Both results
are statistically significant.</p>
      <p>When the number of seeds is normalized per fold and topic by
the number of total seeds found, i.e. the positive seed precision,
the manual efort has an average precision of 66.0% across all folds,
whereas M precision is 64.2%, 76.5%, and 78.1%, respectively across
20/80, 50/50, and 80/20. That is, even though the average number of
documents that M finds on the 50/50 partition is larger (3.5) than
on the 80/20 partition (1.5), the later partition is smaller. The actual
precision goes up slightly.</p>
      <p>hTus in answer to the question: Do portable models initially find
more relevant documents than does human efort, the answer is
mixed. When given 20% of the collection for training, they do not.
When given 50% or 80%, they do. However, the improvement is
modest: a few percentage points, or a few extra documents.
6.3</p>
    </sec>
    <sec id="sec-8">
      <title>RQ3: Portable- vs Human-Seeded Recall</title>
      <p>hTe results for our third and final question are also found in Table 3
under the precision columns. Again, in the interest of symmetric
magnitudes, precision is the percentage point diference between
P-seeded versus H -seeded CAL. While again there is some
variation across topics, on average on the 20/80 partition, P-seeding is
0.4 percentage points worse, while on the 50/50 and 80/20 partitions
P-seeding is 0.6 and 1.2 percentage points beter. However, none
of these results are statistically significant (p=0.303).</p>
      <p>Furthermore, when we look at the raw document count
diference between the two conditions (not shown in the table) another
story emerges. In the 80/20 partition, on those topics for which
Pseeded CAL is beter, is it beter on average by 186 total documents.
Where H -seeded CAL is beter, it is beter by 896 documents. On
the 50/50 partition, P- versus H -seeding is 271 vs 743 documents
beter, and on the 20/80 partition it is 166 versus 655 documents.
hTere is no consistent advantage of either approach over the other,
but the negative consequences of H seeding seems to be far smaller
than those of P seeding.</p>
      <p>
        hTe reason for the topical sort order across all tables should now
become clear: All tables in this paper are sorted by the precision
of the 80/20 partition. This seems to be the partition for which
portable models are the strongest; they have the most training data.
And sorting by precision allows us to see where P-seeding vs
H -seeding each shine. To that end, we introduce one more metric
into the discussion: The WTF “inefectiveness” metric [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]
encapsulates the notion of not only looking at average performance,
but at outliers. A system that has good average performance but
egregious outliers might want to be avoided, especially in
eDiscovery where every case maters and the costs incurred by an outlier
are more significant than in, say, ad hoc web search.
      </p>
      <p>From this perspective, we see that where portable models
perform the strongest, i.e. on the 80/20 partition where they are given
80% of the available positive documents, there are outliers in both
directions. The top three P-advantage outliers show a 60.5, 12.0,
and 5.7 percentage point diference. The top three H -advantage
outliers show a 25.0, 12.6, and 11.1 percentage point diference.
However, in terms of raw document counts these translate to a 670,
667, and 468 documents for the P-advantage, and 9249, 2006, and
892 documents for H -advantage. There appear to be fewer “WTFs”
from H -seeding.
7</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>We have shown that a portable model can be useful. Certainly
relative to linear review, and even relative to randomly seeded CAL
workflows, taking a portable approach ofers a significant
advantage. They are also marginally beter than humans when it comes
to finding initial seed documents. When it comes to sustained
advantage, i.e. precision at 80% recall, the advantages fade. There is no
statistically significant diference in human vs portably seeded CAL
workflows, and slight evidence that the outliers for the portable
approach are worse.</p>
      <p>We note also that porting models carries with it significant risk
in the form of intellectual property rights, data leakage via
membership inference atacks, privacy, and security. It is every party’s own
subjective decision as to whether the advantages of portable models
outweigh the challenges and risk. However, from the results in this
study we would recommend continuing to invest in human-driven
seeding (IA–intelligence augmentation) processes and not going
all in on AI. At least relative to the topics studied in this paper, the
modicum of efort required of the human are a fair trade relative to
risk. Even when portable models are built on a corporation’s own
data, and models are not swapped between diferent owners and
therefore risk is lower, we do not yet find that the portable model
provides a sustained advantage.
8</p>
    </sec>
    <sec id="sec-10">
      <title>FUTURE WORK</title>
      <p>
        Certainly this is but one study and more studies with a wider range
of models, data collections, and human efort are needed. Perhaps
the humans could have done even beter if given more time, were
working on a domain in which they had specific expertise, or were
given more powerful analytics with which to find seed documents.
Conversely, portable models were given all possible advantages
in this experimental structure by building them on documents
drawn from the exact same distribution as the target collect, in
ever increasing amounts (20%, 50%, and 80%). It is not likely that
portable models will ever be trained on prior data as perfectly
similar to the target distribution. Therefore, portable models might
likely have performed much worse in realistic scenarios where the
source and target collections are further apart. E.g. when modeling
fraud or sexual harassment, does what constitute evidence of fraud
or harassment in one collection express itself the same way in
another collection? Future research is needed in three diferent
areas: (a) More and larger collections from similar but not identical
distributions on which to train models, or perhaps the “right” small
collections on which to train models, as per [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], (b) more advanced
models and beter transfer learning, and (c) beter and stronger
baselines against which to compare.
      </p>
      <p>Beter and stronger baselines are not limited to more efective
human efort. They also include other existing, common practices. For
example, many companies dealing with sensitive information keep
lexicons of search terms used to find sensitive information. While
a lexicon could in some sense be thought of as an “unweighted”
portable model, one diference is that it’s manually constructed,
transparent, and can embed human intuition and paterns never
seen in prior data, i.e. lexicons do not need to be trained. Another
common approach for corporations with repeat litigation is the
idea of a “drop in seed”. That is, instead of building large models
based on huge datasets from all possible prior maters, some in
the industry have developed the ad hoc practice of taking a few
coded documents from previous maters, which maters are known
to be similar to the current mater, and using those as the initial
seeds on the target collection. This of course only works behind
the firewall, as companies will not transfer documents to other
companies. But given the security, privacy, and related
membership inference atack risks of portable models, companies might
not want to transfer their own models to a competitor, either. So
in addition to comparing portable models against human seeding,
they should be compared against lexicons and drop-in seeds.
Perhaps these later approaches outperform both portable models and
human-seeded approaches when considering the total cost of a
review and not just the document count.</p>
      <p>hTe cost of the portable model (vendor charge) versus the human
approach (e.g. half an hour of searcher time) needs to be
considered as well, and not just the cost of the subsequent document
review. In short, this is a rich space for the exploration of tradeofs,
risks, and advantages for various human-driven vs machine-driven
eDiscovery processes.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Britany</given-names>
            <surname>Bacon</surname>
          </string-name>
          , Tyler Maddry, and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Pateraki</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Training a Machine Learning Model Using Customer Proprietary Data: Navigating Key IP and Data Protection Considerations</article-title>
          .
          <source>Prat's Privacy and Cybersecurity Law Report 6</source>
          ,
          <issue>8</issue>
          (Oct.
          <year>2020</year>
          ),
          <fpage>233</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Ricardo</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Big Data or Right Data?</article-title>
          .
          <source>In Proceedings of the 7th Alberto Mendelzon International Workshop on Foundations of Data Management (AMW</source>
          <year>2013</year>
          ),
          <source>Loreto Bravo and Maurizio Lenzerini (Eds.)</source>
          , Vol.
          <volume>1087</volume>
          (CEUR Workshop Proceedings). Puebla/Cholula, Mexico. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1087</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Nicholas</given-names>
            <surname>Carlini</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Privacy Considerations in Large Language Models</article-title>
          .
          <source>Retrieved April 29</source>
          ,
          <year>2021</year>
          from https://ai.googleblog.com/
          <year>2020</year>
          /12/privacyconsiderations-in-large.html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Nicholas</given-names>
            <surname>Carlini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Florian</given-names>
            <surname>Tramer</surname>
          </string-name>
          , Eric Wallace, Mathew Jagielski,
          <string-name>
            <surname>Ariel</surname>
            <given-names>HerbertVoss</given-names>
          </string-name>
          , Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and
          <string-name>
            <given-names>Colin</given-names>
            <surname>Rafel</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Extracting Training Data from Large Language Models</article-title>
          . Article arXiv:
          <year>2012</year>
          .07805.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Cormack</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Grossman</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Evaluation of machine-learning protocols for technology-assisted review in electronic discovery</article-title>
          .
          <source>In Proceedings of the 37th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . Gold Coast, Australia,
          <fpage>153</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Ben</given-names>
            <surname>Dixon</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Machine Learning:</article-title>
          <source>What are Membership Inference Atacks? Retrieved April 29</source>
          ,
          <year>2021</year>
          from https://bdtechtalks.com/
          <year>2021</year>
          /04/23/machine-learningmembership
          <article-title>-inference-attacks/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Maura</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Grossman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Gordon</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Cormack</surname>
            , and
            <given-names>Adam</given-names>
          </string-name>
          <string-name>
            <surname>Roegiest</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>TREC 2016 Total Recall Track Overview</article-title>
          . In
          <source>In NIST Special Publication 500-321: The TwentyFifth Text REtrieval Conference Proceedings (TREC</source>
          <year>2016</year>
          ) ,
          <string-name>
            <surname>Ellen</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Voorhees</surname>
          </string-name>
          and Angela Ellis (Eds.). Gaithersburg, Maryland. https://trec.nist.gov/pubs/trec25/ trec2016.html
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Nicholas</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Pace</surname>
            and
            <given-names>Laura</given-names>
          </string-name>
          <string-name>
            <surname>Zakaras</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Where the Money Goes: Understanding Litigant Expenditures for Producing Electronic Discovery</article-title>
          . Rand Corporation, Santa Monica, CA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ahmed</surname>
            <given-names>Salem</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Mathias Humbert, Pascal Berrang, Mario Fritz, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Backes</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>ML-Leaks: Model and Data Independent Membership Inference Atacks and Defenses on Machine Learning Models</article-title>
          .
          <source>In Proceedings of Network and Distributed Systems Security (NDSS) Symposium</source>
          . San Diego, California.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Florian</surname>
            <given-names>Tramer</given-names>
          </string-name>
          , Fan Zhang, Ari Juels,
          <string-name>
            <given-names>Michael K.</given-names>
            <surname>Reiter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Ristenpart</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Stealing Machine Learning Models via Prediction APIs</article-title>
          .
          <source>In Proceedings of the 25th USENIX Security Symposium</source>
          . Austin, Texas,
          <fpage>601</fpage>
          -
          <lpage>618</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Tunkelang</surname>
          </string-name>
          .
          <year>2012</year>
          . WTF!
          <article-title>@k: Measuring Inefectiveness</article-title>
          .
          <source>Retrieved April 29</source>
          ,
          <year>2021</year>
          from https://thenoisychannel.com/
          <year>2012</year>
          /08/20/wtf-k-measuringineffectiveness/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Ellen</given-names>
            <surname>Voorhees</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Measuring Inefectiveness</article-title>
          .
          <source>In Proceedings of the 27th annual ACM SIGIR conference on Research and Development in Information Retrieval. Shefield</source>
          , UK,
          <fpage>562</fpage>
          -
          <lpage>563</lpage>
          . https://doi.org/10.1145/1008992.1009121
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>