<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>H. Weerts);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>and You Will Find It: Fairness-Aware Data Collection through Active Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hilde Weerts</string-name>
          <email>h.j.p.weerts@tue.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renée Theunissen</string-name>
          <email>r.h.theunissen@student.tue.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martijn C. Willemsen</string-name>
          <email>m.c.willemsen@tue.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eindhoven University of Technology</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Machine learning models are often trained on data sets subject to selection bias. In particular, selection bias can be hard to avoid in scenarios where the proportion of positives is low and labeling is expensive, such as fraud detection. However, when selection bias is related to sensitive characteristics such as gender and race, it can result in an unequal distribution of burdens across sensitive groups, where marginalized groups are misrepresented and disproportionately scrutinized. Moreover, when the predictions of existing systems afect the selection of new labels, a feedback loop can occur in which selection bias is amplified over time. In this work, we explore the efectiveness of active learning approaches to mitigate fairnessrelated harm caused by selection bias. Active learning approaches aim to select the most informative instances from unlabeled data. We hypothesize that this characteristic steers data collection towards underexplored areas of the feature space and away from overexplored areas - including areas afected by selection bias. Our preliminary simulation results confirm the intuition that active learning can mitigate the negative consequences of selection bias, compared to both the baseline scenario and random sampling.</p>
      </abstract>
      <kwd-group>
        <kwd>Learning</kwd>
        <kwd>selection bias</kwd>
        <kwd>algorithmic fairness</kwd>
        <kwd>active learning</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Machine learning models are often trained on data sets subject to
selection bias: non-random
selection of instances from the population. When selection bias is related to sensitive
characteristics such as gender and race, it can result in fairness-related harm. For example, in banking,
fraud detection models are typically trained on labeled transaction data, which can be afected
by social biases against sensitive groups. Transactions of customers belonging to those groups
may be flagged more often as suspicious, creating the appearance of a relatively high number of
fraudulent transactions compared to other groups - even when the true fraud rates are similar.
As aptly put by a quote popularly attributed to Sophocles: look and you will find it - what is
unsought will go undetected. Unaddressed, the consequences can be severe: fairness-related
harm linked to selection bias has been observed in various cases such as fraud detection in
welfare benefit applications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], predictive policing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and medical diagnosis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
(M. C. Willemsen)
CEUR
Workshop
Proceedings
      </p>
      <p>Given these implications, it seems imperative to understand and mitigate the harmful efects
of selection bias in machine learning. In this paper, we present the results of a preliminary
simulation study which explores the efectiveness of active learning approaches in mitigating
fairness-related harm caused by selection bias. Active learning approaches aim to select the
most informative instances from unlabeled data. We hypothesize that this characteristic steers
data collection towards underexplored areas of the feature space and away from overexplored
areas – including areas afected by selection bias. Our simulation results confirm the intuition
that active learning can mitigate the negative consequences of selection bias, compared to both
the baseline scenario and random sampling.</p>
      <p>The remainder of this work is structured as follows. In Section 2, we further detail the problem
of selection bias and the moral argument that motivates our work. In Section 3, we motivate
our hypothesis that active learning can be helpful to mitigate harmful efects of selection bias.
In Section 4, we provide a brief overview of related work. Section 5 details our experiment
setup and the results are presented in Section 6. Section 7 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The Unfairness of Selection Bias</title>
      <p>
        A major assumption in machine learning is that the training data is representative of the
population from which it is drawn. In practice, training data is often not sampled at random,
skewing the distribution of the training data set [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In particular, when positives are rare,
(manual) labeling of instances is expensive, and resources are limited, the selection of instances
to be labeled is typically informed by domain expertise or the output of existing statistical
models. In other words, non-random selection is leveraged to achieve reasonable predictive
performance with a small number of labeled instances. However, selection bias in historical
decision-making policies can be afected by social biases and structural injustice, resulting in an
unequal distribution of labeled data across sensitive groups [e.g., 1, 3].
      </p>
      <p>
        Both over- and underrepresentation can result in fairness-related harm. In the context of
clinical prediction models, the underrepresentation of demographic groups in clinical data sets
can lead to underdiagnosis. For example, it is well-known that women have been historically
underrepresented in clinical trials. As a result, there exists limited knowledge regarding adverse
efects, benefits, and risks of treatments for women [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which negatively afects healthcare
outcomes. In the context of fraud detection, selection bias can result in an unequal distribution
of burdens, where marginalized groups are overrepresented and disproportionately scrutinized.
For example, in 2021, it was brought to light that the Dutch Tax and Customs Administration had
unlawfully processed Dutch citizenship in fraud investigations of childcare benefit applications
and even explicitly included Dutch citizenship as a risk factor in the risk assessment model
[
        <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
        ]. As a result, applicants who had a nationality other than Dutch were overly scrutinized
through manual processing of the application and many of them were wrongfully accused of
fraud.
      </p>
      <p>
        The adverse consequences of selection bias are further amplified when the predictions of
existing systems afect the selection of new labels, resulting in a feedback loop. For example, when
a fraud detection model is trained on a data set in which a sensitive group is overrepresented,
the model is more likely to generate alerts of potentially fraudulent activities for members of
this group, resulting in higher testing and a relatively larger number of positives. When the
model is retrained on the newly available data, selection bias is reinforced. This type of feedback
loop has been previously characterized in diferent domains [e.g., 6], most prominently in the
context of criminal justice [
        <xref ref-type="bibr" rid="ref10 ref2 ref7 ref8 ref9">2, 7, 8, 9, 10</xref>
        ].
      </p>
      <p>
        If the goal is to achieve reasonable predictive performance, some form of selection bias seems
unavoidable under limited resources.1 However, from a moral perspective, a disproportionate
distribution of burdens or benefits caused by the misrepresentation of sensitive groups in the
training data violates basic principles of equality [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Moreover, as shown by the case of the
Dutch childcare benefit scandal, a failure to mitigate this form of selection bias may not survive
legal scrutiny under EU law.2
      </p>
      <p>In conclusion, we argue that practitioners have a responsibility to mitigate the harmful efects
of selection bias related to sensitive group membership.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Fairness-Aware Data Collection through Active Learning</title>
      <p>
        Active learning is a form of semi-supervised machine learning, where the objective is to achieve
high predictive performance with a small number of labeled samples [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. To this end, active
learning methods iteratively query an oracle (e.g., a human domain expert) to label specific data
points that are expected to improve the predictive performance of the machine learning model.
Once the labels are obtained, the model is retrained on a data set including the newly labeled
data. This process is repeated several times. Active learning is typically used when it is too
costly or time-consuming to label all data instances manually.
      </p>
      <p>Active learning approaches can generally be divided into two main categories. Pool-based
approaches select one instance (screening-based) or multiple instances (batch-mode) from a
pool of unlabeled data instances. In streaming settings, streaming-based approaches determine
whether a new instance is labeled or discarded on a one-by-one basis as new data arrives. In this
setting, unlabeled instances are not revisited. In this work, we focus primarily on pool-based
approaches, as this setting best matches the typical scenarios where selection bias can occur,
such as fraud detection.</p>
      <p>
        An important component of pool-based active learning approaches is the sampling approach
that determines which sample of data instances will be selected for labeling by the oracle. One
of the most common sampling approaches is uncertainty sampling, in which data instances are
ranked based on how uncertain the algorithm is regarding the true label of the data instance,
typically informed by the confidence score of the machine learning model [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Returning to the problem of selection bias, we hypothesize that active learning could
potentially mitigate fairness-related harm caused by selection bias related to sensitive group
membership. In particular, uncertainty sampling is expected to steer data collection towards
underexplored areas of the feature space and away from overexplored areas – including areas
1Of course, the mere notion of limited resources and the perceived value of predictive performance already embed
important value judgments. In particular, a straightforward alternative mitigation strategy would be to increase the
amount of available resources such that random selection is feasible.
2We would like to emphasize that EU law is highly contextual and judicial decisions related to the Dutch childcare
benefits scandal cannot be readily applied to all cases of unfairness caused by selection bias. We refer to Weerts
et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] for a more elaborate overview of EU non-discrimination law in the context of algorithmic fairness.
afected by selection bias. For example, consider a scenario where some group  is historically
overrepresented in a data set compared to group  . A model trained on this data set is likely to
produce more uncertain confidence scores for instances in group  compared to instances that
belong to group  . As a result, the active learning algorithm is more likely to query instances
in group  , counteracting selection bias.
      </p>
      <p>
        Previous research points towards the potential efectiveness of active learning in the context
of unfairness caused by selection bias. While Branchaud-Charron et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] do not explicitly
study the problems of selection bias and feedback loops, the authors do find that active learning
approaches can efectively improve fairness measured by several fairness metrics compared to
random sampling. Additionally, Richards et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] do not explicitly consider fairness, but do
show that active learning is more efective at mitigating selection bias unrelated to sensitive
group membership than other techniques such as importance weighting. While neither of these
works fit our setting exactly, the results are promising.
      </p>
      <p>Active learning could have several advantages compared to existing fairness-aware machine
learning (fair-ml) approaches. First of all, selection bias is addressed directly at its source:
during data collection. As such, we expect the approach to be more efective compared to
technical interventions that address harmful consequences through fairness constraints during
training. Second, if active learning is shown to consistently and efectively mitigate selection
bias, this implies that unfairness can be mitigated without having access to sensitive features
– an important concern in many eforts towards fair outcomes. However, we would like to
emphasize that even if active learning is efective, it will not be a panacea. Selection bias is
primarily a concern in scenarios where social bias and structural injustice are prevalent. In
such contexts, selection bias is unlikely to be the only source of downstream unfairness and
additional interventions are necessary to ensure equitable outcomes.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Related Work</title>
      <p>Researchers have developed a plethora of fair-ml approaches aimed at mitigating unfairness,
ranging from pre-processing the data that obscure undesirable associations [e.g., 17], enforcing
fairness constraints during model training [e.g., 18], and post-processing (predictions of) existing
models [e.g., 19]. A common denominator of these approaches is that fairness is formulated as
an optimization task, where the objective is to achieve high predictive performance under some
quantitative fairness constraint, such as a maximum diference in error rates across sensitive
groups.</p>
      <p>
        Our work is related to sampling-based pre-processing techniques [e.g., 20], which use specific
sampling schemes to satisfy particular fairness constraints. However, diferent from our work,
these approaches typically assume a fully labeled training data set. More closely related to our
work are fair-ml approaches that adapt or complement existing active learning approaches.
For example, the Fairness-aware Active Learning (FAL) framework [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] aims to create a
balanced data set by selecting new data instances based on total accuracy and improvement of
demographic parity and equalized odds. Similarly, Parity-Constrained Meta Active Learning
(PANDA) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] uses meta-learning to learn a selection policy that optimizes accuracy and
fairness constraints, outperforming random sampling, uncertainty sampling, and FAL in terms of
accuracy and two commonly used fairness metrics.
      </p>
      <p>
        All of the above-mentioned fair-ml approaches are proposed as generic tools to mitigate
unfairness via quantitative fairness constraints and most of them do not make explicit
assumptions about the nature of the bias that lies at the root of unfairness. We argue that formulating
fairness as a black-box optimization task has several limitations. While fairness metrics can be
useful indicators of potential fairness-related harm, they often fail to capture more nuanced
notions of equality, rendering them poor optimization constraints [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Moreover, fair-ml
approaches attempt to address fairness primarily during the modeling stage of the development
process, which often fails to meaningfully address the biases and design choices at the root of
fairness-related harm [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. In contrast to traditional fair-ml approaches, our work thus falls
in a line of recent work [e.g., 8, 9, 24] that shifts focus towards understanding and mitigating
specific types of biases directly.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Simulation Setup</title>
      <p>
        We use a simulation study to explore the potential efectiveness of active learning for mitigating
the harmful efects of selection bias related to sensitive group membership. The source code of
the simulation is available on Github.3
Data Set and Pre-processing The data set that is used for this research is a simulated fraud
data set created using Sparkov Data Generation [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and was published under the CC0 1.0 public
domain license [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. The simulated data set consists of 1.8 million transactions of which roughly
9000 are fraudulent, resulting in a fraud rate of 0.005. The data set contains several features,
including demographics (e.g., gender, date of birth) and characteristics of the transaction (e.g.,
transaction number, amount, time, category).
      </p>
      <p>To make it easier to visualize and observe the harmful efects of selection bias, non-fraudulent
transactions are under-sampled such that the data set has a ground-truth fraud rate of 0.1 for
both males and females. Further pre-processing consists of dropping several features (e.g.,
identifiers such as t r a n s _ n u m and the client’s f i r s t and l a s t name), merging similar categories
(e.g., g r o c e r y _ n e t and g r o c e r y _ p o s into one category g r o c e r y ), one-hot-encoding of categorical
features (e.g., the transaction c a t e g o r y ), and transforming d o b to a new feature a g e .
Simulation Design We simulate a scenario in which a machine learning model is trained on
a data set afected by selection bias. The selection bias in the initial training data is simulated
by sampling diferent levels of observed fraud of two sensitive groups, males and females.</p>
      <p>Subsequently, the model is retrained daily, based on all the labeled data that has been collected
up until that point in time. Newly labeled data originates from two sources: (1) transactions that
are labeled organically through alerts of the fraud detection model (i.e., these are the transactions
for which the model outputs the highest confidence scores), (2) transactions labeled through
explorative sampling (e.g., via active learning).</p>
      <p>Note that this setup implies several important assumptions. First of all, it is assumed that
labels are not noisy: fraud analysts are always able to accurately determine the ground-truth
3https://github.com/reneetheunissen/fraud_detection
labels. In practice, this assumption may not hold, especially when human annotators use
diferent levels of scrutiny for diferent types of alerts. Furthermore, it is assumed that each
label has equal annotation costs, resulting in a fixed amount of alerts that can be labeled each
day. Additionally, we assume no external sources of labels, such as notifications from customers.
That is, we assume that all observed fraudulent transactions are discovered through either fraud
detection system alerts or explorative sampling. Finally, we assume that the data set is not
subject to concept drift.</p>
      <p>Simulation Parameters The simulation has four main parameters that are varied across
simulation scenarios: the observed fraud rates for males and females, the alert rate, the exploration
rate, and the machine learning model class.</p>
      <p>
        • The observed fraud rate indicates the proportion of all transactions that are labeled
as fraudulent for a particular subgroup. The initial observed fraud rates for males and
females difers across simulation scenarios.
• The alert rate is defined as the proportion of incoming daily transactions that will be
labeled. In other words, this parameter determines how many transactions are added to
the training data set each day.
• The exploratory rate is the proportion of alerts that are labeled through exploratory
sampling. All other alerts are labeled organically through alerts of the fraud detection
model. The exploratory rate represents the balance between exploration of unlabeled data
and identifying fraudulent transactions.
• We train two types of machine learning models of diferent levels of complexity:
logistic regression models and random forest classifiers. We leverage the implementations
in scikit-learn [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. All models are trained using the default parameters in scikit-learn 1.2.
      </p>
      <sec id="sec-5-1">
        <title>Explorative Sampling Approaches</title>
        <p>approaches.</p>
        <p>
          We evaluate the following three explorative sampling
• Biased data. The model is trained based solely on instances labeled organically through
alerts. This approach serves as the most extreme baseline where no exploration occurs at
all. This scenario is unlikely to occur in practice, as fraudulent transactions are typically
also discovered via external data sources such as notifications from customers or (manual)
investigations by fraud analysts.
• Random sampling. The model is trained based on both alerts and explorative sampling
via random sampling. This approach serves as a baseline where some exploration is done,
but not via active learning. We hypothesize that while exploration will be able to mitigate
harmful efects of selection bias compared to the biased data approach, the approach may
not identify many fraudulent transactions due to the highly imbalanced nature of the
data set.
• Uncertainty sampling. The model is trained based on both alerts and explorative sampling
via uncertainty sampling. We hypothesize that active learning via uncertainty sampling
performs best, as it steers exploration towards underexplored areas of the feature space
and away from overexplored areas – including areas afected by selection bias.
Evaluation Metrics Selection bias can be viewed as a form of measurement bias in the target
variable: observed base rates are an imperfect proxy of true base rate [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Consequently,
metrics computed over the observed labels can be a poor measure of the true characteristics
of the machine learning model. In practice, we only have access to observed rates. In our
simulation, however, we can compute metrics over the ground-truth target variable, allowing
us to investigate the true efects of selection bias.
        </p>
        <p>In particular, we present the following statistics. The true fraud rate (TFR) indicates the
ground-truth proportion of positives. The observed fraud rate (OFR) indicates the proportion of
predicted positives out of all instances until that point in time. The rate of predicted positives
(RPP) shows the proportion of predicted positives. Additionally, we compute the following
predictive performance metrics: false positive rate (FPR), false negative rate (FNR), false discovery
rate (FDR), false omission rate (FOR), classification accuracy (ACC).</p>
        <p>Experiments The initial training data set contains 6500 instances in all scenarios, which is
created through random sampling from fraudulent and non-fraudulent transactions for males
and females according to the desired observed fraud rate. Each day, 3340 new transactions arrive.
In each scenario, the simulation runs for 25 days. We perform two experiments.
1. The Efect of Selection Bias. In this set of simulations, we explore the efects of
selection bias in the biased data baseline scenario.</p>
        <p>• Observed fraud rates. To study the efects of the magnitude of the initial selection
bias, we perform a set of simulations where the initial observed fraud rate for females
is constant (0.05) and the observed fraud rate of males in the initial training data is
varied between the range of 0 and 1.
• Alert rates. We vary the alert rate to investigate the efect of the magnitude of
selection bias over time. Across simulations, the alert rate is varied between the
values 0.01, 0.05, and 0.10.
2. The Efect of Exploration. In this experiment, we study to what extent exploration
through random sampling and uncertainty sampling can mitigate harmful efects of
selection bias. Across simulations, the exploratory rate is varied between the values 0.10,
0.25, and 0.50. In this experiment, the alert rate is set to 0.05. observed fraud rate is set to
0.3 for the overrepresented group and 0.05 for the underrepresented group.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <sec id="sec-6-1">
        <title>6.1. Experiment 1: The Efect of Selection Bias</title>
        <p>High Observed Fraud Rates Lead to False Alerts and Overscrutinization of the
Overrepresented Group Our first simulation shows the efects of selection bias in the initial
training data (Figure 1). We observe that a high observed fraud rate leads to false alerts and
overscrutinization of the overrepresented group, while a low observed fraud rates leads to
undecteded fraud. As expected, a high observed fraud rate results in an increase of the FPR, FDR,
and RPP, which leads to more transactions incorrectly being predicted as fraudulent, while a
low observed fraud rate results an RPP that is lower than the true proportion of positives (0.10),
which leads to many undetected fraudulent cases.
(a) RPP and ACC for males and (b) FDR and FOR for males and (c) FPR and FNR for males and
fefemales using the logistic re- females using the logistic re- males using the logistic
regresgression classifier. gression classifier. sion classifier.
(d) RPP and ACC for males and fe- (e) FDR and FOR for males and fe- (f) FPR and FNR for males and
females using the random forest males using the random forest males using the random forest
classifier. classifier. classifier.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Increasing the Number of Labeled Instances Reduces Inequalities Figure 2 shows the</title>
        <p>efect of diferent alert rates over time. Low alert rates increases the inequality between the
over and underrepresented groups. Moreover, we observe that these efects are amplified over
time. In particular, low alert rates do not ofer a great variety of alerts, resulting in further
overscrutinization of the overrepresented group. The divergence of the alert distribution is the
driving factor behind the development of other metrics (OFR, TFR, FPR, FNR, FDR, and FOR).
When the alert rate is high, harmful efects of selection bias are dampened. The wider variety
of alerts allows for organic correction of the high initial observed fraud rates as well as other
metrics.
(a) The alert distribution using 1% (b) The alert distribution using 5% (c) The alert distribution using
alerts. alerts. 10% alerts.
(d) The OFR and TFR using 1% (e) The OFR and TFR using 5% (f ) The OFR and TFR using 10%
alerts. alerts. alerts.</p>
        <p>We visualize the results of our simulations for the logistic regression model for random sampling
(Figure 3) and uncertainty sampling (Figure 4). The results of the random forest model are omitted due
to space constraints.
(g) The FPR and FNR using 1% (h) The FPR and FNR using 5% (i) The FPR and FNR using 10%
alerts. alerts. alerts.
(j) The FDR and FOR using 1% (k) The FDR and FOR using 5% (l) The FDR and FOR using 10%
alerts. alerts. alerts.</p>
      </sec>
      <sec id="sec-6-3">
        <title>6.2. Experiment 2: The Efect of Exploration</title>
      </sec>
      <sec id="sec-6-4">
        <title>Uncertainty Sampling Outperforms Random Sampling Figures 3 and Figure 4 show</title>
        <p>the results of our simulations for random sampling and uncertainty sampling, respectively.
The results of the random forest model are omitted due to space constraints. While a low
exploratory rate of random sampling improves the distribution of alerts between the over and
underrepresented group, the improvement is unable to improve predictive performance metrics.
Indeed, without improving the observed fraud rate, selection bias persists and the RPP will
continue to decrease while the FNR will continue to increase.</p>
        <p>Uncertainty sampling, on the other hand, is able to counter negative efects of selection bias
and decrease disparities in predictive performance between groups. Uncertainty sampling leads
to a better alert distribution with balanced alerts (both fraudulent and non-fraudulent) of both
the over- and underrepresented group. In this way, the observed fraud rate is corrected over
time through the additional exploratory data.</p>
        <p>As can be expected, higher exploratory rates (up to 0.50) result in higher levels of mitigation
of harmful efects: a better alert distribution, decreased disparities in the FPR, FNR, FDR, and
FOR, and increase in accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this work, we have studied the problem of selection bias related to sensitive features and
analyzed the efectiveness of active learning to mitigate the ensuing harmful efects.</p>
      <p>The results of our simulations confirm that in the absence of interventions, selection bias
is reinforced over time, resulting in an increase in scrutinization as well as an increase in
the number of false positives for overrepresented groups. Moreover, our experiments show
preliminary evidence that uncertainty sampling can mitigate disparities between groups caused
by selection bias and, in the studied setting, outperforms exploration via random sampling.</p>
      <p>Our results also imply that limited resources are an important bottleneck in the organic
correction of disparities in observed fraud rates, as increasing the number of labeled instances
reduces inequalities between sensitive groups even in absence of interventions. These results
suggest that the problem of mitigating selection bias can be seen as analogous to the
wellknown exploration-exploitation trade-of. That is, under limited resources, decision-makers are
required to balance the mitigation of selection bias through exploration and the exploitation of
existing models for identifying positives in production.</p>
      <p>The preliminary results presented in this paper open up many directions for future work.
First of all, we see several ways in which our simulation study can be extended, including the
replication of our experiments on more (real-world) data sets, uncertainty quantification via
repeated experiments, and relaxation of the simplifying assumptions in our simulation, such
as the lack of noisy labels, equal annotation cost, absence of external sources of labels, and
a lack of concept drift. In particular, the efects of using online or adaptive learners could be
further explored. Additionally, future work could focus on evaluating alternative active learning
sampling techniques. In particular, an important limitation of uncertainty sampling is that it
relies solely on the model’s confidence score as a proxy for uncertainty. This could result in
repeated selection of instance types that are inherently dificult to predict based on the available
features, even when an additional instance is unlikely to improve the predictive performance of
the model. Finally, we envision future work that tackles the development of active learning
techniques that are specifically designed to tackle selection bias related to sensitive group
membership. In particular, more research is needed to identify a suitable trade-of between
exploration through active learning and exploitation of organically generated alerts across
scenarios.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Nederlandse</given-names>
            <surname>Autoriteit</surname>
          </string-name>
          <string-name>
            <surname>Persoonsgegevens</surname>
          </string-name>
          , Onderzoeksrapport Belastingdienst/Toeslagen - De verwerking van de nationaliteit van aanvragers van kinderopvangtoeslag,
          <year>2021</year>
          . URL: https://www.autoriteitpersoonsgegevens.nl/sites/default/files/atoms/files/ onderzoek_belastingdienst_kinderopvangtoeslag.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Lum</surname>
          </string-name>
          , W. Isaac, To predict and serve?,
          <source>Significance</source>
          <volume>13</volume>
          (
          <year>2016</year>
          )
          <fpage>14</fpage>
          -
          <lpage>19</lpage>
          . doi: h t t p s : / / d o i .
          <source>o r g / 1 0 . 1 1 1 1 / j . 1 7</source>
          <volume>4 0 - 9 7 1 3 . 2 0 1 6 . 0 0 9 6 0</volume>
          . x .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Daitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Turjeman</surname>
          </string-name>
          , I. Poran,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Ayalon-Dangur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nashashibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yahav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Paul</surname>
          </string-name>
          , L. Leibovici,
          <article-title>Underrepresentation of women in randomized controlled trials: a systematic review and meta-analysis</article-title>
          ,
          <source>Trials</source>
          <volume>23</volume>
          (
          <year>2022</year>
          ).
          <source>doi:1 0 . 1 1 8 6 / s 1 3</source>
          <volume>0 6 3 - 0 2 2 - 0 7 0 0 4 - 2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dal Pozzolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Boracchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Caelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Alippi</surname>
          </string-name>
          , G. Bontempi,
          <article-title>Credit card fraud detection: A realistic modeling and a novel learning strategy</article-title>
          ,
          <source>IEEE Transactions on Neural Networks and Learning Systems</source>
          <volume>29</volume>
          (
          <year>2018</year>
          )
          <fpage>3784</fpage>
          -
          <lpage>3797</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>0</volume>
          <fpage>9</fpage>
          <string-name>
            <surname>/ T N N L S</surname>
          </string-name>
          .
          <volume>2 0 1 7 . 2 7 3 6 6 4 3 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Nederlandse</given-names>
            <surname>Autoriteit</surname>
          </string-name>
          <string-name>
            <surname>Persoonsgegevens</surname>
          </string-name>
          , Besluit tot boeteoplegging,
          <year>2021</year>
          . URL: https://autoriteitpersoonsgegevens.nl/sites/default/files/atoms/files/boetebesluit_ belastingdienst.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Gleich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ramani</surname>
          </string-name>
          ,
          <article-title>Deconvolving feedback loops in recommender systems</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>29</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ensign</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Friedler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Neville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Scheidegger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Venkatasubramanian</surname>
          </string-name>
          ,
          <article-title>Runaway feedback loops in predictive policing</article-title>
          , in: S. A.
          <string-name>
            <surname>Friedler</surname>
          </string-name>
          , C. Wilson (Eds.),
          <source>Proceedings of the 1st Conference on Fairness, Accountability and Transparency</source>
          , volume
          <volume>81</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>171</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kallus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Residual unfairness in fair machine learning from prejudiced data</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>2439</fpage>
          -
          <lpage>2448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fogliato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chouldechova</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>G'Sell, Fairness evaluation in presence of biased noisy labels</article-title>
          ,
          <source>in: International Conference on Artificial Intelligence and Statistics</source>
          , PMLR,
          <year>2020</year>
          , pp.
          <fpage>2325</fpage>
          -
          <lpage>2336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>I.</given-names>
            <surname>Pastaltzidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Dimitriou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Quezada-Tavarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aidinlis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Marquenie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gurzawska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tzovaras</surname>
          </string-name>
          ,
          <article-title>Data augmentation for fairness-aware machine learning: Preventing algorithmic bias in law enforcement systems</article-title>
          , in: 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT '22,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2022</year>
          , p.
          <fpage>2302</fpage>
          -
          <lpage>2314</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 5 3 1 1 4 6 . 3 5 3 4 6 4 4 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Weerts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Royakkers</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Pechenizkiy, Does the end justify the means? On the moral justification of fairness-aware machine learning</article-title>
          ,
          <source>arXiv preprint arXiv:2202.08536</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Weerts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xenidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tarissan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. P.</given-names>
            <surname>Olsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pechenizkiy</surname>
          </string-name>
          ,
          <article-title>Algorithmic unfairness through the lens of EU non-discrimination law</article-title>
          ,
          <source>in: 2023 ACM Conference on Fairness, Accountability, and Transparency</source>
          , ACM,
          <year>2023</year>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 5 9 3 0 1 3 . 3 5 9 4 0 4 4 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Settles</surname>
          </string-name>
          ,
          <article-title>Active Learning Literature Survey</article-title>
          ,
          <source>Computer Sciences Technical Report 1648</source>
          , University of Wisconsin-Madison,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>D. D. Lewis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Catlett</surname>
          </string-name>
          ,
          <article-title>Heterogeneous uncertainty sampling for supervised learning</article-title>
          ,
          <source>in: Machine learning proceedings 1994, Elsevier</source>
          ,
          <year>1994</year>
          , pp.
          <fpage>148</fpage>
          -
          <lpage>156</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Branchaud-Charron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atighehchian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Abuhamad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lacoste</surname>
          </string-name>
          ,
          <article-title>Can active learning preemptively mitigate fairness issues?</article-title>
          ,
          <source>arXiv preprint arXiv:2104.06879</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Richards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Starr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Brink</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bloom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Butler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>James</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Long</surname>
          </string-name>
          , a. Rice,
          <article-title>Active learning to overcome sample selection bias: Application to photometric variable star classification</article-title>
          ,
          <source>The Astrophysical Journal</source>
          <volume>744</volume>
          (
          <year>2011</year>
          )
          <article-title>192. doi: 1 0 . 1 0 8 8 / 0 0 0 4 - 6 3 7 X / 7 4 4 / 2 / 1 9 2</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Kamiran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Calders</surname>
          </string-name>
          ,
          <article-title>Data preprocessing techniques for classification without discrimination</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>33</volume>
          (
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          .
          <source>doi:1 0 . 1 0 0 7 / s 1 0</source>
          <volume>1 1 5 - 0 1 1 - 0 4 6 3 - 8</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>M. B. Zafar</surname>
            , I. Valera,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Rogriguez</surname>
            ,
            <given-names>K. P.</given-names>
          </string-name>
          <string-name>
            <surname>Gummadi</surname>
          </string-name>
          ,
          <article-title>Fairness Constraints: Mechanisms for Fair Classification</article-title>
          , in: A.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          Zhu (Eds.),
          <source>Proceedings of the 20th International Conference on Artificial Intelligence and Statistics</source>
          , volume
          <volume>54</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>962</fpage>
          -
          <lpage>970</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Srebro</surname>
          </string-name>
          ,
          <article-title>Equality of opportunity in supervised learning</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>29</volume>
          (
          <year>2016</year>
          )
          <fpage>3315</fpage>
          -
          <lpage>3323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>F.</given-names>
            <surname>Kamiran</surname>
          </string-name>
          , T. Calders,
          <article-title>Classification with no discrimination by preferential sampling</article-title>
          ,
          <source>in: Proc. 19th Machine Learning Conf. Belgium and The Netherlands</source>
          , volume
          <volume>1</volume>
          ,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Anahideh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Asudeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thirumuruganathan</surname>
          </string-name>
          ,
          <article-title>Fair active learning</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>199</volume>
          (
          <year>2022</year>
          )
          <article-title>116981</article-title>
          . doi:h t t p s : / / d o i .
          <source>o r g / 1 0 . 1 0</source>
          <volume>1 6</volume>
          / j . e
          <source>s w a . 2 0</source>
          <volume>2 2 . 1 1 6 9 8 1 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharaf</surname>
          </string-name>
          , H.
          <string-name>
            <surname>Daume</surname>
            <given-names>III</given-names>
          </string-name>
          , R. Ni,
          <article-title>Promoting fairness in learned models by learning to active learn under parity constraints</article-title>
          , in: 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT '22,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2022</year>
          , p.
          <fpage>2149</fpage>
          -
          <lpage>2156</lpage>
          . URL: https://doi.org/10.1145/3531146.3534632.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 5 3 1 1 4 6 . 3 5 3 4 6 3 2 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>A. D. Selbst</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          <string-name>
            <surname>Friedler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Venkatasubramanian</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Vertesi</surname>
          </string-name>
          ,
          <article-title>Fairness and abstraction in sociotechnical systems</article-title>
          ,
          <source>in: Proceedings of the Conference on Fairness, Accountability, and Transparency</source>
          , FAT* '
          <volume>19</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>59</fpage>
          -
          <lpage>68</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 2 8 7 5 6 0 . 3 2 8 7 5 9 8 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Guerdan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Coston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Holstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Counterfactual prediction under outcome measurement error</article-title>
          ,
          <source>in: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency</source>
          , FAccT '23,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , p.
          <fpage>1584</fpage>
          -
          <lpage>1598</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>H.</given-names>
            <surname>Brandon</surname>
          </string-name>
          , Sparkov data generation,
          <year>2022</year>
          . URL: https://github.com/namebrandon/ Sparkov_Data_Generation.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>K.</given-names>
            <surname>Shenoy</surname>
          </string-name>
          ,
          <article-title>Credit card transactions fraud detection dataset</article-title>
          ,
          <year>2020</year>
          . URL: https://www.kaggle. com/datasets/kartik2112/fraud-detection?select=fraudTrain.csv.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , E. Duchesnay,
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>A. Z.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <article-title>Measurement and fairness</article-title>
          ,
          <source>in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</source>
          , FAccT '21,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2021</year>
          , p.
          <fpage>375</fpage>
          -
          <lpage>385</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 4 4 2 1 8 8 . 3 4 4 5 9 0 1 .</volume>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>