<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bias Amplification Chains in ML-based Systems with an Application to Credit Scoring⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro G. Buda</string-name>
          <email>alessandro.buda@iusspavia.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Greta Coraglia</string-name>
          <email>greta.coraglia@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco A. Genco</string-name>
          <email>francesco.genco@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Manganini</string-name>
          <email>chiara.manganini@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Primiero</string-name>
          <email>giuseppe.primiero@unimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>3rd Workshop on Bias, Ethical AI, Explainability and the Role of Logic and Logic Programming (BEWARE24)</institution>
          ,
          <addr-line>co-located with AIxIA 2024</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LUCI Lab, Department of Philosophy, University of Milan</institution>
          ,
          <addr-line>via Festa del Perdono 7, 20122 Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Machine Learning (ML) systems, whether predictive or generative, not only reproduce biases and stereotypes but, even more worryingly, amplify them. Strategies for bias detection and mitigation typically focus on either ex post or ex ante approaches, but are always limited to two steps analyses. In this paper, we introduce the notion of Bias Amplification Chain (BAC) as a series of steps in which bias may be amplified during the design, development and deployment phases of trained models. We provide an application to such notion in the credit scoring setting and a quantitative analysis through the BRIO tool.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ML Fairness</kwd>
        <kwd>Bias Amplification</kwd>
        <kwd>Responsible AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In recent years, research in the field of artificial intelligence has seen its impact grow in the lives of
many people. In particular, algorithmic fairness, and its negative counterpart, algorithmic bias, have
become hot topics: many experts, as well as ordinary people, have realized that systems based on
Machine Learning (ML), whether predictive or generative, not only reproduce biases and stereotypes
but, even more worryingly, amplify and reinforce them. This phenomenon, known as bias amplification,
urgently requires addressing to start relying on automatic systems.</p>
      <p>
        In academic research, risks of bias have been outlined for a decade now and solutions are being
sought, following two main approaches to algorithmic fairness:
• an ex post approach, in which fairness metrics are defined to identify and mitigate the presence
of bias [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], and
• an ex ante approach, which sees algorithmic bias “as evidence of the underlying social and technical
conditions that (re)produce it” [3, p.2], focusing on eXplainable Artificial Intelligence (XAI).
      </p>
      <p>
        Both of these approaches, despite their great diferences, are characterized by identifying an ideal,
abstract, probabilistic model to which an empirical, non-deterministic result must come as close as
possible. In the former approach, the evaluation of distance between the two is performed without
assuming any knowledge of the underlying model. In the latter approach, the analysis seeks to identify
a transparent underlying structure which minimizes distance from a given justification. Standard
theoretical computer science and philosophy of computing terminology [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ] have called these two
terms of comparison Levels of Abstractions (LoAs); the complex structure of ML systems has already
required a reconsideration of such terminology in terms of a User Level approach [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which in turn
motivates a future redesign of the whole ontology of such systems.
      </p>
      <p>
        However, when dealing with complex AI systems, and especially in the case of generative AI, there
are several levels to consider, and at each of these levels bias amplification phenomena can occur.
We will refer to the bias amplification phenomenon in a system with more than one level as Bias
Amplification Chains (BAC). Consider the case of text-to-image generation: one should first take into
account that “[f]airness [...] is not a purely technical construct, having social, political, philosophical
and legal facets" [7, p.1], and therefore consider how bias is reproduced and pre-amplified from society
to training datasets, via image captioning [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]; secondly, one should take into consideration how bias
present in training datasets is amplified in ML models due to model accuracy, model capacity, model
overconfidence, amount of training data, and how such bias can vary throughout the training process
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]; then, one should focus on bias amplification occurring in generated images, primarily caused by
discrepancies between training captions and text prompts [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]; lastly, one should consider the surprising
fact that bias mitigation operations may themselves unexpectedly boost bias [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This behavior can also
be seen in action in the predictive context, where a ML model is trained to predict properties that are
relevant for taking decisions, especially in critical fields like finance, healthcare, or social justice.
      </p>
      <p>
        A perspective similar to ours on the way bias propagates and amplifies throughout the life cycle of
machine learning was given by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], although the authors did not ofer any detailed methodology to
quantify the divergence between the bias in input and the one in output (respectively, the first and
the last links of the BAC). On the other hand, our point of view difers from that of [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and others,
who argue for a directional approach, meaning one where causality of the amplification is taken into
account: our analysis is causal in the sense that one can use the methods presented here to further
investigate conditioning, but it should be noted that a compositional approach – meaning one allowing
us to propagate that information through the various levels presented in Section 6 – is still lacking.
      </p>
      <p>
        In the present work we describe and exemplify the case of bias amplification in the sense of [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] in a
credit scoring setting, conducting our analysis by means of a recent bias detection tool called BRIO1,
which is a model-agnostic tool designed to assess the bias and related risk of unfairness of prediction
tasks on tabular data. For a technical presentation of its basic features and a discussion on validation,
we refer to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The most recent developments of BRIO in the direction of assessing bias and risk can be
instead found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This proprietary software has been developed on the basis of a family of logics
for trustworthiness assessment [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16">13, 14, 15, 16</xref>
        ], and has already been used to assess fairness in credit
scoring models [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>To show our methodology, we examine a simple amplification chain consisting of just three links.
The first link identifies the amplification of bias occurring in sampling the training dataset from the
real population, thus identifying the input bias of our ML pipeline. The second link corresponds to the
divergence between the distribution on which the model is trained and the one to which it is applied
(test set). Lastly, the third link identifies the amplification of bias outputted by our model, that is, how
the bias is amplified from the dataset used to perform the analysis to the predictions produced by the
model. Our goal is achieved through the following functionalities:
• FreqVsRef : a functionality of BRIO which considers the first and the second link of the chain,
by investigating how the dataset population distribution (Freq) amplifies bias on a property of
interest as observed in the real population (Ref);
• FreqVsFreq: a functionality of BRIO which considers the second and third links of the chain, by
investigating how the distribution divergence between the sensitive groups increases or decreases
in the predicted outcome, with respect to the true distribution;
• an in-depth final analysis to inspect the entire chain, quantifying the contribution of each link to
amplification, and identifying which of them most urgently requires mitigation interventions.</p>
      <p>
        In practice, we conduct our analysis in the paradigmatic context of credit scoring. We use the UCI
German Credit Dataset [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], despite several reported limitations [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ], because our purpose is not to
1The open source code is available at https://github.com/DLBD-Department/BRIO_x_Alkemy.
provide a comprehensive data analysis, but rather to show a novel methodology. The full outputs of the
tool will remain available in the open repository of this project.
      </p>
      <p>The paper is structured as follows. First, we introduce BRIO’s main functionalities in Section 2. These
will be used in subsequent Sections 3, 4, and 5 – each of which focuses on a specific link of the BAC – to
compute the corresponding bias amplification. Finally, some final insights on the possible refinements
of the notion of BAC will be given in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The BRIO Method</title>
      <sec id="sec-2-1">
        <title>BRIO’s bias detection method takes in input</title>
        <p>• the predictions of a ML model encoded as a set of datapoints with relative features,
• a set of parameters, including the designation of one or more sensitive features, also called protected
attributes,
• and a distribution of reference, which might be automatically computed on the input dataset of
the model, or externally provided,</p>
        <p>On this input, it returns an evaluation of the possibility that the model under consideration is unfair
with respect to the designated features when comparing the predictions and the reference distribution.
To do so, the system allows conducting two kinds of analyses, consisting in:
1. FreqVsRef : comparing the behaviour of the AI system against a desirable one;
2. FreqVsFreq: comparing the behaviour of the AI system with respect to a sensitive group 1 against
another sensitive group 2 related to the same feature  , with  = 1, 2.</p>
        <p>If the second analysis alerts of a possibly biased behaviour, one can conduct a subsequent check on
some (or all the) subclasses of the considered sensitive classes. This second check is meant to verify
whether the bias encountered at the level of the classes can be explained away by non-sensitive features
of the individuals that it is morally acceptable to use for taking decisions.</p>
        <p>One can use the result of analyses 1 and 2 and, keeping track of the (sub)classes where they fail,
compute:
3. Hazard: an aggregate measure of the number of datapoints where a test of type 1 or 2 fails, of
how far it is from not failing on some datapoints (when so), and of how dificult was it for it to
fail.</p>
        <p>Age
0-24 years old
25-39 y.o.
40-59 y.o.</p>
        <p>Over 60 y.o.</p>
        <sec id="sec-2-1-1">
          <title>Gender</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>Female</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>Male</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>Marital status</title>
        </sec>
        <sec id="sec-2-1-5">
          <title>Single</title>
        </sec>
        <sec id="sec-2-1-6">
          <title>Married/separated/divorced/widowed</title>
        </sec>
        <sec id="sec-2-1-7">
          <title>Total population 45.48% 54.52%</title>
        </sec>
        <sec id="sec-2-1-8">
          <title>Training set 54.00% 46.00%</title>
          <p>Of this latter measure, fully detailed in [2, S5], two options are provided, one focusing on group fairness
and one on individual fairness. In the present work, we only consider the one involving group fairness2.</p>
          <p>Finally, it should be noted that, since the tool is model-agnostic, the same analyses 1, 2, and 3 described
above can be applied to study the distribution of the ground truth (as it was the prediction of a perfect
classifier). This feature will be used in Section 5.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Bias in Sampling the Training Set</title>
      <p>In this section, we apply the FreqVsRef bias detection module of the BRIO tool in order to evaluate
whether the process of labelling and sampling from the general population3 to produce the training set
is biased in the sense that it inaccurately represents the real distribution Bias is in fact mathematically
defined as a systematic deviation from the true estimation. Although in the ML literature this term is
typically used in relation to the decisions of a model, when it comes to the training set, we are interested
in checking for bias contained in the features, rather than in the labels (or decisions). In fact, as Figure 1
illustrates, the reference distribution against which we want to compare the training set contains no
labels.</p>
      <p>
        For our experiments, we use the UCI German Credit Dataset [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which ofers a comprehensive
compilation of attributes relevant for creditworthiness evaluation. The dataset comprises 1, 000
instances, each characterised by a set of 20 input variables and an associated binary label representing the
occurrence or not of the default event. Our reference distributions are provided by statistics concerning
the German society in 2024, see Table 1. It should be noted that this will eventually bring us to compare
data from 2024 (our reference distribution) and data from the 1970s4 (our training set based on UCI
German Credit). While this is unfortunate, as obviously the two distribution difer in many sensible
ways, we could not gather high quality data about general demographics in the desired period of time.
Since the present work is expository of a general methodology, we shelf this issue as accidental.
2This is merely a choice in perspective, as we aim to describe possible discrimination of groups of people, instead of single
individuals. From a technical point of view, BRIO allows for both.
3References for general demographic data come from
• https://www.statista.com/statistics/454349/population-by-age-group-germany/ (age)
• https://www.statista.com/statistics/454338/population-by-gender-germany/ (gender)
• https://www.destatis.de/EN/Themes/Society-Environment/Population/Current-Population/Tables/
population-by-marital-status.html (marital status)
as of September 2024.
4One of the many problems with this dataset, is that it is usually attributed to the 1990s, while in reality its data was collected
between 1973 and 1975 [18, §3.1] – when there were two Germanies!
      </p>
      <sec id="sec-3-1">
        <title>Result of the test violation violation violation</title>
        <p>The first step of the analysis consists in running the FreqVsRef bias detection module of the BRIO tool
in order to evaluate whether the distribution of individuals in the diferent classes of the training set
diverges too much from the distribution indicated by the statistics about German society. The method
returns the respective Kullback-Leibler divergences5 [20] with Laplacian smoothing, and compares it
with the BRIO automated threshold set to high sensitivity, see Table 2.</p>
        <p>The results show that the distributions of both gender and age in the training set do not match the
corresponding distribution in real population: when further looking at the training set, in fact, one
ifnds that males are extremely over-represented, and that people in age groups 40-59 and over 60 are
greatly under-represented. The respective distributions of marital statuses are almost swapped.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Bias in the Test Set</title>
      <p>The sampling to produce a test set for the prediction task is the next link of our BAC. By test set we mean
the set of unseen data used to evaluate how the final model performs on unseen data, as opposed to the
validation set which can be used to fine-tune the model itself. A major assumption of the ML approach
is that the training and the test data share the same statistical distribution, being Independent and
Identically Distributed (i.i.d.). However, in many practical applications, the i.i.d. assumption fails due to
unforeseen distributional changes occurring once the ML system is deployed and used on new inputs
(test set). This becomes relevant for our discussion because, in addition to a possible degradation of
performance, the failure of i.i.d. might also amplify bias. We here distinguish two diferent perspectives
from which the bias amplification occurring at the test set level can be thought of.</p>
      <p>In a first sense, sampling unseen inputs for the test set may be amplifying bias in the properties of
interest as they occur in the general population. In this case, an evaluation of this bias can be computed
in the same vein as done for bias in the sampling of the training set.</p>
      <p>In a second sense, the test set may amplify or compensate bias with respect to the training set. This
is evidence of the fact that the sampling process of the test set was diferent from that producing the
training set. Put diferently, a model trained on a given distribution is then used to predict on a dataset
which sensibly difers for some property of interest. This relates to the Out-of-Distribution problem,
which occurs when the ML system is used on a distribution that significantly diverges from the one
learnt during the training phase [21]. Evaluating such bias diference can be done through the BRIO
tool in the same vein as done in the next phase of our analysis (Section 5).</p>
      <p>Note that:
a. assuming a neutral bias amplification in sampling the training set, the bias input of the chain at
the next stage will be ascribed entirely to the first sense of bias in the test set;
b. assuming a neutral bias diference in the second sense for the test set, the bias input of the chain
at the next stage will be ascribed entirely to the second sense of bias in the test set.
5The Kullback–Leibler divergence KL is mathematically expressed by the following equation:</p>
      <p>KL( ‖ ) = ∑∈︁  () · log( (()) )
and quantifies the discrepancy of a considered probability distribution  with respect to a reference probability distribution
 .</p>
    </sec>
    <sec id="sec-5">
      <title>5. Bias in Model Output</title>
      <p>To determine the extent to which the model’s predictions amplify bias relative to the actual distribution
of the sensitive groups of interest, we run the bias detection module of the BRIO tool on both the
model’s predictions and the ground truth and compare the former result against the latter. For the
present example, we run experiments for two diferent sensitive features: age and gender. Our predictive
model is obtained through a standard credit scorecard modelling process based on the methodology of
OptBinning6.</p>
      <p>The bias detection module returns – along with a list of fairness violations alerts – a hazard value
for each dataset and each sensitive feature that indicates how dangerous the violation is. This is the
value that we will employ to compute the bias amplification between the ground truth and the result of
passing it through the model.</p>
      <p>The bias detection tests have been conducted with the following arguments.</p>
      <p>• Target feature: ground truth labels (Table 3), and model predictions (Table 4).
• Sensitive features:
– gender (male and female);
– age (four age groups: 0-24, 25-39, 40-59, over 60);
– marital status (single and married/separated/divorced/widowed).
• Divergence used to measure the discrepancy between frequencies: Jensen–Shannon [22].
• Threshold: automatically computed with sensitivity set to high.
• Function to aggregate divergences between pairs of sensitive classes (only for the test on age):
arithmetical mean.
• Conditioning variables for double-checks on subclasses:
– Attribute 3, “credit history” (no credits taken / all credits paid back duly, all credits at this
bank paid back duly, existing credits paid back duly till now, delay in paying of in the past,
critical account / other credits existing but not at this bank);
– Attribute 10, “other debtors / guarantors” (none, co-applicant, guarantor).</p>
      <p>Our results are collected in Table 3 and in Table 4. We display hazard values, in particular cumulative
hazard is the sum of all estimated hazard values including without conditioning. The aim is to assess
whether looking at subclasses increases the diference in behaviour or not. Recall that a full list of
violations with their respective estimated hazards is available on Git7.</p>
      <p>First, it should be noted that the hazard measure for gender and marital status is quite similar – it is
not actually identical, see the full report in the Git repo, since it difers from the 15ℎ decimal point on.
This is partly explained by the fact that all women in our dataset are married.8</p>
      <p>Now we move on to the significance of our analysis. The most notable diference in comparing
our analysis on the ground truth with that on the model predictions occurs with gender: while the
ground truth appears to be non-biased with respect to gender,9 this is not the case for the model which</p>
      <sec id="sec-5-1">
        <title>6http://gnpalencia.org/optbinning/scorecard.html</title>
        <p>
          7https://github.com/DLBD-Department/BRIO_x_Alkemy/tree/main/BEWARE_2024
8This is perhaps another argument in favour of the thesis of [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
9Notice that having an hazard value be equal to 0 does not mean that the ground truth is perfectly balanced, but that it is
unbalanced within acceptable limits according to the selected threshold. More on this is detailed in [2, §5].
introduces “new” bias even while conditioning with possible predictors of good credit behaviour. While
less strinkingly so, the model seems to introduce new bias with respect to age as well. In order to
compute by how much this is the case, we represent bias amplification by a percentage change in hazard:
amplification =
(predictions_hazard − ground_truth_hazard) · 100
        </p>
        <p>ground_truth_hazard
whose results are collected in Table 5.</p>
        <p>Of course, as for the hazard values related to gender and marital status, we would need to divide
0.1673 by 0, so we cannot compute an actual percentage. When this is the case, in general, we have a
situation in which the model introduces a bias which is not at all present in the ground truth. While we
cannot compute a percentage value, we can still evaluate the seriousness of the situation by considering
the hazard values computed by the BRIO tool relatively to the predictions of the model. For instance,
as far as gender and marital status are concerned, the bias introduced is comparable: the tool yielded
for both sensitive features an hazard value (without conditioning) of 0.0695 and a cumulative hazard
value of 0.1673. As already mentioned, this is probably due to the extreme correlation between the
gender value associated to women and the marital status value associated to married people. We leave
as future work the task of defining a mathematical measure of bias amplification that enables us to
homogeneously treat all these cases.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Bias amplification describes the process by which bias in the data propagates throughout the ML
pipeline, ultimately leading to unfair outcomes. By building on this notion, in this paper, we introduced
the concept of the Bias Amplification Chain (BAC), illustrated three links thereof, and used the BRIO
framework to quantify each of these.</p>
      <p>An interesting avenue for future research involves exploring more complex configurations of the
Bias Amplification Chain itself. In particular, refining our BAC with additional links to the chain
could provide a more fine-grained and realistic analysis of bias amplification. For instance, we want
to investigate how bias is amplified during the phase of data curation, when in order to optimize the
training certain variables are discarded (feature selection), and synthetic features are automatically
crafted by the system from the initial predictors (feature engineering).</p>
      <p>More generally, the notion of BAC provides a model to quantitatively reason about where and when
unfairness is produced along the ML life cycle. Breaking down the phenomenon of ML unfairness in
diferent links of a chain representing the ML pipeline can potentially yield valuable insights on the
efectiveness of certain mitigation methods (e.g., increasing the data quality, increasing the sample
size, learning fair representations, etc.). Finally, a major open question remains how to meaningfully
aggregate the preliminary results here obtained in order to gain a holistic insight on the phenomenon
of ML unfairness.
2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, Association
for Computing Machinery, New York, NY, USA, 2024, p. 642–659. URL: https://doi.org/10.1145/
3630106.3658931. doi:10.1145/3630106.3658931.
[20] S. Kullback, R. A. Leibler, On Information and Suficiency, The Annals of Mathematical Statistics 22
(1951) 79 – 86. URL: https://doi.org/10.1214/aoms/1177729694. doi:10.1214/aoms/1177729694.
[21] J. Liu, Z. Shen, Y. He, X. Zhang, R. Xu, H. Yu, P. Cui, Towards out-of-distribution generalization: A
survey, 2023. URL: https://arxiv.org/abs/2108.13624. arXiv:2108.13624.
[22] J. Lin, Divergence measures based on the shannon entropy, IEEE Transactions on Information
Theory 37 (1991) 145–151. doi:10.1109/18.61115.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Coraglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>A. D'Asaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Genco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Giannuzzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Posillipo</surname>
          </string-name>
          , G. Primiero,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quaggio</surname>
          </string-name>
          ,
          <article-title>Brioxalkemy: a bias detecting tool</article-title>
          , in: BEWARE@AI*IA,
          <year>2023</year>
          . URL: https://api.semanticscholar. org/CorpusID:267200510.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Coraglia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Genco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Piantadosi</surname>
          </string-name>
          , E. Bagli,
          <string-name>
            <given-names>P.</given-names>
            <surname>Giufrida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Posillipo</surname>
          </string-name>
          , G. Primiero,
          <article-title>Evaluating ai fairness in credit scoring with the brio tool</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2406</volume>
          .
          <fpage>03292</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ziosi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          ,
          <article-title>A genealogical approach to algorithmic bias</article-title>
          ,
          <source>Minds and Machines</source>
          <volume>34</volume>
          (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11023-024-09672-2.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Floridi</surname>
          </string-name>
          ,
          <article-title>The method of levels of abstraction, Minds Mach</article-title>
          .
          <volume>18</volume>
          (
          <year>2008</year>
          )
          <fpage>303</fpage>
          -
          <lpage>329</lpage>
          . URL: https: //doi.org/10.1007/s11023-008-9113-7. doi:
          <volume>10</volume>
          .1007/s11023-008-9113-7.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Primiero</surname>
          </string-name>
          ,
          <article-title>Information in the philosophy of computer science</article-title>
          , Routledge,
          <year>2016</year>
          . URL: https:// www.routledgehandbooks.com/doi/10.4324/9781315757544.ch10. doi:
          <volume>10</volume>
          .4324/9781315757544. ch10.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Buda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Primiero</surname>
          </string-name>
          ,
          <article-title>A pragmatic theory of computational artefacts</article-title>
          ,
          <source>Minds Mach</source>
          .
          <volume>34</volume>
          (
          <year>2024</year>
          )
          <fpage>139</fpage>
          -
          <lpage>170</lpage>
          . URL: https://doi.org/10.1007/s11023-023-09650-0. doi:
          <volume>10</volume>
          .1007/S11023-023-09650-0.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Foulds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Keya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <article-title>An intersectional definition of fairness, 2019</article-title>
          . URL: https: //arxiv.org/abs/
          <year>1807</year>
          .08362. arXiv:
          <year>1807</year>
          .08362.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hirota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nakashima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <article-title>Quantifying societal bias amplification in image captioning</article-title>
          ,
          <source>in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>13440</fpage>
          -
          <lpage>13449</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR52688.
          <year>2022</year>
          .
          <volume>01309</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          , L. van der Maaten, L. Gustafson,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Adcock</surname>
          </string-name>
          ,
          <article-title>A systematic study of bias amplification</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv.org/abs/2201.11706. arXiv:
          <volume>2201</volume>
          .
          <fpage>11706</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Seshadri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Elazar</surname>
          </string-name>
          ,
          <article-title>The bias amplification paradox in text-to-image generation</article-title>
          ,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2308.00755. arXiv:
          <volume>2308</volume>
          .
          <fpage>00755</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Suresh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. V.</given-names>
            <surname>Guttag</surname>
          </string-name>
          ,
          <article-title>A framework for understanding sources of harm throughout the machine learning life cycle, Equity and Access in Algorithms</article-title>
          , Mechanisms, and
          <string-name>
            <surname>Optimization</surname>
          </string-name>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Russakovsky</surname>
          </string-name>
          ,
          <article-title>Directional bias amplification</article-title>
          , in: M.
          <string-name>
            <surname>Meila</surname>
          </string-name>
          , T. Zhang (Eds.),
          <source>Proceedings of the 38th International Conference on Machine Learning</source>
          , volume
          <volume>139</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10882</fpage>
          -
          <lpage>10893</lpage>
          . URL: https://proceedings.mlr.press/v139/wang21t. html.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>A. D'Asaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Primiero</surname>
          </string-name>
          ,
          <article-title>Probabilistic typed natural deduction for trustworthy computations</article-title>
          ,
          <source>in: TRUST@AAMAS</source>
          ,
          <year>2021</year>
          . URL: https://api.semanticscholar.org/CorpusID:245423393.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Genco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Primiero</surname>
          </string-name>
          ,
          <article-title>A typed lambda-calculus for establishing trust in probabilistic programs</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>00958</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>A. D'Asaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Genco</surname>
          </string-name>
          , G. Primiero,
          <article-title>Checking trustworthiness of probabilistic computations in a typed natural deduction system</article-title>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2206</volume>
          .
          <fpage>12934</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kubyshkina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Primiero</surname>
          </string-name>
          ,
          <article-title>A possible worlds semantics for trustworthy non-deterministic computations</article-title>
          ,
          <source>International Journal of Approximate Reasoning</source>
          <volume>172</volume>
          (
          <year>2024</year>
          )
          <article-title>109212</article-title>
          . URL: https://www. sciencedirect.com/science/article/pii/S0888613X24000999. doi:https://doi.org/10.1016/j. ijar.
          <year>2024</year>
          .
          <volume>109212</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <article-title>Statlog (German Credit Data)</article-title>
          ,
          <source>UCI Machine Learning</source>
          Repository DOI: https: //doi.org/10.24432/C5NC77 (
          <year>1994</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>U.</given-names>
            <surname>Grömping</surname>
          </string-name>
          ,
          <article-title>South German Credit Data: Correcting a Widely Used Data Set</article-title>
          .,
          <string-name>
            <surname>Department</surname>
            <given-names>II</given-names>
          </string-name>
          , Beuth University of Applied Sciences, Berlin,
          <year>2019</year>
          . URL: https://www1.beuth-hochschule.de/FB_ II/reports/Report-2019-004.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Simson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <article-title>Lazy data practices harm fairness research</article-title>
          , in: Proceedings of the
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>