<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Integrating SQuaRE data quality model with ISO 31000 risk management to measure and mitigate software bias</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Simonetta</string-name>
          <email>alessandro.simonetta@gmail.com</email>
          <email>alessandro.simonetta@gmail.com ORCID: 0000-0003-2002-9815</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Torchiano</string-name>
          <email>marco.torchiano@polito.it</email>
          <email>marco.torchiano@polito.it ORCID: 0000-0001-5328-368X</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Vetrò</string-name>
          <email>antonio.vetro@polito.it</email>
          <email>antonio.vetro@polito.it ORCID: 0000-0003-2027-3308</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Cristina Paoletti</string-name>
          <email>mariacristina.paoletti@gmail.com</email>
          <email>mariacristina.paoletti@gmail.com ORCID: 0000-0001-6850-1184</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Enterprise Engineering, University of Rome Tor Vergata</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Control and Computer Eng., Politecnico di Torino</institution>
          ,
          <addr-line>Turin</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>- In the last decades the exponential growth of available information, together with the availability of systems able to learn the knowledge that is present in the data, has pushed towards the complete automation of many decisionmaking processes in public and private organizations. This circumstance is posing impelling ethical and legal issues since a large number of studies and journalistic investigations showed that software-based decisions, when based on historical data, perpetuate the same prejudices and bias existing in society, resulting in a systematic and inescapable negative impact for individuals from minorities and disadvantaged groups. The problem is so relevant that the terms data bias and algorithm ethics have become familiar not only to researchers, but also to industry leaders and policy makers. In this context, we believe that the ISO SQuaRE standard, if appropriately integrated with risk management concepts and procedures from ISO 31000, can play an important role in democratizing the innovation of software-generated decisions, by making the development of this type of software systems more socially sustainable and in line with the shared values of our societies. More in details, we identified two additional measure for a quality characteristic already present in the standard (completeness) and another that extends it (balance) with the aim of highlighting information gaps or presence of bias in the training data. Those measures serve as risk level indicators to be checked with common fairness measures that indicate the level of polarization of the software classifications/predictions. The adoption of additional features with respect to the standard broadens its scope of application, while maintaining consistency and conformity. The proposed methodology aims to find correlations between quality deficiencies and algorithm decisions, thus allowing to verify and mitigate their impact.</p>
      </abstract>
      <kwd-group>
        <kwd>ISO Square</kwd>
        <kwd>ISO31000</kwd>
        <kwd>data ethics</kwd>
        <kwd>data quality</kwd>
        <kwd>data bias</kwd>
        <kwd>algorithm fairness</kwd>
        <kwd>discrimination risk</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Software nowadays replace most human decisions in
many contexts [1] ; the rapid pace of innovation suggests that
this phenomenon will further increase in the future [2]. This
trend has been enabled by the large availability of data and of
the technical means to analyze them for building the
predictive, classification, and ranking models that are at the
core of automated decision making (ADM) systems.
Advantages for using ADM systems are evident and they
concern mainly scalability, efficiency, and removal of
decision makers’ subjectivity. However, several critical
aspects have emerged: lack of accountability and transparency
[3], massive use of natural resources and low-unpaid labor to
building extensive training sets [
        <xref ref-type="bibr" rid="ref2">4</xref>
        ] , the distortion of the public
sphere of political discussion [
        <xref ref-type="bibr" rid="ref3">5</xref>
        ], and the amplification of
existing inequalities in society [
        <xref ref-type="bibr" rid="ref4">6</xref>
        ]. This paper focuses on the
latter problem, which occurs when automated software
decisions “systematically and unfairly discriminate against
certain individuals or groups of individuals in favor of others
[by denying] an opportunity for a good or [assigning] an
undesirable outcome to an individual or groups of individuals
on grounds that are unreasonable or inappropriate” [
        <xref ref-type="bibr" rid="ref5">7</xref>
        ] . In
practice, software systems may perpetuate the same bias of
our societies, systematically discriminating the weakest
people and exacerbating existing inequalities [
        <xref ref-type="bibr" rid="ref6">8</xref>
        ]. A recurring
cause for this phenomenon is the use of incomplete and biased
data, because of errors or limitations in the data collection
(e.g., under-sampling of a specific population group) or
simply because the distributions of the original population are
skewed. From a data engineering perspective, this translates
into imbalanced data, i.e. a condition with an unequal
distribution of data between the classes of a given attribute,
which causes highly heterogeneous accuracy across the
classifications [
        <xref ref-type="bibr" rid="ref7">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">10</xref>
        ]. Imbalanced data is known to be
problematic in the machine learning domain since long time
[
        <xref ref-type="bibr" rid="ref9">11</xref>
        ]. In fact, imbalanced datasets may lead to imbalanced
results, which in the context of ADM systems means
differentiation of products, information and services based on
personal characteristics. In applications such as allocation of
social benefits, insurance tariffs, job profiles matching, etc.,
such differentiations can lead to unjustified unequal treatment
or discrimination.
      </p>
      <p>
        For this reason, we maintain that imbalanced and
incomplete data shall be considered as a risk factor in all the
ADM systems that rely on historical data and operate in
relevant aspects of the lives of individuals. Our proposal relies
on the integration of the measurement principles of the ISO
SQuaRE [
        <xref ref-type="bibr" rid="ref10">12</xref>
        ] with the risk management process defined in
ISO 31000 [
        <xref ref-type="bibr" rid="ref11">13</xref>
        ] to assess the potential risk of discriminating
software output and take action for remediations. In the paper,
we describe the theoretical foundations, and we provide a
workflow of activities. We believe that the approach can be
useful to a variety of stakeholders for assessing the risk of
discriminations, including the creators or commissioners of
software systems, researchers, policymakers, regulators,
certification or audit authorities. Assessments should prompt
taking appropriate action to prevent adverse effects.
      </p>
      <p>II.</p>
    </sec>
    <sec id="sec-2">
      <title>METHODOLOGY</title>
      <p>Figure 1 gives an overview of the proposed methodology.
The process begins with the common subdivision of the
original data into training and test data. At this point, it is
possible to measure the quality in the training data (balance
and completeness) and the fairness in the results obtained on
the test data. Data balance measures extend the characteristics
of the data quality model (ISO/IEC 25012), while
completeness measures complement it. Data quality measures
give rise to an indicator for unbalanced or incomplete data for
the sensitive characteristics, which implicates a risk of biased
classifications by the algorithm. In this circumstance, it is
necessary to also assess the fairness of the algorithms used
through the measures outlined in this paper. The presence of
unfair results from the point of view of sensitive features in
correspondence with poor quality data leads to the necessary
data enrichment step to try to mitigate the problem. Thus, our
proposed methodology is composed by two main blocks:
A. Risk analysis: measuring the risk that a training
set could contain unbalanced data, integrating the
SQuaRE approach with ISO 31000 risk
management principles;
B. Risk evaluation: verify that a high level of risk
corresponds to unfairness, and in positive case
enrich original data with synthetic data to
mitigate the problem.</p>
      <p>A. Risk analysis: where SQuaRE and ISO3100 meet</p>
      <p>We integrate the SQuaRE theoretical framework with the
ISO 31000 risk management principles to measure the risk
that an unbalanced or incomplete training set might cause
discriminating software output. Since the primary recipients
of this document are the participants of the “3rd International
Workshop on Experience with SQuaRE series and its Future
Direction”1, we do not describe here the standard, however we
summarize the most important aspects for the scope of the
paper. Firstly, we remind that SQuaRE includes quality
modeling and measurements of software products2, data and
software services. According to the philosophy and
organization of this family of standards, quality is categorized
into one or more quantifiable characteristics and
subcharacteristics. For example, the standard ISO/IEC
25012:2011 formalizes the product quality model as
composed of eight characteristics, which are further
subdivided into sub-characteristics. Each (sub) characteristic
relates to static properties of software and dynamic properties
of the computer system3. The ISO/IEC 25012:2008 on data
quality has 15 characteristics: 5 of them belongs to the
“inherent” point of view (i.e., the quality relies only on the
characteristics of the data per se), 3 of them are
systemdependent (i.e., the quality depends on the characteristics of
the system hosting the data and making it available), the
remaining 7 belonging to both points of view. Data balance is
not recognized as a characteristic of data quality in ISO/IEC
25012:2008: it is proposed here as an additional inherent
characteristic. Because of its role in the generation of biased
1See http://www.sic.shibaura-it.ac.jp/~tsnaka/iwesq.html
2 A software product is a “set of computer programs, procedures,
and possibly associated documentation and data” as defined in
ISO/IEC 12207:1998. In SQuaRE standards, software quality
stands for software product quality.
3 A system is the “combination of interacting elements organized
to achieve one or more stated purposes” (ISO/IEC 15288:2008),
for example the aircraft system. It follows that a computer system
is “a system containing one or more components and elements
such as computers (hardware), associated software, and data”, for
example a conference registration system. An ADM system that
determines eligibility for aid for drinking water is a software
system.
software output, data
balance reflects the
propagation
principle of SQuaRE: the quality of the software product,
service and</p>
      <p>data affects the quality in use. Therefore,
evaluating and improving product/service/data quality is one
mean of improving the system quality in use. A simplification
of this concept is the GIGO principle (“garbage in, garbage
out”): data that is outdated, inaccurate and incomplete make
the output of the software unreliable. Similarly, unbalanced
data
will
probably
cause
unbalanced
software
output,
especially in the context of machine learning and AI systems
trained
with that
data.</p>
      <p>This
principle
applies
also to
completeness, which is already an inherent characteristic of
data quality in SQuaRE: in this work we propose an additional
metric to those proposed in ISO/IEC 25024:2015, that is more
suitable for the problem of biased software.</p>
      <p>To better address the problem of biased software output,
we consider the measures of data balance and completeness
not only as extension of SQuaRE data quality modelling but
also as risk factors. Here comes the integration of SQuaRE
theoretical and
measurement framework
with the ISO
31000:2018 standard for risk
management. The standard
defines guiding principles and a process of three phases: risk
identification, risk analysis and risk evaluation. Here, we
briefly describe them and specify the relation
with our
approach.</p>
      <p>Risk
identification
refers to finding, recognizing
and
describing risks within a certain context and scope, and with
respect to specific criteria defined prior to risk assessment. In
this paper, this is implicitly contained in the motivations and
in the problem formulation: it is the risk of discriminating
individuals or groups of individuals by operating software
systems that automate high-stake decisions for the lives of
people.</p>
      <p>Risk analysis is the understanding of the characteristics and
levels of the risk. This is the phase where measures of data
balance and completeness are used as indicators, due to the
propagation effect previously described.</p>
      <p>Risk evaluation, as the last step, is the process in which the
results of the analysis are taken into consideration to decide
whether additional action is required. If affirmative, this
process would then outline available risk treatment options
and the need for conducting additional analyses. In our case,
specific thresholds for the measures should be decided for the
specific prediction/classification algorithms used, the social
context, the legal requirements of the domain, and other
relevant factors for the case at hand. In addition to the
technical actions, the process would define other types of
required actions (e.g., reorganization of decision processes,
communication to the public, etc.) and the actors who must
undertake them.</p>
      <sec id="sec-2-1">
        <title>1) Completeness measure</title>
        <p>The completeness measure proposed is agnostic with respect
to classical ML data classification because for our purposes
we are interested in evaluating those columns that assume
values in finite and discrete intervals, which we will call
categorical with respect to the row data. This characteristic
will allow us to consider the set of their values as the digits
constituting a number in a variable base numbering system.
The idea of the present study is based on the principle that a
learning system provides predictions consistent with the data
with which it has been trained. Therefore, if it is fed with
nonhomogeneous
data
it
will
provide
unbalanced
and
discriminatory predictions with respect to reality. For this
reason, the methodology we propose starts with the analysis
phase of the reality of interest and of the dataset, an activity
that must be carried out even before starting the pre-training
phase in line with previous studies where some of the authors
proposed the use of balance measures in automated decision</p>
        <sec id="sec-2-1-1">
          <title>Index</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Gini</title>
      <sec id="sec-3-1">
        <title>Formula Normalized formula</title>
        <p>= 1 −

∑   2
 =1
  =</p>
        <p>− 1
∙ (1 −

∑   2)
 =1
 
=</p>
        <p>1
 − 1
( 
∑</p>
        <p>1
 =1 
 2 − 1)</p>
        <p>Notes
m is the number of classes
f is the relative frequency of each class
  =
 

∑
 =1  
ni= absolute frequency
The higher G and Gn, the higher is the
heterogeneity: it means that categories
have similar frequencies</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>The lower the index, the lower is the heterogeneity: a few classes account for majority of instances</title>
    </sec>
    <sec id="sec-5">
      <title>For m, f, fi and ni check Gini</title>
      <p>Higher values of D and Dn indicate higher
diversity in terms of probability of
belonging to different classes</p>
    </sec>
    <sec id="sec-6">
      <title>The lower the index, the lower is the diversity, because frequencies are concentrated in a few classes</title>
      <p>
        making systems [
        <xref ref-type="bibr" rid="ref12">14</xref>
        ][
        <xref ref-type="bibr" rid="ref13">15</xref>
        ][
        <xref ref-type="bibr" rid="ref14">16</xref>
        ][
        <xref ref-type="bibr" rid="ref15">17</xref>
        ]. In particular, during this
phase, it is necessary to identify all the independent columns
that define whether the instance belongs to a class or category.
Suppose we have a structured dataset as follows:
      </p>
      <p>DS = { C0, C1, ... , Cn−1 }</p>
      <p>with the set S the positions of the columns
categorising the instances, functionally independent of the
other columns in the dataset:</p>
      <p>S ⊆ { 0, 1, ... , n – 1 } , dim(S) = m , m ≤ n
(1)
(2)
we can analyze the new dataset consisting of the columns
CS(j) with j ∈ [0, m − 1].</p>
      <p>Having said that, we can decide to use two different notions
of completeness: maximum or minimum. In the first case the
presence in the dataset of a greater number of distinct
instances that belong to the same categorising classes
constitutes a constraint for all the other instances of the
dataset. That is, one must ensure that one has the same number
of replicas of distinct class combinations for distinct instances.
Instead, in the second case it is sufficient to have at least one
combination of distinct classes among those possible for each
instance. For simplicity, but without loss of generality of the
procedure, we will explore the minimum completeness of the
dataset, then we will reduce the dataset to just the columns ( )
by removing duplicate rows. We will use the Python language
to
explicate
the
calculation
formulas
and
make
the
mathematical logic implied less abstract. The Python language
has the pandas library, which makes it possible to carry out
analysis and data manipulation in a fast, powerful, flexible and
easy-to-use
manner. Through the</p>
      <p>DataFrame class it is
possible to load data frames from a simple csv file:
import pandas as pd
df=pd.read_csv(&lt;file_name&gt;)
The ideal value of minimum
completeness for the
combinatorial metric is when in the dataset there is at least one
instance that
belongs to
each
distinct
combination
of
categories. The absence of some combination could create the
lack of information that we do not want to exist. To calculate
the total number of distinct combinations we need to calculate
the product of the distinct replicas per single category.
k=( df['CS0'].unique().size *
df['CS1'].unique().size *...*
df['CSm-1'].unique().size )</p>
      <p>On the other hand, in the dataset we only have the
characterising columns so we can derive the true number of
distinct instances in order to determine how far the data
deviates from the ideal case.</p>
      <p>len (df.drop_duplicates())/ k</p>
      <p>The value for maximum completeness is calculated from
the maximum number of duplicates of the same combinations
of characterising columns. For this reason it is necessary to
maintain in the dataset in addition to the columns ( ) a
discriminating identification field of the rows with the same
values in these columns. To determine the potential total, once
to all other classes.</p>
      <p>M= df.groupby(['CS0',...,'CSm-1']).</p>
      <p>size().reset_index(name='counts').</p>
      <p>counts.max()
len(df)/(M∗k)</p>
      <sec id="sec-6-1">
        <title>2) Balance measures</title>
        <p>
          Since imbalance is defined as an unequal distribution
between classes [
          <xref ref-type="bibr" rid="ref7">9</xref>
          ], we focus on categorical data. In fact,
most of the sensitive attributes are considered categorical
data, such as gender, hometown, marital status, and job.
Alternatively, if they are numeric, they are either discrete and
within a short range, such as family size, or they are
continuous but often re-conducted to distinct categories, such
as information on “age” which is often discretized into ranges
such as “&lt; 18”, “19-35”, “36-50”, “51-65”, etc. We show two
examples of measures in Table 1, retrieved from the literature
of social and natural sciences, where imbalance is known in
terms of (lack of) heterogeneity and diversity. They are
normalized in the range 0-1, where 1 correspond to maximum
balance and 0 to minimum balance, i.e. imbalance. Hence,
lower level of balance measures mean a higher risk of bias in
the software output.
        </p>
        <p>
          B. Risk evaluation with fairness measures
The majority of fairness measures in machine learning
literature rely on the comparison of accuracy, computed for
each population group of interest [
          <xref ref-type="bibr" rid="ref16">18</xref>
          ]. For computing the
accuracy, two different approaches can be adopted: the first
attempts to measure the intensity of errors, i.e. the deviation
between prediction and actual value (precision), while the
other measures the general direction of the error. Indicating
with ei the ith error, with fi and di respectively the ith forecast
and demand, we have:
At this point we can add up all the errors with sign and find
the average error:
        </p>
        <p>=   −  


= 1 ∑  
However, this
measure is very crude because
error
compensation phenomena may be present so generically it is
preferred to use the mean of the absolute error or the square
root of the mean square error:


= 1 ∑ |  |
= √

1
∑   2
(3)
(4)
(5)
(6)</p>
        <p>RMSE is sensitive to important errors, while from this
point of view MAE is fairer because it considers all errors at
the same level. Moreover, if our prediction tends to the
median it will get a good value of MAE, vice versa if it
approaches the mean it will get a better result on RMSE.</p>
        <p>Under conditions where the median is lower than the
mean, for example in processes where there are peaks of
demands compared to normal steady state operation, it will
not be convenient to use MAE which will introduce a bias
while it will be more convenient to use RMSE. Things are
reversed if outliers are present in the distribution as MAE is
less sensitive than RMSE.</p>
        <p>To measure model performance, you can choose to
measure error with one or more KPI.</p>
        <p>In the case of classification algorithms instead you can use
the confusion matrix that allows you to compute the number
of true positives (TPs), true negatives (TNs), false positives
(FPs) and false negatives (FNs):</p>
        <p>At this point, the single values could be computed for each
population
subgroup
(e.g., “Asian”
vs “Caucasian”
vs
“African-American” etc. , or “Male” vs “Female”, etc.) and
the same applies to the concepts of precision, recall, and
accuracy, known from the literature and reported here:
 = [ …
 11
  1
…
. . .
…  
 1
… ]
 ( ) =  
 ( ) =</p>
        <p>( ) =
 ( ) =</p>
        <p>∑
 =1, ≠</p>
        <p>∑
 =1, ≠</p>
        <p>∑
 =1, ≠</p>
        <p>
          You can use the following equations to calculate the
following values:

=
Fairness
measures
should be then compared
with
appropriate thresholds selected with respect to social context
in which the software application is used. If the unfairness is
higher than the
[
          <xref ref-type="bibr" rid="ref17">19</xref>
          ]); other rebalancing techniques have been proposed in the
literature (e.g. SMOTE [
          <xref ref-type="bibr" rid="ref18">20</xref>
          ], ROSE [
          <xref ref-type="bibr" rid="ref19">21</xref>
          ]).
        </p>
        <p>III.</p>
        <p>
          RELATION WITH LITERATURE AND OUR PAST STUDIES
An approach similar to ours is the work of Takashi
Matsumoto and Arisa Ema [
          <xref ref-type="bibr" rid="ref20">22</xref>
          ], who proposed a risk chain
model (RCM) for risk reduction in Artificial Intelligence
services: the authors consider both data quality and data
imbalance as risk factors. Our work can be easily integrated
into the RCM framework, because we offer a quantitative
way to measure balance and completeness, and because it is
natively related to the ISO/IEC standards on data quality
requirements and risk management.
4 It is a joint initiative of MIT Media Lab and Berkman Klein
Center at Harvard University: https://datanutrition.org/.
        </p>
        <p>Other approaches which can be connected to ours are in
the direction of labeling datasets: for example “The Dataset</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Nutrition</title>
    </sec>
    <sec id="sec-8">
      <title>Label</title>
      <p>Project” 4
aims to identify
the
“key
ingredients” in a dataset such as provenance, populations, and
missing data. The label takes the form of an interactive
visualization
mentioned</p>
      <p>that
aspects
allows for
exploring the</p>
      <p>
        previously
and
spot
flawed,
incomplete,
or
problematic data. One of the author of this paper took
inspiration from this study in previous works for “Ethically
and socially-aware labeling” [
        <xref ref-type="bibr" rid="ref14">16</xref>
        ] and for a data annotation
and
visualization
schema based
on
      </p>
    </sec>
    <sec id="sec-9">
      <title>Bayesian</title>
      <p>
        statistical
inference [
        <xref ref-type="bibr" rid="ref15">17</xref>
        ] always for the purpose of warning about the
risk of discriminatory outcomes due to poor quality of
datasets.
      </p>
      <p>
        We
started
from
that experience to
conduct
preliminary case studies on the reliability of the balance
measures [
        <xref ref-type="bibr" rid="ref12">14</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">15</xref>
        ]: in this work we continue in that direction
by adding a measure of completeness and proposing an
explicit workflow of activities for the combination of SQuaRE
with ISO 31000.
      </p>
      <p>IV.</p>
      <sec id="sec-9-1">
        <title>CONCLUSION AND FUTURE WORK</title>
        <p>We propose a methodology that integrates SQuaRE
measurement framework with the ISO 31000 process with the
goal of evaluating balance and completeness in a dataset as
risk factors of discriminatory outputs of software systems. We
believe that the methodology can be a useful instrument for all
the actors involved in the development and regulation of
ADM systems, and, from a more general perspective, it can
play an important role in the collective attempt of placing
democratic control on the development of these systems, that
should be more accountable and less harmful than how they
are now. In fact, the adverse effects of ADM systems are
posing a significant danger for human rights and freedoms as
our societies increasingly rely on automated decision making.
It must be stressed that this is still at a prototypical stage and
further studies are necessary to improve the methodology and
to assess the reliability of the proposed measures, for example
to find meaningful risk thresholds in relation to the context of
use and the severity of the impact on individuals. The current
paper is also way to seek engagement from other researchers
in a community effort to test the workflow in real settings,
improve it and build an open registry of additional measures
combined</p>
        <p>with evaluations benchmark. Finally, we are
conscious that technical adjustments are not enough, and they
should be integrated with other types of actions because of the
socio-technical nature of the problem.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Available:</title>
      <p>https://automatingsociety.algorithmwatch.org
[2]</p>
      <p>E. Brynjolfsson and A. McAfee, The Second Machine
[3]</p>
      <p>Age: Work, Progress, and Prosperity in a Time of
Brilliant Technologies, Reprint edition. New York
London: W. W. Norton &amp; Company, 2016.
F. Pasquale, The Black Box Society: The Secret
Algorithms That Control Money and Information.</p>
    </sec>
    <sec id="sec-11">
      <title>Cambridge: Harvard Univ Pr, 2015.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Spielkamp</surname>
          </string-name>
          , “
          <source>Automating Society Report</source>
          <year>2020</year>
          ,” Berlin, Oct.
          <year>2020</year>
          . Accessed: Nov.
          <volume>10</volume>
          ,
          <year>2020</year>
          . [Online].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Crawford</surname>
          </string-name>
          , Atlas of AI: Power, Politics, and
          <source>the Planetary Costs of Artificial Intelligence: The Real Worlds of Artificial Intelligence. New Haven: Yale Univ Pr</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Howard</surname>
          </string-name>
          , Lie Machines:
          <article-title>How to Save Democracy from Troll Armies</article-title>
          , Deceitful Robots, Junk News Operations, and
          <string-name>
            <given-names>Political</given-names>
            <surname>Operatives</surname>
          </string-name>
          . New Haven ; London: Yale Univ Pr,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>V.</given-names>
            <surname>Eubanks</surname>
          </string-name>
          ,
          <article-title>Automating Inequality: How High-Tech Tools Profile, Police, and Punish the Poor</article-title>
          . New York, NY: St. Martin's Press,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Friedman</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Nissenbaum</surname>
          </string-name>
          , “Bias in Computer Systems,”
          <source>ACM Trans Inf Syst</source>
          , vol.
          <volume>14</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>330</fpage>
          -
          <lpage>347</lpage>
          , Jul.
          <year>1996</year>
          , doi: 10.1145/230538.230561.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [8]
          <string-name>
            <surname>C. O'Neil</surname>
          </string-name>
          , Weapons of Math Destruction:
          <article-title>How Big Data Increases Inequality and Threatens Democracy, Reprint edition</article-title>
          . New York: Broadway Books,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>He</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Garcia</surname>
          </string-name>
          , “
          <article-title>Learning from Imbalanced Data,”</article-title>
          <source>IEEE Trans. Knowl. Data Eng.</source>
          , vol.
          <volume>21</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>1263</fpage>
          -
          <lpage>1284</lpage>
          , Sep.
          <year>2009</year>
          , doi: 10.1109/TKDE.
          <year>2008</year>
          .
          <volume>239</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Krawczyk</surname>
          </string-name>
          , “
          <article-title>Learning from imbalanced data: open challenges and future directions,” Prog</article-title>
          . Artif. Intell., vol.
          <volume>5</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>221</fpage>
          -
          <lpage>232</lpage>
          , Nov.
          <year>2016</year>
          , doi: 10.1007/s13748-016-0094-0.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Japkowicz</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Stephen</surname>
          </string-name>
          , “
          <article-title>The class imbalance problem: A systematic study</article-title>
          ,
          <source>” Intell. Data Anal.</source>
          , vol.
          <volume>6</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>429</fpage>
          -
          <lpage>449</lpage>
          , Oct.
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [12] International Organization for Standardization,
          <source>“ISO/IEC 25000:2014 Systems and software engineering - Systems and software Quality Requirements</source>
          and
          <string-name>
            <surname>Evaluation (SQuaRE) - Guide to</surname>
            <given-names>SQuaRE</given-names>
          </string-name>
          ,” ISO-International Organization for Standardization,
          <year>2014</year>
          . https://www.iso.org/standard/64764.html (accessed
          <year>Nov</year>
          .
          <volume>10</volume>
          ,
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [13] International Organization for Standardization, “ISO 31000:
          <year>2018</year>
          <article-title>Risk management</article-title>
          - Guidelines,” ISO - International Organization for Standardization,
          <year>2018</year>
          . https://www.iso.org/cms/render/live/en/sites/isoorg/co ntents/data/standard/06/56/65694.html (accessed
          <year>Nov</year>
          .
          <volume>10</volume>
          ,
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mecati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. E.</given-names>
            <surname>Cannavò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetrò</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Torchiano</surname>
          </string-name>
          , “
          <article-title>Identifying Risks in Datasets for Automated Decision-Making,” in Electronic Government</article-title>
          , Cham,
          <year>2020</year>
          , pp.
          <fpage>332</fpage>
          -
          <lpage>344</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -57599-1_
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetrò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Torchiano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mecati</surname>
          </string-name>
          , “
          <article-title>A data quality approach to the identification of discrimination risk in automated decision making systems,”</article-title>
          <string-name>
            <surname>Gov. Inf. Q.</surname>
          </string-name>
          , p.
          <fpage>101619</fpage>
          ,
          <string-name>
            <surname>Sep</surname>
          </string-name>
          .
          <year>2021</year>
          , doi: 10.1016/j.giq.
          <year>2021</year>
          .
          <volume>101619</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>E.</given-names>
            <surname>Beretta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetrò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lepri</surname>
          </string-name>
          , and J.
          <string-name>
            <surname>C. De Martin</surname>
          </string-name>
          , “Ethical and
          <string-name>
            <surname>Socially-Aware Data</surname>
            <given-names>Labels</given-names>
          </string-name>
          ,” in Information Management and
          <string-name>
            <given-names>Big</given-names>
            <surname>Data</surname>
          </string-name>
          , Cham,
          <year>2019</year>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Beretta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetrò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lepri</surname>
          </string-name>
          , and J.
          <string-name>
            <surname>C. D. Martin</surname>
          </string-name>
          , “
          <article-title>Detecting discriminatory risk through data annotation based on Bayesian inferences</article-title>
          ,”
          <source>in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</source>
          , New York, NY, USA, Mar.
          <year>2021</year>
          , pp.
          <fpage>794</fpage>
          -
          <lpage>804</lpage>
          . doi:
          <volume>10</volume>
          .1145/3442188.3445940.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Barocas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hardt</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <source>Fairness and Machine Learning. fairmlbook.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [19] “
          <article-title>Handling imbalanced data sets with synthetic boundary data generation using bootstrap re-sampling and AdaBoost techniques - ScienceDirect</article-title>
          .” https://doi.org/10.1016/j.patrec.
          <year>2013</year>
          .
          <volume>04</volume>
          .019 (
          <issue>accessed</issue>
          <year>Sep</year>
          .
          <volume>18</volume>
          ,
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Chawla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. W.</given-names>
            <surname>Bowyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. O.</given-names>
            <surname>Hall</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. P.</given-names>
            <surname>Kegelmeyer</surname>
          </string-name>
          , “
          <article-title>SMOTE: synthetic minority oversampling technique,”</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Artif</surname>
          </string-name>
          .
          <source>Intell. Res.</source>
          , vol.
          <volume>16</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          , Jun.
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Menardi</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Torelli</surname>
          </string-name>
          , “
          <article-title>Training and assessing classification rules with imbalanced data,” Data Min</article-title>
          . Knowl. Discov., vol.
          <volume>28</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>122</lpage>
          , Jan.
          <year>2014</year>
          , doi: 10.1007/s10618-012-0295-5.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>T.</given-names>
            <surname>Matsumoto</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Ema</surname>
          </string-name>
          , “
          <article-title>RCModel, a Risk Chain Model for Risk Reduction in</article-title>
          AI Services,” ArXiv200703215 Cs, Jul.
          <year>2020</year>
          , [Online]. Available: http://arxiv.org/abs/
          <year>2007</year>
          .03215
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>