<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Can Justice be a measurable value for AI? Proposed evaluation of the relationship between NLP models and principles of Justice</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lidia Marassi</string-name>
          <email>lidia.marassi@unina.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Narendra Patwardhan</string-name>
          <email>narendraprakash.patwardhan@unina.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Gargiulo</string-name>
          <email>francesco.gargiulo@icar.cnr.i</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for High Performance Computing and Networking of National Research Council, ICAR-CNR</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Naples Federico II, Department of Electrical Engineering and Information Technology</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>NLP models, such as chat-based generative ones, have gained widespread use and influence in various fields. Considering AI as a valuable resource, it seems critical to determine the ethical principles to which these models should adhere to be considered usable. Applying the philosophical concept of justice to the evaluation of NLP models can help improve their functioning, performance, and overall ethical implications. Although strict adherence to traditional codes of justice may not be sufficient, it is proposed that the concept be adapted to evaluate the performance of NLP systems. Creating a rating scale based on these values may help estimate the "amount of Justice" or fairness demonstrated by different NLP systems. By emphasizing the importance of certain values central to assessing the fairness of NLP models, it becomes possible to develop parameters for evaluation. As these models become increasingly pervasive, this research can help improve this models' performance and their socio-cultural significance. The awareness of building more equitable technologies raises consciousness about the potential consequences of irresponsible use of these models, highlighting the significance of ethical considerations. This approach can lead to a more comprehensive understanding of these systems, promoting their improvement and fostering greater knowledge of the ethical implications associated with their functioning.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Justice</kwd>
        <kwd>NLP</kwd>
        <kwd>trustworthy</kwd>
        <kwd>ethics</kwd>
        <kwd>fairness</kwd>
        <kwd>measure 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>"What is the right measure of justice?". If this question is the basis of historical philosophical
debate, today the issue of determining Justice must come to terms with technological progress.
That is, what does justice mean in the field of AI? By what metrics can one assess whether a
system is fair or just? More importantly, how should this fairness be measured?</p>
      <p>These questions are important not only from a moral point of view in the classical sense, as
we consider that different philosophical theories offer different approaches to measurement, but
also ethically from a more practical point of view. Artificial intelligence is now within everyone's
reach (or nearly so), and technological solutions are part of our everyday lives. NLP models have
made massive progress in recent years, finding application in a variety of different fields. In recent
times, chat-based generative models have experienced an incredible surge in interest and use
thanks to numerous companies releasing their language models for free or open-source
(ChatGPT, You Ask, Microsoft Bing are some of the early examples of this phenomenon).</p>
      <p>The growing interest in NLP models and their pervasiveness suggest that it would be prudent
to evaluate their performance in terms of fairness. The philosophical concept of Justice, in this
case, could be applied precisely to the evaluation of these technological solutions. In this paper
we propose a possible theoretical basis for the development of a system for evaluating the degree
of justice of NLP models. By asking ourselves to consider the "rightness" of these models, we can
help not only improve its operation and performance, but also to make society, and AI, more
ethical.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Justice theories and AI</title>
      <p>
        The topic of Justice is not new to the field of AI. In search of inspiration and practicality, several
researchers have turned to theories of distributive justice, and in particular, the theories of John
Rawls [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], recently dubbed as "the favorite philosopher of artificial intelligence" [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Drawing
from the philosophical realm, researchers aim to mitigate social injustices by quantifying,
measuring, and redistributing resources among individuals in society [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Theories of distributive
justice propose that justice can be assessed, at least in part, by examining how benefits and
burdens are equitably distributed among individuals in society [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These theories are valuable
for community members, as researchers and policy makers, as they provide proper guidance for
understanding the concept of Justice and determining the relative fairness of different societies
or systems. In the field of machine learning, an area called fair machine learning (fair ML) has
been emerging with the goal of addressing "algorithmic injustice" or bias, as it seeks to apply
theories of distributive justice to the design of machine learning systems [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Researchers have also attempted to formalize these theories from a quantitative perspective,
focusing on concepts such as equality of opportunity [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and elements of John Rawls' influential
distributive justice theory [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Despite the great influence that theories of distributive justice seem to have in the field of AI,
here we instead propose an approach closer to the so-called Capability Approach. As argued by
Amartya Sen [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Rawls' measure of justice would indeed seem to prove insensitive to the
heterogeneities of people and social context.
      </p>
      <p>Especially when applied to the field of AI, distributive justice approach may not consider the
different ways in which some people are (or are not) able to convert resources into well-being. It
is considered that what usually represents an opportunity for most people, namely converting a
resource into a valuable state of being or acting, may be precluded to others because of their
individual differences and social context. For example, for a blind person, the purchase of a
computer does not result in a means to enjoy the benefits of Internet access if the computer is not
compatible with screen reader technology. In this situation, the computer becomes a false
opportunity to achieve the corresponding valuable state, which is surfing the Internet.</p>
      <p>NLP models, indeed, can be used for text creation (e.g., articles, reports, product
descriptions...); however, if the model does not consider cultural sensitivities or lexical
appropriateness, it could generate inappropriate or offensive text for certain user groups. This
would make the automatic generation resource a false opportunity to create relevant and
respectful content for different communities of readers. Capability theorists argue that the real
support for well-being lies not in the resources themselves, but in what people are able to achieve
through their use. Thus, Justice does not reside in resources, but in people's abilities to use them.</p>
      <sec id="sec-2-1">
        <title>2.1. Ethical principles</title>
        <p>
          For individuals to be able to use resources to improve their own well-being, it seems necessary
that access to them be guaranteed by a set of ethical principles. If we consider AI as a resource (to
date, perhaps one of the richest), then the question about the principles it should adhere to be
usable is indeed fundamental. While sticking strictly to classical codifications of the concept of
Justice can lead to a dead end, in this paper, we suggest instead adapting the concept to
performance evaluation of NLP systems. Referring to the capability approach theory, it considers
several ethical principles [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], among which we consider:
 Freedom: It is argued that Justice should be evaluated since people's ability to realize
their own choices. This implies promoting a wide range of freedoms and opportunities
to enable individuals to pursue their aspirations.
 Human Dignity: as the inherent dignity of every human being. Justice should ensure
that people have the capabilities and have access to the resources they need to live a
life of dignity.
 Equality of Opportunity: this principle emphasizes the importance of providing all
individuals with equitable access to the opportunities they need to realize their
potential.
 Social inclusion: it calls for the removal of discrimination and inequalities that limit
participation and inclusion.
 Sustainability: both environmentally and socially. Indeed, justice requires the
responsible management of resources and the preservation of opportunities for future
generations.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Ethical principles as measurable values for NLP models</title>
      <p>In the field of AI, with reference to NLP models, there are several values that seem to be
significant with respect to this view of Justice. We believe that it would be possible to assess a
model's ability to satisfy these principles of justice if they were adapted to NLP models. As
chatbots are artificial intelligence systems designed to interact with people, they can have a
significant impact on people's lives, both in terms of well-being and justice.</p>
      <p>The identification of values, in terms of variables, can be the basis for creating a rating scale to
estimate the "amount of justice" (which we will mean in terms of fairness) of different NLP
systems. Based on the attribution of a numerical estimate of the NLP model passed under
consideration, it will be possible to quantify its "fairness" in practical terms. This result would not
only allow for a more in-depth analysis of these systems to work toward improving their
performance but would also have significance in socio-cultural terms.</p>
      <p>Indeed, awareness of the need to build more equitable technologies seems to be the basis for
increased awareness of the possible consequences of ill-considered use of these models,
especially if they took ethical aspects into account. With this approach, an in-depth and
quantitative evaluation of each parameter considered in NLP models is proposed, helping to
understand how the values identified affect the fairness of the model.</p>
      <p>The process proposed in the research involves:
(A) Identification of Reference Values and Contextualization in NLP Systems:
o Identifying the value (e.g., fairness): understanding what this value represents within NLP
systems (e.g., fairness as fairness of treatment);
o Relevance of Value to Model Fairness: understand why the identified value is relevant to
greater fairness in the model (e.g., a fair model does not discriminate and provides accurate
and relevant results for all users);
o Choice of Evaluation Metrics: identify suitable metrics for quantitative measurement of the
identified value, also drawing on relevant scientific literature (e.g., to assess fairness, an
important metric might be the degree to which the model is able to identify personal
information regarding the interlocutor, such as gender, after interaction with the algorithm.
A fair model should avoid making assumptions or discriminating based on these personal
characteristics. Thus, the more the model can predict who it is talking to, the less fair it would
be considered).</p>
      <p>(B) Quantitative Assessment of Model Fairness:
o Measurement of Identified Values: measuring the values identified in the performance of the</p>
      <p>NLP model;
o Assigning Numerical Scores: assigning a numerical score to each evaluated value,
representing the amount of fairness present in each aspect considered;
o Creating a Graph: representing the measured values through a graph that will show the
amount of equity for the different aspects considered.</p>
      <p>In the end, a graph with the measured values will be obtained, assigning a number to each
evaluated value. On this basis, it will be possible to have a quantitative measure of the fairness of
the model, differentiated for the different aspects considered. Here, we will focus on explaining
the importance of certain values, central to the evaluation of the fairness of NLP models, which
can later be developed in terms of parameters to be evaluated.</p>
      <sec id="sec-3-1">
        <title>3.1. Transparency – Freedom</title>
        <p>
          Transparency is a fundamental pillar in evaluating the justice of NLP systems. Transparency
in AI is often used interchangeably with explainable AI, but it’s more focused on ensuring that a
model is open and visible [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Users must have the ability to understand how these systems make
decisions and process their data. In the Capabilities Approach, freedom of choice is considered a
fundamental element of human well-being. It is not enough to have different options available,
but it is essential that people have a real opportunity to understand and evaluate these options
so that they can make informed and well-informed choices. This often implies the need for access
to clear, understandable and transparent information about the options available and the
consequences of the choices that can be made.
        </p>
        <p>Providing access to model and training data internals enables users to examine the inner
workings of NLP systems and ensures accountability. Model cards provide detailed information
about the structure, design choices, and components of the system. They offer insights into the
underlying mechanisms of the model, such as the types of neural network layers used, attention
mechanisms, and pre-processing steps. Understanding the model architecture empowers users
to assess the strengths, limitations, and potential biases of the system.</p>
        <p>Users should also have the right to remove their personal data from the training corpus,
respecting their privacy and building trust. Accommodating data removal requests is essential
for ethical practices.</p>
        <p>Additionally, detailed algorithm documentation empowers users with a deeper
understanding of decision-making, data processing, and output generation. This knowledge helps
users assess the system's reliability, consistency, and potential vulnerabilities.</p>
        <p>This connection is important because if an NLP model is opaque and does not provide
explanations or justifications for its predictions, users' freedom of choice could be compromised
because they would not understand how decisions were made and what options were considered.
When users understand how a model makes decisions and operates on their data, they can assess
whether the model meets their values, expectations, and preferences. This awareness enables
users to make informed decisions about whether to trust the model.</p>
        <p>Greater transparency in NLP models helps to ensure that users could make informed and
well-informed choices by having a better understanding of how the model works and how it
reaches its conclusions. This is especially important in critical applications where model
decisions can have a significant impact on people's lives, such as in medical, legal or public policy
decisions. In addition, greater transparency can help identify and address any bias or prejudice
in the model, promoting greater equity and inclusion in NLP applications. Additionally,
transparency facilitates the identification and mitigation of biases and discriminatory tendencies,
allowing for continuous improvement and the establishment of trust between users and
developers.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Fairness of treatment – Equality of Opportunity</title>
        <p>Despite the remarkable accuracy of AI solutions in various applications, learning algorithms
can still rely on social biases encoded in training data to make predictions. This can happen even
when information about gender and ethnicity is not explicitly provided to the system. Therefore,
machine learning algorithms potentially risk encouraging unfair and discriminatory decision
making. Fairness of treatment, in the context of a natural language processing system, refers to
the system's ability to treat the users involved in the interaction fairly and impartially. Indeed,
the system should not discriminate against or unfairly favour any individual because of personal
characteristics such as gender, ethnicity, age, or sexual orientation. This value relates directly to
the fundamental principle of equality of opportunity and the prohibition of discrimination. An
NLP model that promotes fair treatment will have to provide homogeneous and impartial
responses, regardless of their individual user characteristics. The system that meets this
parameter must avoid stereotyping, bias, or discriminatory treatment in processing users'
requests.</p>
        <p>
          Measuring fairness of treatment in an NLP system can be problematic, involving analysing
training data, identifying any biases embedded in the model and implementing techniques to
mitigate them. Careful consideration must be given to data input, evaluation metrics, and
mitigation strategies to ensure that the system maintains fair and non-discriminatory treatment
of all users. Assessing bias is critical to better understanding and addressing unfairness in NLP
models. This is often done through equity metrics [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] which quantify differences in the
behaviour of a model across a range of demographic groups. Including the fair treatment
parameter in an NLP system is intended to ensure that all people have equal access, opportunity
and fair treatment while using the system, to incentivize an inclusive and unbiased experience
for users.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Accessibility – Human Dignity</title>
        <p>The purpose of accessibility is to remove any barriers that might prevent certain people from
accessing information and interacting with the system efficiently. In the context of an NLP system,
the system should be designed to be usable by people with visual or hearing disabilities. This may
require the implementation of text-to-speech capabilities or support for reading text by assistive
devices.</p>
        <p>Also, the system ought to be designed in a way that allows people with motor or speech
disabilities to interact with it effectively (such as using voice commands or alternative interfaces).
Human dignity is considered as a fundamental principle that recognizes the inherent worth and
dignity of every individual, regardless of their personal characteristics.</p>
        <p>An accessible NLP system respects human dignity by ensuring that all people have an equal
opportunity to access and interact with the system. If an NLP system is not accessible, it may
discriminately treat some people or exclude them from accessing information or services
provided by the system. This can be harmful to the dignity of the individuals involved, resulting
in a sense of marginalization, discrimination, or disadvantage.</p>
        <p>
          While training multimodal architectures can be costly due to the need of balanced data
between modalities, text based pretrained models can be extended for accessibility.
Parameterefficient transfer learning techniques can be employed to adapt pretrained models to specific
accessibility requirements. Fine-tuning models on specialized datasets that capture the nuances
of accessibility-related interactions can improve the system's performance and responsiveness
to diverse user needs. Pretrained models that support different modalities can be chained, such
as Whisper [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], a speech-to-text model to a text-to-text model, to extend their capabilities.
        </p>
        <p>Accessibility is also connected to the value of social inclusion as the main aim of social
inclusion is to guarantee that all people, regardless of their differences and skills, have the chance
to participate fully and meaningfully in society. This encompasses access to resources, services,
opportunities and basic rights. Accessibility is a crucial aspect in the context of NLP systems;
proper attention to accessibility can ensure that the system can be used by a wide range of users
and that no one is marginalized or disadvantaged because of limitations or disabilities.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Equity in access – Social Inclusion</title>
        <p>Closely related to the value of accessibility, we consider equity in access. By this expression,
we refer to the principle that ensures equal opportunity to use NLP models, regardless of the
differences, abilities, or socio-cultural backgrounds of users.</p>
        <p>To promote equity in access, the system should be designed and implemented in such a way
that it is accessible to people with different levels of technological competence or language
proficiency. Another important aspect of equity in access is to address inequalities in access
caused by socioeconomic or geographic factors, to ensure that people from disadvantaged groups
or marginalized communities also could access NLP systems. As social inclusion refers to the
process of engaging and participating all people in society, equitable access to NLP systems is a
key element in fostering social inclusion. When NLP systems are designed with equity of access
in mind, barriers that may marginalize certain individuals or groups are reduced, thus promoting
a more inclusive social environment and technology space.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Long-term utility – Sustainability</title>
        <p>In this context, the principle of sustainability ties in with justice insofar as it aims to ensure
equitable access and benefits for all stakeholders, both in the present and in future
generations.NLP models can indeed be effective and relevant over time without becoming
obsolete or less useful. To be sustainable over the long term, they must be able to maintain their
accuracy, understanding, and language-generating capacity even in the face of new scenarios and
contexts (such as evolving languages, changes in user needs).</p>
        <p>
          NLP models, to meet this value, should be designed with their possible flexibility and
adaptability to context in mind, thus avoiding model obsolescence. For example, considering
adaptation to change in long-term utility for NLP systems requires a strategic and flexible
approach throughout the life cycle of the system [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. It is essential to continuously monitor the
operation of NLP systems and evaluate performance against relevant metrics. This enables early
identification of problems or changes in the context and user needs.
        </p>
        <p>To ensure equity over the long term, it is important to proactively explore the system and
continuously assess its relevance and impact on different populations. This could include close
monitoring of the model's performance on different demographic groups and early identification
of any disparities or biases in the model's operation. If the model shows unequal or
discriminatory performance on certain groups, corrective measures should be taken to improve
its fairness. Thus, long-term usefulness is an important aspect of sustainability in NLP model, as
ensuring the long-term usefulness of NLP models means developing models that can both remain
relevant and helpful over time, thus allowing equitable access to advanced language technologies
to be preserved for future generations.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>The aim of this position paper was to propose a possible theoretical basis for the development
of a system for evaluating the degree of justice of NLP models. Here, we have brought attention
to the need to assess the fairness of these algorithms more reliably by analyzing what could be
the core values in evaluating their adherence to the philosophical principles of Justice. This
emphasizes the importance of fuzzy evaluation based on ethical and moral principles in AI
development.</p>
      <p>This work is part of the Human Centered Artificial Intelligence Master (HCAIM) program,
sponsored by the University of Naples Federico II and in collaboration with the National Research
Council (CNR) of Naples. The HCAIM program involves four prominent European universities and
specialists from various fields to provide students with a comprehensive and interdisciplinary
education on human-centered AI. The research aims to create an evaluation system that can
classify AI solutions' adherence to philosophical principles of Justice. It seeks to develop socially
responsible, technically sound, and ethically acceptable NLP models by considering values like
freedom, human dignity, equal opportunity, social inclusion, and sustainability. The proposed
evaluation framework provides a basis for assessing the Justice of these models and guiding their
improvement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Rawls</surname>
          </string-name>
          , John.
          <source>A Theory of Justice</source>
          . Belknap Press of Harvard University Press, Cambridge, Mass.,
          <source>1971. ISBN 978-0-674-04260-5</source>
          . http://site.ebrary.com/id/10318418.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Procaccia</surname>
          </string-name>
          , Ariel.
          <source>AI Researchers Are Pushing Bias Out of Algorithms</source>
          . Bloomberg Opinion,
          <year>March 2019</year>
          . https://www.bloomberg.com/opinion/articles/2019-03-07/
          <article-title>ai-researchersare-pushing-bias-out-of-algorithms.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Hoffmann</surname>
            ,
            <given-names>Anna</given-names>
          </string-name>
          <string-name>
            <surname>Lauren</surname>
          </string-name>
          .
          <article-title>Where Fairness Fails: Data, Algorithms, and the Limits of Antidiscrimination Discourse</article-title>
          . Information,
          <source>Communication &amp; Society</source>
          ,
          <volume>22</volume>
          (
          <issue>7</issue>
          ):
          <fpage>900</fpage>
          -
          <lpage>915</lpage>
          ,
          <year>2019</year>
          . https://doi.org/10.1080/ 1369118X.
          <year>2019</year>
          .
          <volume>1573912</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Lundgard</surname>
            ,
            <given-names>Alan.</given-names>
          </string-name>
          <article-title>Measuring justice in machine learning</article-title>
          .
          <year>2020</year>
          . DOI.org (Datacite), https://doi.org/10.48550/ARXIV.
          <year>2009</year>
          .
          <volume>10050</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Binns</surname>
          </string-name>
          , Reuben.
          <source>Fairness in Machine Learning: Lessons from Political Philosophy. page 11</source>
          ,
          <year>2018</year>
          . http://proceedings.mlr.press/v81/binns18a.html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Hardt</surname>
          </string-name>
          , Moritz, Eric Price, ecprice, and Nati Srebro.
          <article-title>Equality of Opportunity in Supervised Learning</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          <volume>29</volume>
          , pages
          <fpage>3315</fpage>
          -
          <lpage>3323</lpage>
          ,
          <year>2016</year>
          . http://papers.nips. cc/paper/6374-equality
          <article-title>-of-opportunity-in-supervised-learning</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I</given-names>
            <surname>Hashimoto</surname>
          </string-name>
          , Tatsunori, Megha Srivastava, Hongseok Namkoong, and
          <string-name>
            <given-names>Percy</given-names>
            <surname>Liang. Fairness Without</surname>
          </string-name>
          <article-title>De mographics in Repeated Loss Minimization</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          , pages
          <fpage>1929</fpage>
          -
          <lpage>1938</lpage>
          ,
          <year>July 2018</year>
          . http://proceedings.mlr.press/v80/hashimoto18a.html.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Sen</surname>
          </string-name>
          , Amartya. Commodities and
          <string-name>
            <surname>Capabilities.</surname>
          </string-name>
          North-Holland, Amsterdam,
          <year>1985</year>
          . https://scholar. harvard.edu/sen/publications/commodities-and
          <article-title>-capabilities.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Alkire</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Deneulin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2009</year>
          , “
          <article-title>The Human Development and Capability Approach”</article-title>
          , in Shahani and Deneulin (eds.),
          <article-title>An Introduction to the Human Development and Capability Approach: Freedom and Agency</article-title>
          , London: Earthscan.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Santosh</surname>
            ,
            <given-names>K. C.</given-names>
          </string-name>
          , e Casey Wall.
          <source>AI</source>
          ,
          <string-name>
            <surname>Ethical</surname>
            <given-names>Issues</given-names>
          </string-name>
          and Explainability --
          <source>Applied Biometrics</source>
          . Springer,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Paula</surname>
            <given-names>Czarnowska</given-names>
          </string-name>
          , Yogarshi Vyas, Kashif Shah;
          <article-title>Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <year>2021</year>
          ;
          <volume>9</volume>
          <fpage>1249</fpage>
          -
          <lpage>1267</lpage>
          . doi: https://doi.org/10.1162/tacl_a_
          <fpage>00425</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alec</surname>
          </string-name>
          , et al.
          <article-title>"Robust speech recognition via large-scale weak supervision</article-title>
          .
          <source>" arXiv preprint arXiv:2212.04356</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Huiting</surname>
          </string-name>
          , et al.
          <source>Model Stability with Continuous Data Updates</source>
          .
          <year>2022</year>
          . DOI.org (Datacite), https://doi.org/10.48550/ARXIV.2201.05692.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>