<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AI in Finance: AI Model Validation Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Seçil Arslan</string-name>
          <email>secil.arslan@prometeia.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>(Associate Partner)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Buğra Akyüz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>(Manager)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff8">8</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Artificial Intelligence</institution>
          ,
          <addr-line>Model Validation, Finance, Interpretability, Bias, Robustness, Fairness</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gradient Boosting Algorithms</institution>
          ,
          <addr-line>and Neural Networks</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Learning</institution>
          ,
          <addr-line>Deep Learning, Natural Language Processing</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Operational Risk</institution>
          ,
          <addr-line>Natural Language Processing methods</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Prometeia</institution>
          ,
          <addr-line>River Plaza, Kat 19, Büyükdere Caddesi Bahar Sokak No. 13, 34394, Istanbul/</addr-line>
          <country country="TR">TURKEY</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Prometeia</institution>
          ,
          <addr-line>River Plaza, Kat 19, Büyükdere Caddesi Bahar Sokak No. 13, 34394, Istanbul/</addr-line>
          <country country="TR">TURKEY</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>have started to replace models such as Logistic Regres-</institution>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>of banks. For example; Decision Trees</institution>
          ,
          <addr-line>Random Forests</addr-line>
        </aff>
        <aff id="aff8">
          <label>8</label>
          <institution>sion as far as Credit Risk is concerned. In the domain of</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>Validation of Artificial Intelligence (AI) Models in the finance sector has been one of the most crucial phases of the AI models' life cycle. Although the finance sector is highly regulated and already familiar with validating traditional statistical methods in credit risk, they need an extension and adaptation to their as-is validation standards and frameworks for advanced AI algorithms. The extension is not only limited to credit risk but can also apply to divergent business domains. This paper highlights the risks of using AI in finance applications and provides significant motivations for having an AI validation framework to control and eliminate those risks. Besides, we underline the details of our framework's pillars by mapping them to well-known validation contexts like conceptual soundness, model performance, and model usage.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction
applications.
framework.
cessibility to powerful processing units, the applications
of Artificial Intelligence (AI) have been increasing
tremendously in the finance sector. Although the finance sector
gies, it is still a blue ocean to use Machine Learning (ML)
or AI-based systems and trust them in mission-critical</p>
    </sec>
    <sec id="sec-2">
      <title>It is crucial to recall the definition and life cycle of</title>
    </sec>
    <sec id="sec-3">
      <title>Artificial Intelligence (AI) systems, according to OECD,</title>
      <p>is defined up to some parameters, and ”learning” is the
execution of a computer program to optimize the model’s
dictive to make predictions in the future, descriptive to
gain knowledge from data or both.</p>
    </sec>
    <sec id="sec-4">
      <title>Advanced AI approaches difer from traditional (sta</title>
      <p>Regression. These traditional models are designed to
make inferences about the relations between variables,
following models and variables defined by human
experts. These models can make reliable predictions, yet
models, they are designed to make the most accurate
predictions possible, as well as other inferences working on
a data set and similar new data. They might sacrifice
internative data sources [1]. They use machine and/or
human-based data and inputs to (i) perceive real and/or
virtual environments; (ii) abstract these perceptions into
models through analysis in an automated manner (e.g.,
with machine learning), or manually; and (iii) use model
inference to formulate options for outcomes [1].</p>
    </sec>
    <sec id="sec-5">
      <title>In another definition by [ 2]: ”Machine Learning is programming computers to optimize a performance criterion using example data or past experience”. A model</title>
      <p>the development of Chatbots and conversational
interfaces for direct communication with clients. As far as
ifnancial fraud is concerned, AI approaches have
signifiof data samples, and learning from past data experience is
the main driver of those algorithms. Also, unlike classical
programming approaches, there is not a 100% expected
outcome precision in those approaches. Although Banks
are very familiar with model outputs that reflect a
predictive approach, traditional well-known methods like
Linear Regression and Logistic Regression are far more
diferent than advanced ML techniques applied
nowadays. Today, most of the ML approaches are more
blackbox and they require careful examination to create trust
in robustness, fairness, data privacy, and bias concerns.</p>
      <p>Figure 1: Model Life Cycle In addition, unlike the linear methods, new algorithms
require new methods to provide feature interpretability
and model explainability like SHAPley or LIME [1].
tion of both credit card fraud and application fraud or The major risks of ML/AI models are about the
responAnti-Money Laundering owing to their capacity to model sibility and accountability of the models. The discussion
complex patterns within the data. is on who is accountable for unfair, biased results of a</p>
      <p>A typical life cycle of an ML/AI model has the fol- model; the historical data including biased information
lowing steps: (i) planning and design, data collection or the model developer not taking the necessary
precauand processing, and model building and interpretation; tions, or the validation team not detecting the possible
(ii) verification and validation; (iii) deployment; and (iv) bias and fairness weaknesses.
operation and monitoring [1] (Figure 1). Another risk is related to the typical problem of ML/AI</p>
      <p>This paper aims to provide a full-fledged AI model vali- models where they seem to perform very well in
labodation framework that serves as a guideline to one of the ratory environments and cannot reflect the same
percritical life cycle phases of ML/AI models in the finance formance in production environments. This may result
sector. Our approach discusses in detail the controls re- from various reasons; one is that the data distribution or
garding conceptual soundness (model documentation, quality patterns may difer in production compared to
data validation, model design) and model performance the training data or the model may have overfitted on
suitability required for the entire life cycle process of an the training phase and no one has detected it.
AI model as well as the controls necessary for the model All these risk factors afect the trust of ML/AI within
usage in finance systems (production environment, us- institutions and compliance with legislation standards.
age and controls, monitoring approach). Those three con- Divergent applications including back, middle, and/or
cepts of validation framework refer to the 4 main pillars front-ofice related to credit, asset management, or even
to be controlled and validated: Data, Methodology, Pro- algorithmic trading [1] are at the core of those risks, and
cess, and Governance. Our framework aims to underline validation of these models has become the key point in
how to eliminate the risks of AI in finance applications managing the risks.
by providing guides to validate models in terms of data
bias, quality, and privacy issues; robustness and fair- 3. Motivation
ness of algorithms; preventing and detecting overfitting
or underfitting performances and interpretability of
ML/AI models and features.</p>
    </sec>
    <sec id="sec-6">
      <title>In order to alleviate the risks that have already been</title>
      <p>mentioned, companies need a standardized guideline for
validating AI models. Our main motivations for creating
2. Risks of AI in Finance a validation framework are to (i) create trust for AI, (ii)
guarantee compliance with legislation frameworks, and
With the increasing number of ML/AI applications on (iii) improve internal procedures.
credit risk, CRM/Marketing analytics, operational risk, First of all, improving the adoption of AI by creating
process automation, fraud detection, and robo-advisory; “Trust” is a significant dimension of the need for an AI
ifnancial companies need to take care of the risks of Model Validation framework. The adoption of AI in
bankadoption of those ML/AI models in their as-is processes ing and finance applications is, although limited,
increasand workflows. ing. Creating more awareness within the companies is</p>
      <p>Due to high competition among financial institutions, possible by standardization of model validation processes
many banks and insurance companies investing in appli- that can fasten the early adoption of AI. Tier-1 banks
precations of ML/AI in their core processes. However, the fer in-house developments of AI models, whereas, other
nature of ML/AI models depends highly on the selection banks may be limited in in-house capacities and prefer
vertical start-ups. Both cases require developing ”Trust”
in the adoption of AI Models. Reports underline that 44%
of models are in the pre-deployment phase and only 56%
of them are in the deployment phase [3]. One way to
improve trust is to validate the models before going live
in production environments.</p>
      <p>Secondly, it has become a good opportunity to create
awareness and the need for an internal AI procedure
(either developed or COTs services) that ensures AI Models
are validated and checked with respect to legislations
like EU ”AI Act” [4], or ”EBA Discussion Papers on
Machine Learning for IRB Models” [5]. This will strengthen
the arguments to convince customers of the need and
empower ownership mechanisms.</p>
      <p>Finally, although validation or audit teams of the AI
Models are the main stakeholders of the AI Model
Validation framework, they are not the only ones. The
framework serves CDOs and CTOs of Banks to overview their
available model development procedures. Risk awareness
will be handled with a standard validation approach so
that deployment processes will be smoother and aligned.
4. Approach</p>
      <sec id="sec-6-1">
        <title>4.1. Validation Landscape</title>
        <p>Our focus on the AI model’s validation is to question and
validate models including all steps from data creation to
deployment [6]. Each step of a typical ML/AI model is
subject to validation including data curation, training,
adaptation, and deployment. Those concepts are
represented in our framework in 4 main pillars of validation:
(i) Data, (ii) Methodology, (iii) Processes, and (iv)
Governance. Those validation steps respectively, validate
models both in quantitative and qualitative perspectives.</p>
        <p>The main actors within a model life cycle are model
developers, validation teams, and end-users of the models
(Figure 2). Each one has diferent roles and
responsibilities in the life cycle of a model. Our framework primarily
serves validation teams to support an internal defense
mechanism within the company to check and approve
the go/no-go decision of ML/AI models before going into
production. Besides, the validation team is the main
responsible to monitor regularly and repeat periodically
some of the validation steps.</p>
      </sec>
      <sec id="sec-6-2">
        <title>4.2. Validating on Three Concepts</title>
      </sec>
      <sec id="sec-6-3">
        <title>Mapped into Four Main Pillars</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Most banks are more familiar with validation approaches</title>
      <p>and guidelines since the models are highly regulated by
regulation and supervision agencies. (ECB, EBA). These
agencies provide the rules and steps to be followed for
the validation of regulative credit risk models. Since our</p>
    </sec>
    <sec id="sec-8">
      <title>AI Model Validation framework includes the validation</title>
      <p>of ML/AI models, including but not limited to, credit risk
models, we mapped the new paradigms of validation to
the as-is validation concepts that banks are already using
internally [7].</p>
      <p>Figure 3 demonstrates the mapping of our validation
pillars into three concepts of validation.</p>
      <p>Under the conceptual soundness dimension; the
quality of model design, construction, and
documentation is assessed. In addition to conventional steps like
data validation, we need to focus specifically on ML-only
steps in methodology suitability: selection of the correct depending on the rule sets defined on those trafic lights.
model, feature extraction/selection, and hyperparameter
optimization. Even in data validation, we enlarge the
typical validation phases of raw data, model data, and 5. Conclusion
target variable quality controls with privacy and bias
considerations on the selected data. Besides, the approach In this paper, we briefly discuss the definition of AI
sysof choosing the correct data splits for utilizing the steps tems and their burgeoning usage in finance applications.
of training, parameter optimization, and testing the final We emphasize the possible risks of using AI in finance and
results are questioned. underline the importance of the validation phase within</p>
      <p>In the model performance part of the validation the overall life cycle of a model. The driving factors
beframework, model outputs are compared against the out- hind preparing an end-to-end validation framework for
comes observed. In ML/AI, many diferent metrics and AI models are the need for appropriate control over them
tests can be derived to quantify results. The important and for creating trust in terms of bias, robustness, and
part is to provide guidelines to compare outcomes on the fairness of the models.
objective of models and define feature interpretability/- Furthermore, we described diferent validation types
explainability with advanced methods. Our framework and the logic of the trafic light approach for scoring
underlines the possibility of several diferent problem do- models both initially and periodically in pre-deployment,
mains and algorithms that can be under validation. Each and production environments, respectively.
and every algorithm is questioned by choosing suitable Finally, our future works will focus on converting our
performance metrics to compare results, feature selection AI model validation framework into an automated
Softand extraction methods, feature interpretability/explain- ware as a Service (SaaS) approach that is embedded into
ability, and bias/variance concepts to detect models that our Model Risk Management (MRM™) tool.
are underfitting or overfitting.</p>
      <p>The model usage phase focuses on not only validat- References
ing outcomes and soundness but also the adaptation to
processes and applications, the design of reflecting and [1] OECD, Artificial intelligence, machine learning and
digitizing processes with ML models, and also considers big data in finance: Opportunities, challenges, and
the ”human-in-the-loop” strategy for the end-users. In implications for policy makers (2021).
addition, the model governance precautions, monitor- [2] E. Alpaydın, Introduction to Machine Learning, The
ing and reporting mechanisms, and continuous-learning MIT Press, 2020.
techniques are examined. [3] B. of England, Machine learning in uk financial
services (2022).
4.3. Validation Types and Triggers [4] E. Commission, Proposal for a regulation of the
european parliament and of the council laying down
harmonised rules on artificial intelligence (artificial
intelligence act) and amending certain union
legislative acts (2021).
[5] EBA, Eba discussion paper on machine learning for</p>
      <p>irb models (2021).
[6] R. Bommasani, et al., On the
opportunities and risks of foundation models, CoRR
abs/2108.07258 (2021). URL: https://arxiv.org/abs/
2108.07258. a r X i v : 2 1 0 8 . 0 7 2 5 8 .
[7] E. C. Bank, Targeted review of internal models
(2021).</p>
      <p>Validation procedures divide into two, according to their
content and scope: Initial and Periodic. Initial validation
is the end-to-end examination of a model after it has been
developed and before it is released into the production
environment. During initial validation, all controls
under the four pillars mentioned above are completed. In
periodic validation, however, changes in the population
subject to the model and their efects on model
performance are monitored in order to monitor the health of the
model in general. Thus, it is aimed to detect models that
are aging or whose performance is seriously deteriorated.</p>
      <p>In both initial and periodic validation, the results of
the tests applied are expressed by trafic lights. The green
light indicates that the test has been passed, the red light
indicates that the test has failed, while the yellow light
indicates that the result is good enough, but can be
improved. Since the question sets used for the validation
processes consist of many questions under many
categories, the use of trafic lights is important for these
results to be clearly understood by relevant parties and
for the final result of the validation to be determined</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>