<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>model for Loan Amount Prediction and Distribution</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hitesh K. Sharma</string-name>
          <email>hkshitesh@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tanupriya Choudhury</string-name>
          <email>tanupriya1986@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prashant Ahlawat</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sachi N. Mohanty</string-name>
          <email>sachinandan09@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarika</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haryana</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India.</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>!Department of Computer Science, Singidunum University, Serbia and and School of Computer Science &amp;</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Application, National Institute of Technology</institution>
          ,
          <addr-line>Kurukshetra,136119</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept. of Computer Science, Chandigarh University</institution>
          ,
          <addr-line>Punjab 140413</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Engineering, VIT-AP University</institution>
          ,
          <addr-line>Amaravati, Andhra Pradesh</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Machine Learning</institution>
          ,
          <addr-line>Decision Tree, Logistic Regression, Random Forest, Loan, Mine, Train</addr-line>
        </aff>
      </contrib-group>
      <fpage>210</fpage>
      <lpage>219</lpage>
      <abstract>
        <p>With the enhancement of the technology, expand of businesses and thoughts more and more people are applying for the loans. Both for their personal use and for business use. But due to the limited amount of assets the bank cannot grant loans to each and every person. Finding out those right people is a typical and time-consuming process. Banks desire to deliver the loan to an individual who can be recompensate the loan on time and can afford maximum profit to the bank. So there is a need of a system which could do this analysis and save the banks time and resources. This can be done using Machine Learning. The objective of this paper is to create a more accurate loan prediction model using machine learning to reduce the risk behind selecting of appropriate people for the loan. For this we'll mine the previous records of the people to whom the bank has granted loan. Using these records variables and bank loan rules we will train a machine learning model which will predict that a person is eligible for loan or not. We will use sklearn for our model and train_test_split for splitting the data set into train dataset and test dataset. Here we are going to use various models like Logistic Regression, Decision Tree(DT) and Random Forest(RF) so as to fetch more accurate results as the given problem is a supervised classification problem.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Prediction. Machine Learning, Decision Tree, Logistic Regression, Random Forest, Loan, Mine, Train,</title>
      <sec id="sec-1-1">
        <title>1. Introduction</title>
        <p>Maximum profit of the bank comes from lending loans to the people. So distribution of loans is a
very important part for every bank. Loan needs to be given to the right person otherwise the bank can
face financial trouble and lack of profits. Banks aim to invest their assets in safe hands and from where
they will get maximum interest. There are various factors that banks investigate before lending a loan
EMAIL:
(A.1);
(A.2);
(A.3);</p>
        <p>2020 Copyright for this paper by its authors.
such as whether a person will be able to repay the loan, what is his financial condition, for what purpose
he/she wants to have the loan, etc. Many banks undergo regress processes for choosing a right applicant
for loan even though there is still no assurance that the applicant who is selected is the deserving
applicant from all others candidates. This wastes the bank as well as the customers time. Through this
system we will be able to predict whether the applicant is safe to lend a loan or not to a level of accuracy
using the machine learning technique. As the whole process is automated no one will be able to alter
the results, it will save the banks and customers time and will lead to quick paperwork. Instead of
waiting for a few days to get the whole process done. This machine learning model will really help bank
employees and the customers. There are various factors on which this system makes the prediction, but
sometimes in real life only one strong factor could be enough for granting the loan to the person. The
data of previous records of customers will be mined so as to train our model. Data mining is the process
to extract useful information from the large dataset. Classification, clustering and association are the
types of data mining. Classification is the main type. There are various classification techniques such
as decision tree, neural network, , support vector machine and logistic regression, k-nearest neighbour
etc. In machine learning we need two types of datasets: training dataset which will be used to train the
model and other is test dataset which will be used to test the model. These both datasets are taken from
one dataset. In this paper we will use the train_test_split function of model_selection which will split
the data into training dataset and test dataset. Tree representation is used to solve the problem in which
each and every leaf node represents a class label and internal nodes represent the attributes. Any boolean
function on discrete attributes can be represented using a decision tree. Random forest, which also fits
to the supervised learning category. Both classification and regression problems can be solved by this.
The concept of ensemble learning is used, which in turn is the process of combining multiple classifiers
in order to solve a complex problem and also to progress the performance. It basically produces decision
tree set from a random subset selected from the training set and then gathers the votes from different
decision trees to make the final prediction. Logistic Regression, it is a supervised classification
Algorithm. Logical regression forecasts the categorical dependent variables. Therefore, the outcome
should be a categorical or discrete value. It is much similar to Linear regression. The core objective of
this paper is to create a less complex system for prediction of loan model. This model has been
implemented in python language by using Google Colab software, pandas library for data manipulation
and seaborn library for data visualization.</p>
      </sec>
      <sec id="sec-1-2">
        <title>2. Literature Review</title>
        <p>
          We studied various research papers on loan prediction models. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] A research paper by G. Arutjothi
and Dr. C. Senthamaria explained that predicting credit defaulters is a complex task so there is a need
for a machine learning model for this to save time and resources. Using the R software they have
proposed that the combination of Min-Max normalization and K- Nearest Neighbor (K-NN) classifier
will be good to accurately predict loan approvals. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] Aboobyda and Tagir from University of
Khartoum, Sudan used j48, bayesNet and naiveBayes algorithms for loan prediction and concluded that
j48 would be best for the accurate prediction of credit approvals. They have used Weka application for
implementation and testing of the model and compared the results of all the three algorithms.[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] A
research paper by Kumar Arun, Garg Ishan and Kaur Sanmeet explained the use of various classification
models such as Random Forest, SVM, LM, Nnet and ADB for the prediction of the loan approvals. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
Glorfeld and Hardgrave had projected a high-performance model with optimum design for the use of
neural network thus estimating the credit value of the applications of loans. 75% of the loan applicants
were correctly predicted by their designed model. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] Andy Liaw and Matthew Wiener in their research
paper told about the classification and regression by Random Forest. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] Stephan Dreiseitl and Lucila
Ohno-Machado told about the artificial neural network classification and logistic regression models,
How logical regression is useful in building various systems for predictions. Machine learning
predictive models are also a good choice for loan predictions in this market scenario[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ][
          <xref ref-type="bibr" rid="ref8">8</xref>
          ][
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>3. Proposed Model</title>
      </sec>
      <sec id="sec-1-4">
        <title>3.1.1. Data analysis</title>
        <p>Firstly, we will be doing exploratory data analysis then preprocessing and then finally we will be
testing different machine learning models on this data. The data set (Fig. 1) that we are using consists
of the following columns.</p>
        <p>We will import (Fig. 2) certain important libraries into the code such as seaborne for the visualization
purpose and pandas library for data manipulation purpose. We will also be loading the dataset in the
code. Using the head function (Fig. 3,4) we can see a few rows of our dataset.
Let’s see the data types of all the columns of our train dataset (Fig. 5).</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>As shown above there are 3 data types:</title>
    </sec>
    <sec id="sec-3">
      <title>1. object: This means that the variables are categorical</title>
      <p>2. int64: This shows the integer variables
3. float64: This data type shows variables have decimal values</p>
      <p>Now let's study about the distribution of numerical variables. Let's see the applicant income and loan
amount. We will be doing this with the help seaborn's visualisation.</p>
      <p>By seeing the graph (Fig. 6) we can say that distribution is suddenly changing its position and also
has few deviations. This may be due to the missing values in the dataset. So, we can drop those missing
values and again plot the graph with loan amount. We can achieve this with the dropna function (Fig.
7).</p>
    </sec>
    <sec id="sec-4">
      <title>Now let’s see for the Co-applicant income distribution (Fig 8).</title>
      <p>It looks similar to the applicant income distribution. As people who are more educated should be
having higher income than people who are not. So let's plot the graph with education level and income
(Fig 9).</p>
      <p>From the above graph we can see that the graduate people are having more deviation which shows
that the people with high income are well educated (Fig. 9). Loan history can also be another interesting
variable which can affect the loan prediction. We can calculate the mean of each value of loan history
by turning loan status to 0 or 1. Values closer to 1 will indicate the high loan success rates.</p>
      <p>The results tell that the loan history variable will play an important role in the loan prediction in our
model (Fig. 10).</p>
      <sec id="sec-4-1">
        <title>3.1.2. Missing data values &amp; processing the data:</title>
        <p>The dataset that we are using may or may not have missing value. But from the above graphs we
can say that our dataset is having some missing values. Let's check how many missing values (Fig. 11)
are there for each variable in our dataset.</p>
        <p>One solution for categorical values is that we can fill the missing values (Fig. 12,13) with the mode,
which means filling the values with the highest frequency. For the numerical values we can use any
mean or median, but as we have seen above there are outliers so using medium will be a proper approach
in fill all the missing numerical values (Fig. 14).</p>
        <p>Now let's work on the deviations. We can remove them by log transformation where we will nullify
their effect. We will also combine applicant income and co-applicant income to total income column.</p>
        <p>The above graph (Fig. 15) is a histogram of total income log and it seems to be much closer to
normal distribution. We will divide our dataset in dependent and independent variables. X will show
independent variables and y will show all the dependent variables (Fig. 16).</p>
        <p>Now we will split our dataset into train dataset and test dataset using the train_test_split (Fig. 17).</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.1.3. Creating the model &amp; Testing:</title>
        <p>For creating our model we will be using sklearn. But before creating, we will need to change all the
categorical variables to numbers. We can do this easily by using LabelEncoder which is present in
sklearn.</p>
        <p>All the categorical variables are now turned into numbers (Fig. 18). So, now we will scale our data
as it improves our prediction.Now we will test different classification models for accuracy. First model
we will test is the Decision tree (Fig. 19).</p>
        <p>This model has a greater accuracy than the decision tree. Now let’s test for Logistic Regression model
(Fig. 21).</p>
        <p>Accuracy for logistic regression is higher than the DT algorithm and RF algorithm. It is a good model
to be used for loan prediction.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4. Conclusion</title>
        <p>All in all, in this paper the three models that is Logistic Regression, Decision Tree and Random Forest
were applied so as to build loan prediction models that will forecast the loan approval status of applicants
as Yes or No. After training and testing all the three models with train dataset and test dataset we had the
accuracy values on the basis of which we came to a conclusion that Logistic Regression algorithm will be
accurate in predicting loan as it had a high accuracy value. Using this algorithm we will be able to predict
the right applicants for the loan approval with proper accuracy.</p>
      </sec>
      <sec id="sec-4-4">
        <title>5. References</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Arutjothi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dr. C.</given-names>
            <surname>Senthamaria</surname>
          </string-name>
          . “
          <article-title>Prediction of Loan Status in Commercial Bank using Machine Learning Classifier”</article-title>
          ,
          <source>International Conference on Intelligent Sustainable Systems(ICISS</source>
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1109/iss1.
          <year>2017</year>
          .
          <volume>8389442</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Aboobyda</given-names>
            <surname>Jafar</surname>
          </string-name>
          <article-title>Hamid and Tarig Mohammed Ahmed. “DEVELOPING PREDICTION MODEL OF LOAN RISK IN BANKS USING DATA MINING”</article-title>
          ,
          <source>Machine Learning and Applications: An International Journal (MLAIJ)</source>
          Vol.
          <volume>3</volume>
          , No.1,
          <string-name>
            <surname>March</surname>
          </string-name>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Kumar</given-names>
            <surname>Arun</surname>
          </string-name>
          , Garg Ishan, Kaur Sanmeet. “
          <source>Loan Approval Prediction based on Machine Learning Approach'', National Conference on Recent Trends in Computer Science and Information Technology (NCERT CSIT-2016).</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Louis</surname>
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Glorfeld</surname>
            and
            <given-names>Bill C.Hardgrave.</given-names>
          </string-name>
          “
          <article-title>An improved method for developing neural networks: The case of evaluating commercial loan creditworthiness”</article-title>
          ,
          <source>Computers &amp; Operations Research</source>
          , Volume
          <volume>23</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>10</given-names>
          </string-name>
          ,
          <year>October 1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Andy</given-names>
            <surname>Liaw</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Wiener</surname>
          </string-name>
          . “
          <article-title>Classification and Regression by randomForest”</article-title>
          ,
          <source>ISSN 1609-3631</source>
          , Vol.
          <volume>2</volume>
          /3,December2002.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Stephan</given-names>
            <surname>Dreiseitl and Lucila</surname>
          </string-name>
          Ohno-Machado.
          <article-title>“Logistic regression and artificial neural network classification models: a methodology review”</article-title>
          ,
          <source>Journal of Biomedical Informatics</source>
          , Volume
          <volume>35</volume>
          ,
          <string-name>
            <surname>Issues</surname>
          </string-name>
          5-6,
          <year>October 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , G. Dangi,
          <string-name>
            <given-names>T. P.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chauhan</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          ,
          <article-title>"An Efficient Way to Detect Credit Card Fraud Using Machine Learning Methodologies,"</article-title>
          <source>2018 Second International Conference on Green Computing and Internet of Things (ICGCIoT)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>591</fpage>
          -
          <lpage>597</lpage>
          , doi: 10.1109/ICGCIoT.
          <year>2018</year>
          .
          <volume>8753077</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Taneja</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Tarun Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>The Machine Predicted Market," 2018 International Conference on Computational Techniques, Electronics and Mechanical Systems (CTEMS)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>256</fpage>
          -
          <lpage>260</lpage>
          , doi: 10.1109/CTEMS.
          <year>2018</year>
          .
          <volume>8769306</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>H.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toe</surname>
            ,
            <given-names>T.T.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Machine Learning Based Predictive Analytics: A Use Case in Insurance Sector</article-title>
          . In: Jeyanthi,
          <string-name>
            <given-names>P.M.</given-names>
            ,
            <surname>Choudhury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Hack-Polay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.P.</given-names>
            ,
            <surname>Abujar</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>(eds) Decision Intelligence Analytics and the Implementation of Strategic Business Management</article-title>
          . EAI/Springer Innovations in Communication and Computing. Springer, Cham. https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -82763-2_
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>