<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Lazy Classication with Interval Pattern Structures: Application to Credit Scoring</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexey Masyutin</string-name>
          <email>alexey.masyutin@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yury Kashnitsky</string-name>
          <email>ykashnitsky@hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergei Kuznetsov</string-name>
          <email>skuznetsov@hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University Higher School of Economics Scientic-Educational Laboratory for Intelligent Systems and Structural Analysis Moscow</institution>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Pattern structures allow one to approach the knowledge extraction problem in case of arbitrary object descriptions. They provide the way to apply Formal Concept Analysis (FCA) techniques to nonbinary contexts. However, in order to produce classication rules a concept lattice should be built. For non-binary contexts this procedure may take much time and resources. In order to tackle this problem, we introduce a modication of the lazy associative classication algorithm and apply it to credit scoring. The resulting quality of classication is compared to existing methods adopted in bank systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Banks and credit institutions face classication problem each time they
consider a loan application. In the most general case, a bank aims to have a tool to
discriminate between solvent and potentially delinquent borrowers, i.e. the tool
to predict whether the applicant is going to meet his or her obligations or not.
Before 1950s such a decision was expert driven and involved no explicit
statistical modeling. The decision whether to grant a loan or not was made upon an
interview and after retrieving information about spouse and close relatives [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
From the 1960s, banks have started to adopt statistical scoring systems that were
trained on datasets of applicants, consisting of their socio-demographic factors
and loan application features. As far as mathematical models are concerned, they
were typically logistic regressions run on selected set of attributes. Apparently,
a considerable amount of research was done in the eld of alternative machine
learning techniques seeking the goal to improve the results of the wide-spread
scorecards [
        <xref ref-type="bibr" rid="ref10 ref11 ref7 ref8 ref9">7,8,9,10,11</xref>
        ].
      </p>
      <p>
        All mentioned methods can be divided into two groups: the rst one provides
the result dicult for interpretation, so-called black box models, the second
group provides interpretable results and clear model structure. The key feature
of risk management practice is that, regardless of the model accuracy, it must
not be the black box. That is why methods such as neural networks and SVM
classiers did not earn much trust within the banking community [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The
dividing hyperplane in an articial high-dimensional space (dependent on the chosen
kernel) cannot be easily interpreted in order to claim the reject reason for the
client. As far as neural networks are concerned, they also do not provide the
user with a set of reasons why a particular loan application has been approved
or rejected. In other words, these algorithms do not provide the decision maker
with knowledge. The predicted class is generated, but no knowledge is retrieved
from data.
      </p>
      <p>
        On the contrary, alternative methods such as association rules and decision
trees provide the user with easily interpretable rules which can be applied to the
loan application. FCA-based algorithms also belong to the second group since
they use concepts in order to classify objects. The intent of the concept can be
interpreted as a set of rules that is supported by the extent of the concept.
However, for non-binary context the computation of the concepts and their relations
can be very time-consuming. In case of credit scoring we deal with numerical
context, as soon as categorical variables can be transformed into a set of dummy
variables. Lazy classication [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] seems to be appropriate to use in this case
since it provides the decision maker with the set of rules for the loan application
and can be easily parallelized. In this paper, we modify the lazy classication
framework and test it on credit scoring data of a top-10 Russian bank.
      </p>
      <p>The paper is structured as follows: section 2 provides basic denitions. Section
3 argues why the original setting can be inconsistent in case of a large numerical
context and describes the proposed modication and its parameters. Section
4 describes voting schemes that can be used to classify test objects. Section
5 describes the data in hand and some experiments with parameters of the
algorithm. Finally, section 6 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Main Denitions</title>
      <p>
        First, we recall some standard denitions related to Formal Concept Analysis,
see e.g. [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ].
      </p>
      <p>
        Let G be a set (of objects), let (D, u) be a meet-semi-lattice (of all possible
object descriptions) and let : G ! D be a mapping. Then (G, D , ), where
D =(D, u), is called a pattern structure [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], provided that the set
(G) := { (g)|g 2 G} generates a complete subsemilattice ( D , u) of (D, u), i.e.,
every subset X of (G) has an inmum uX in (D, u). Elements of D are called
patterns and are naturally ordered by subsumption relation v:
given c, d 2 D one has c v d $ c u d = c. Operation u is also called a similarity
operation. A pattern structure (G; D; ) gives rise to the following derivation
operators ( ) :
      </p>
      <p>A = l (g)
g2A</p>
      <p>for A 2 G;
d = fg 2 G j d v (g)g
for d 2 (D; u):</p>
      <p>
        These operators form a Galois connection between the powerset of G and
(D; u). The pairs (A; d) satisfying A G, d 2 D, A = d, and A = d are called
pattern concepts of (G,D, ), with pattern extent A and pattern intent d.
Operator ( ) is an algebraical closure operator on patterns, since it is idempotent,
extensive, and monotone [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The concept-based learning model for standard object-attribute
representation (i.e., formal contexts) is naturally extended to pattern structures. Suppose
we have a set of positive examples G+ and a set of negative examples G w.r.t.
a target attribute, G+ \ G = ;, objects from
G = G n(G+ [ G ) are called undetermined examples. A pattern c 2 D is an
- weak positive premise (classier) i:
jjc \ G jj</p>
      <p>jjG jj
jjh \ G jj
jjG jj
and 9A</p>
      <p>G+ : c v A
and 9A</p>
      <p>G+ : h = A
A pattern h 2 D is an</p>
      <p>- weak positive hypothesis i:</p>
      <p>
        In case of credit scoring we work with pattern structures on intervals as
soon as a typical object-attribute data table is not binary, but has many-valued
attributes. Instead of binarizing (scaling) data, one can directly work with
manyvalued attributes by applying interval pattern structure. For two intervals [a1; b1]
and [a2; b2], with a1; b1; a2; b2 2 R the meet operation is dened as [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]:
[a1; b1] u [a2; b2] = [min(a1; a2); max(b1; b2)].
      </p>
      <p>
        The original setting for lazy classication with pattern structures can be
found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Modication of lazy classication algorithm</title>
      <p>In credit scoring the object-attribute context is typically numerical. Factors
can have arbitrary distributions and take wide range of values. At the same time
categorical variables and dummies can be present. With relatively large number
of attributes (over 30-40) it produces high-dimensional space of continuous
variables. That is when the result of the meet operator tends to be very specic, i.e.
for almost every g 2 G only g and gn have the descriptin (gn) u (g). This
happens due to the fact that numerical variables, ratios especially, can have unique
values for every object. This results in that for test object gn the number of
positive and negative premises is close to the number of observations in those
context correspondingly. In other words, too specic descriptions are usually not
falsied (i.e. there are no objects of opposite class with such description) and
almost always form either positive or negative premises. Therefore, the idea of
voting scheme for lazy classication in the case of high dimensional numerical
context may turn out to be obscure. Thus, it seems reasonable to seek the
concepts with larger extent and with not too specic intent. At the same we would
like to preserve the advantages of lazy classication, e.g. no need to compute a
full concept lattice, easy parallelization etc. The way to increase the extent of the
generated concepts is to consider intersection of the test object with more than
one element from the positive (negative) context. What is the suitable number
of objects to take for intersection? In our modication we consider this as a
parameter subsample size and perform grid search. The parameter is expressed
as percentage of the observations in the context. As subsample size grows, the
resulting intersection (g1) u : : : u (gk) u (g) becomes more generic and it is
more frequently falsied by the objects from the opposite context. Strictly
speaking, in order to replicate the lazy classication approach, one should consider all
possible combinations of the chosen number of objects from the positive
(negative) context. Apparently, this is not applicable in the case of large datasets.
For example, having 10 000 objects in positive context and having subsample
size equal to only two objects will produce almost 50 mln combinations for
intersection with the test object. Therefore, we randomly take the chosen number
of objects from positive (negative) context as candidates for intersection with
the test object. The number of times (number of iterations) we randomly pick a
subsample from the context is also tuned through grid search. Intuition says , the
higher the value of the parameter the more premises are mined from the data.
However, the obvious penalty for increasing the value of this parameter is time
and resources required for computing intersections. As mentioned before, the
greater the subsample size, the more it is likely that ( (g1) u : : : u (gk) u (g))
contains the object of the opposite class. In order to control this issue, we add a
third parameter which is alpha-threshold. If the percentage of objects from the
positive (negative) context that falsify the premise (g1) u : : : u (gk) u (g) is
greater than alpha-threshold of this context then the premise will be considered
as falsied, otherwise the premise will be supported and used in the classication
of the test object.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Voting schemes</title>
      <p>The nal classication of a test object is based on a voting scheme among
premises. In most general case voting scheme F is a mapping:</p>
      <p>F (gtest; h1+; :::; hp+; h1 ; :::; hn ) ! [ 1; 1; ;]
where gtest is the test object with unknown class, hi+ is a positive premise 8i =
1; p and hj is a negative premise 8j = 1; n , -1 is a label for negative class, and 1
is a label for positive class (i.e. defaulters). In other words, F is an aggregating
rule that takes premises as input and gives the classication label as an output.
Note that we allow for an empty label. If the label is empty it is said that the
voting rule abstains from classication. There may be dierent approaches to
build up aggregating rules. The voting scheme is built upon weighting function
!( ), aggregation operator A( ) and comparing operator .</p>
      <p>F (!( ); A( ); ) =
= (Aip=1[!(hi+)])</p>
      <p>(Ajn=1[!(hj )])</p>
      <p>In order to congure a new weighting scheme it is sucient to dene the
operators and the weighting function. In this paper we use the number of positive
versus negative premises. In this case the rule allows the test object to satisfy
both positive and negative premises which decreases the rejection from
classication. The weighting function, aggregation operator and comparing operator
are dened as follows:</p>
      <p>A(h) =</p>
      <p>X h
!(h) =
a
b =
(1; if (gtest) v h</p>
      <p>0; otherwise
(sign(b
;;
a); if a 6= b
a = b
So the label for a test object gn is dened by the following mapping:
F (gtest; h1+; :::; hp+; h1 ; :::; hn ) =</p>
      <p>p n
= (X[ (gtest) v hi+]) (X[ (gtest) v hj ])</p>
      <p>i=1 j=1</p>
      <p>However, one can think of margin b a as a measure for discrimination
between two classes and consider the decision boundary based on receiver
operating characteristic analysis, for instance. This approach is good for decreasing
the number of rejects from classication, but it does not account for the
support of the premises. Naturally, one would give more weight to the premise with
large image (with higher support). Also, if the number of positive and negative
premises is equal the rule rejects from classication.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>
        The data we used for the computation represent the customers and their
metrics assessed on the date of loan application. The applications were approved by
the bank credit policy and the clients were granted the loans. After that the
loans were observed for the fact of delinquency. The dataset is divided into two
contexts positive and negative. The positive context is the set of loans where
the target attribute is present. The target attribute in credit scoring is
typically dened as more than 90 days of delinquency within the rst 12 months
after the loan origination. So, the positive context is the set of bad borrowers,
and the negative context consists of good ones. Each context consists of 1000
objects in order that voting scheme concerned in the second section was
applicable. The test dataset consists of 300 objects and is extracted from the same
population as the positive and negative contexts. Attributes represent various
metrics such as loan amount, term, rate, payment-to-income ratio, age of the
borrower, undocumented-to-documented income, credit history metrics etc. The
set of attributes used for the lazy classication trials contained 28 numerical
attributes. In order to evaluate the accuracy of the classication we calculate the
Gini coecient for every combination of parameters based on 300 predictions on
the test set. Gini coecient is calculated based on the margin between the
number of objects within positive premises and negative ones. In fact, the margin is
the analog for the score value in credit scorecards. Gini coecient was chosen as
performance metric because it is conventionally used to evaluate the quality of
classication models in credit scoring [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. When the subsample size is low, the
intersections of the test object description and the members of positive
(negative) context tend to be more specic. That is why, a relatively high number of
premises are mined and used for the classication. As subsample size increases,
the candidates for premises start being generic and it is likely that there exists
certain amount of objects from the opposite context which also satisfy the
description. If alpha-threshold is low, the frequency of rejects from classication is
high. The dynamics of premise mining is demonstrated on the following graphs:
      </p>
      <p>The average number of premises mined for a test object is dropping as
expected with the increase in the subsample size and the drop is quicker for higher
alpha-thresholds. This supports the idea, that if lazy classication is run in its
original setting upon the numerical context (i.e. when subsample size consists
of only one object) the number of premises generated is close to the number of
objects in the context, so the premises can be considered as too specic. The
descriptive graph above allows one to expect that the proposed parameters of
the algorithm can be tuned (grid searched), so as to tackle the trade-o between
the high number of premises used for classication and the size of their support.
The average number of positive premises tends to fall slightly faster compared
to negative premises. Below we present the classication accuracy obtained for
dierent combinations of parameters (grid search).
We observe the area with zero Gini coecients where the alpha-threshold is
zero and the subsample size is relatively high. That is due to the fact that almost
no premises were mined during the lazy classication run. It is quite intuitive
because as the subsample size grows, the intersection of the subsample with a
test object results in a generic description, which is very likely to be falsied at
least by one object from the opposite context. In this case the rejection from
classication takes place almost for all test objects. The rst thing that is quite
intuitive is that the more iterations are produced, the higher is the Gini on
average:</p>
      <p>The more times the subsamples are randomly extracted the more knowledge
(in terms of premises) is generated. By increasing the number of premises used for
classication according to voting scheme, we are likely to capture the structure of
the data in more detail. However, the number of iterations is not the only driver
of the classication accuracy in our case. We nd a range with relatively high
Gini in the area of mild alpha-threshold and relatively high subsample size. It
also seems natural as soon as the support of a good predictive rule (i.e. premise)
is expected to be higher than its support in the opposite context. We elaborate
further and run additional grid search in range of parameters providing high
Gini coecient:</p>
      <p>
        According to performed grid search the range with the highest Gini
(55%56%) on the test sample is in range with following parameter values:
alphathreshold = 0,3%, number of iterations = 10000, subsample size = 1,0%. The
result was compared to three benchmarks that are traditionally used in the credit
scoring within the bank system: logistic regression, scorecard and decision tree.
It should be cleared what is implied by the scorecard classier. Mathematical
architecture of the scorecard is based on logistic regression which takes the
transformed variables as input. The transformation of the initial variables which is
typically used is weight of evidence transformation (WOE-transformation [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]).
It is wide-spreaded in credit scoring to apply such a transformation to the input
variables as soon as it accounts for non-linear dependencies and it also provides
certain robustness coping with potential outliers. The aim of the transformation
is to divide each variable into no more than k categories. The thresholds are
derived so as to maximize the information value of a variable [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Having each
variable binned into categories, the log-odds ratio is calculated for each category.
Finally, instead of initial variables the discrete valued variables are considered as
input in logistic regression. The properties of the decision tree were as follows:
we ran CART with two possible child nodes from each parent node. The
criterion for optimal threshold calculation was the greatest entropy reduction. The
number of terminal nodes was not explicitly restricted; however, the minimum
size of the terminal node was set to 50. As far as logistic regression is concerned,
the variable selection was performed based on stepwise approach [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. As for
scorecard, the variables were initially selected based on their information value
after the WOE-transformation. The comparison of the classiers performance
based on test sample of 300 objects is given in Table 3.
      </p>
      <p>
        When dealing with large numerical datasets, lazy classication may be
preferable to classication based on explicitly generated classiers, since it requires less
time and memory resources [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, the original lazy classication setting
in case of high dimensional numerical feature space meets certain limitation.
The limitation is that, when intersecting descriptions of a test object and every
object from the context, one is likely to acquire premises with image consisting
only of those two objects. In other words, the premises tend to be very specic
for the context and, therefore, the number of positive and negative premises is
likely to be equal to the number of the objects in the contexts. The weighting
cannot be considered helpful in this case as soon as the premises will have very
similar low support. In this paper, we modied the original lazy classication
setting by making it, in fact, a stochastic procedure with three parameters:
subsample size, number of iterations and alpha-threshold. In eect, the modied
algorithm mines the premises with relatively high support that will be used for
the classication of the test object. The classication is then carried out upon
the predened voting scheme. We applied the introduced procedure to the retail
loan classication problem. The data we used for was provided during the pilot
project with one of the top-10 banks in Russia, the details are not provided due
to non-disclosure agreement. The positive and negative contexts both had 1000
objects with 28 numerical attributes. The accuracy of the algorithm was
evaluated on the test dataset consisting of 300 objects. Gini coecient was chosen as
accuracy metric. We performed the basic grid search by running the modied
lazy classication algorithm with dierent parameter values. The classication
accuracy of the algorithm was compared to the conventionally adopted models
used in the bank. The benchmark models were logistic regression, scorecard and
decision tree. The proposed algorithm outperforms the logistic regression the
scorecard with the subsample size parameter around 1%, alpha-threshold equal
to 0,3% and with number of iterations over 5000. The performance of the decision
tree is at the comparable level with the proposed algorithm, however, the
modied lazy classication is slightly better in terms of Gini coecient. As an area
for further research, one can consider and compare accuracy when other voting
schemes are used. It is expected that taking into account premises’ specicity
one can improve overall accuracy of the classication algorithm or, alternatively,
one will reach the same accuracy given less number of iterations, which can save
the time resources required for the calculations.
Input: fP osdata; N egdatag positive and negative numerical contexts.
N +; N number of objects in the contexts. It is preferable that the positive and
negative contexts are of the same size.
      </p>
      <p>M number of attributes.
sub:smpl percentage of the context randomly used for intersection with the test
object (parameter).
num:iter number of iterations (resamplings) during the premise mining (parameter).
alpha:threshold is the maximum allowable percentage of the opposite context for that
the premise is not falsied (parameter).
t test object.
- weak premises
- weak negative</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Ganter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sergei</given-names>
            <surname>Kuznetsov</surname>
          </string-name>
          ,
          <article-title>Pattern structures and their projections, in Conceptual Structures: Broadening the Base, Harry Delugach</article-title>
          and Gerd Stumme, Eds., vol.
          <volume>2120</volume>
          of Lecture Notes in Computer Science, pp.
          <fpage>129142</fpage>
          . Springer, Berlin/Heidelberg,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ganter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wille</surname>
          </string-name>
          , R.:
          <source>Formal concept analysis: Mathematical foundations</source>
          . Springer, Berlin,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Sergei O. Kuznetsov</surname>
          </string-name>
          ,
          <article-title>Scalable knowledge discovery in complex data with pattern structures</article-title>
          ., in PReMI, Pradipta Maji, Ashish Ghosh,
          <string-name>
            <given-names>M. Narasimha</given-names>
            <surname>Murty</surname>
          </string-name>
          , Kuntal Ghosh, and Sankar K. Pal, Eds.
          <year>2013</year>
          , vol.
          <volume>8251</volume>
          of Lecture Notes in Computer Science, pp.
          <fpage>3039</fpage>
          , Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Thomas</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edelman</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crook</surname>
            <given-names>J</given-names>
          </string-name>
          . (
          <year>2002</year>
          )
          <article-title>Credit Scoring</article-title>
          and
          <string-name>
            <given-names>Its</given-names>
            <surname>Applications</surname>
          </string-name>
          ,
          <source>Monographs on Mathematical Modeling and Computation</source>
          , SIAM: Pliladelphia, pp.
          <fpage>107</fpage>
          <lpage>117</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bigss</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ville</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Suen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>1991</year>
          ).
          <article-title>A Method of Choosing Multiway Partitions for Classication and Decision Trees</article-title>
          .
          <source>Journal of Applied Statistics</source>
          ,
          <volume>18</volume>
          ,
          <issue>1</issue>
          ,
          <fpage>49</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Naeem</given-names>
            <surname>Siddiqi</surname>
          </string-name>
          ,
          <article-title>Credit Risk Scorecards: Developing and Implementing Intelligent Credit Scoring</article-title>
          , WILEY,
          <source>ISBN: 978-0-471-75451-0</source>
          ,
          <fpage>2005</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>B</given-names>
            <surname>Baesens</surname>
          </string-name>
          , T Van Gestel,
          <string-name>
            <given-names>S</given-names>
            <surname>Viaene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Stepanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Suykens</surname>
          </string-name>
          ,
          <article-title>Benchmarking stateof-the-art classication algorithms for credit scoring</article-title>
          ,
          <source>Journal of the Operational Research Society</source>
          <volume>54</volume>
          (
          <issue>6</issue>
          ),
          <fpage>627</fpage>
          -
          <lpage>635</lpage>
          ,
          <year>2003</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ghodselahi</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A Hybrid</given-names>
            <surname>Support</surname>
          </string-name>
          <article-title>Vector Machine Ensemble Model for Credit Scoring</article-title>
          ,
          <source>International Journal of Computer Applications (0975 8887)</source>
          , Volume
          <volume>17</volume>
          No.
          <issue>5</issue>
          ,
          <string-name>
            <surname>March</surname>
            <given-names>2011</given-names>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>K. K.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>An intelligent agent-based fuzzy group decision making model for nancial multicriteria decision support: the case of credit scoring</article-title>
          .
          <source>European journal of operational research</source>
          . vol.
          <volume>195</volume>
          . pp.
          <fpage>942</fpage>
          -
          <lpage>959</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gestel</surname>
            ,
            <given-names>T. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baesens</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suykens</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          , Van den Poel, D.,
          <string-name>
            <surname>Baestaens</surname>
          </string-name>
          , D.-E. and
          <string-name>
            <surname>Willekens</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Bayesian kernel based classication for nancial distress detection</article-title>
          .
          <source>European journal of operational research</source>
          . vol.
          <volume>172</volume>
          . pp.
          <fpage>979</fpage>
          -
          <lpage>1003</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>P.</given-names>
            <surname>Ravi Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ravi</surname>
          </string-name>
          ,
          <article-title>Bankruptcy Prediction in Banks and Firms via Statistical and Intelligent Techniques-A Review</article-title>
          ,
          <source>European Journal of Operational Research</source>
          , Vol.
          <volume>180</volume>
          , No.
          <volume>1</volume>
          ,
          <issue>2007</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Sergei</surname>
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Kuznetsov</surname>
            and
            <given-names>Mikhail V.</given-names>
          </string-name>
          <string-name>
            <surname>Samokhin</surname>
          </string-name>
          ,
          <article-title>Learning closed sets of labeled graphs for chemical applications</article-title>
          .,
          <string-name>
            <surname>in</surname>
            <given-names>ILP</given-names>
          </string-name>
          , Stefan Kramer and Bernhard Pfahringer, Eds.
          <year>2005</year>
          , vol.
          <volume>3625</volume>
          of Lecture Notes in Computer Science, pp.
          <fpage>190</fpage>
          <lpage>208</lpage>
          , Springer
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. SAS Institute Inc. (
          <year>2012</year>
          ),
          <article-title>Developing Credit Scorecards Using Credit Scoring for</article-title>
          SAS R Enterprise MinerTM
          <volume>12</volume>
          .1,
          <string-name>
            <surname>Cary</surname>
          </string-name>
          , NC: SAS Institute Inc.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hocking</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          (
          <year>1976</year>
          )
          <article-title>"The Analysis and Selection of Variables in Linear Regression,"</article-title>
          <source>Biometrics</source>
          ,
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mehdi</surname>
            <given-names>Kaytoue</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sergei O. Kuznetsov</surname>
          </string-name>
          , Amedeo Napoli, and Sebastien Duplessis,
          <article-title>Mining gene expression data with pattern structures in formal concept analysis</article-title>
          ,
          <source>Information Sciences</source>
          , vol.
          <volume>181</volume>
          , no.
          <issue>10</issue>
          , pp.
          <fpage>19892001</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Veloso</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Jr.</surname>
            ,
            <given-names>W. M.</given-names>
          </string-name>
          (
          <year>2011</year>
          ),
          <article-title>Demand-Driven Associative Classication</article-title>
          ., Springer.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>yt class labels predicted for the test object. for iter from 1 to num:iter do S=random.sample(P osdata,size=sub:smpl N +) mine positive descr = (g1) u ::: u (gs) u (t) N egimage = fx 2 descr jx 2 N egdatag if jjN egimagejj &lt; alpha:threshold N then Add descr to positive - weak premises set else Do nothing end if S=random.sample(N egdata,size=sub:smpl N ) mine premises descr = (g1) u ::: u (gs) u (t) P osimage = fx 2 descr jx 2 P osdatag if jjP osimagejj &lt; alpha:threshold N + then Add descr to negative - weak premises set else Do nothing end if end for p = dim(set of positive - weak premises) n = dim(set of negative - weak premises) Choose voting scheme: A( ); w( ); pos:power = Aip(w(hi+)) neg:power = Ajn(w(hj )) margin = pos:power neg:power yt = pos:power neg:power</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>