<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Stability of Software Defect Prediction in Relation to Levels of Data Imbalance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>TIHANA GALINAC GRBAC</string-name>
          <email>tgalinac@riteh.hr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>GORAN MAU SˇA</string-name>
          <email>gmausa@riteh.hr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>University of Rijeka BOJANA DALBELO-BASˇ IC´</string-name>
          <email>bojana.dalbelo@fer.hr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>University of Zagreb</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Additional Key Words and Phrases: Software Defect Prediction</institution>
          ,
          <addr-line>Data Imbalance, Feature Selection, Stability</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <abstract>
        <p>Software defect prediction is an important decision support activity in software quality assurance. Its goal is reducing veri cation costs by predicting the system modules that are more likely to contain defects, thus enabling more e cient allocation of resources in veri cation process. The problem is that there is no widely applicable well performing prediction method. The main reason is in the very nature of software datasets, their imbalance, complexity and properties dependent on the application domain. In this paper we suggest a research strategy for the study of the performance stability using di erent machine learning methods over di erent levels of imbalance for software defect prediction datasets. We also provide a preliminary case study on a dataset from the NASA MDP open repository using multivariate binary logistic regression and forward and backward feature selection. Results indicate that the performance becomes unstable around 80% of imbalance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>1:2
Several solutions are offered for the data imbalance problem. However, these solutions are not equally
effective in all application domains. Moreover, there is still an open question regarding the extent to
which imbalanced learning methods help with learning capabilities. This question should be answered
with extensive and rigorous experimentation across all application domains, including software defect
prediction, aiming to explore underlaying effects that would lead to fundamental understandings [He
and Garcia 2009].</p>
      <p>The work presented in this paper is a step in that direction. We present an research strategy that
aims to explore performance stability of software defect prediction models in relation to levels of data
imbalance. As an illustrative example we present an experiment taken to Stability of Software Defect
Prediction in Relation to Levels of Data Imbalance our strategy. We observed how learning
performance, with and without stepwise feature selection, in case of logistic regression learner, is changing
over a range of imbalances in the context of software defect prediction. The findings are just indicative
and are to be explored by exhausting experimenting aligned with proposed strategy.</p>
    </sec>
    <sec id="sec-2">
      <title>1.1 Complexity of software defect prediction data</title>
      <p>Software defect prediction (SDP) is concerned with early prediction of system modules (file, class,
module, method, component, or something else) that are likely to have a critical number of faults
(above certain threshold value, THR). In numerous studies it is identified that these modules are not so
common. In fact, they are special cases, and that is why they are harder to find. Dependent variable
in learning models is usually a binary variable with two classes labeled as ’fault–prone’ (FP) and
’not–fault–prone’ (NFP). The number of FP modules usually is much lower, and represents a minority
class, than the number of NFP modules which represents a majority class. Datasets with significantly
unequal distributions of minority over majority class are imbalanced. Independent variables used
in SDP studies are numerous. In this paper we will address SDP based on the static code metrics
[McCabe 1976].</p>
      <p>
        In SDP datasets the level of class imbalance varies for various software application domains. We
reviewed the software engineering publications dealing with software defect prediction and we noticed
that the percentage of the non-fault prone modules (%NFP) in the datasets varies a lot
        <xref ref-type="bibr" rid="ref2">(from 1% in
medical record system [Andrews and Stringfellow 2001] to more then 94% in telecom system
[Khoshgoftaar and Seliya 2004])</xref>
        for various software application domains (telecom industry, aeronautics,
radar systems, etc.). Since there are SDP initiatives on datasets with a whole range of imbalance
percentages, we are motivated to determine the percentage at which data imbalance becomes a problem,
i.e., learners become unstable.
      </p>
      <p>
        As already mentioned above, the random variables measured in software engineering usually do not
follow any distribution in general, and the applicability of classical mathematical modeling methods
and techniques is limited. Hence, algorithms from the machine learning have been widely adopted.
Among various learning methods used in the defect prediction approaches, this paper will explore the
capabilities of multivariate binary logistic regression (LR). Our ultimate goal is not to validate
different learning algorithms but to explore learning performance stability over different levels of
imbalance. The LR has shown very good performance in the past and is known to be a simple but
robust method. In [Lessmann et al. 2008] it is the 9th best classifier among 22 examined (9/22) and at
the same time it is the 2nd best statistical classifier among 7 of them (2/7). The stepwise regression
classifier was the most accurate classifier (1/4) and was outperformed only in cases with many outliers
in [Shepperd and Kadoda 2001]. Very good performance of logistic regression was also observe
        <xref ref-type="bibr" rid="ref3">d in
[Kaur and Kaur 2012</xref>
        ] (3/12 it terms of accuracy an
        <xref ref-type="bibr" rid="ref3">d AUC), [Banthia and Gupta 2012</xref>
        ] (1/5 both with
and without preprocessing of 5 raw NASA datas
        <xref ref-type="bibr" rid="ref12">ets), [Giger et al. 2011</xref>
        ] (1/8 in terms of median AUC
from 15 open source projects), [Jiang et al. 2008] (2/6 in terms of AUC and 3/6 according to Nemenyi
post-hoc test), etc. However, neither of the studies has analyzed the performance of logistic regression
classifier in relation to data imba
        <xref ref-type="bibr" rid="ref7">lance. The study [Provost 2000</xref>
        ] assumes that in majority of published
work the performance of logistic learner would be significantly improved, if it is adequately used. We
will refer to this issue in more detail in Section 3.
      </p>
      <p>
        As in the whole software engineering field, an important problem in software defect prediction is the
lack of quality industrial data, and therefore generalization ability and further propagation of research
results is very limited. The problem is usually that this data are considered as confidential by the
industry, or the data are not available at all for industry with low maturity. To overcome these obstacles,
there are initiatives for open source repositories of datasets aligned with the goal of improving
generalization of research results. However, the problem of generalization still remains, because usually
the open repositories contain data from a particular type of software (e.g. NASA MDP repository, open
source software repositories, etc.) and/or of questionabl
        <xref ref-type="bibr" rid="ref12">e quality [Gray et al. 2011</xref>
        ].
      </p>
      <p>
        In this study we used NASA MDP datasets and have carefully addressed all the potential issues,
i.e. remove
        <xref ref-type="bibr" rid="ref3">d duplicates [Gray et al. 2012</xref>
        ]. This selection is motivated by simple comparison of results
with the related work, so that our contribution can be easily incorporated to the existing knowledge
base of imbalance problem in the SDP area.
      </p>
    </sec>
    <sec id="sec-3">
      <title>1.2 Experimental approach</title>
      <p>Our goal is to explore stability of evaluation metrics for learning SDP datasets with machine learning
techniques across different levels of imbalance. Moreover, we want to evaluate potential sources of
bias in study design by constructing number of experiments in which we diverse one parameter per
experiment. Parameters that are subject of change are explained briefly in Sect.2.</p>
      <p>
        To integrate conclusions obtained from each experiment a meta–analytic statistical analysis is
proposed. These methods are suggested by number of authors as tool for generalizing the results and
integrating knowledge
        <xref ref-type="bibr" rid="ref8">across many studies [Brooks 1997</xref>
        ]. We propose the following steps:
(1) Acquiring data. A sample S of independent random variables X1; : : : ; Xn measuring different
features of a system module, and a binary dependent variable Y measuring fault–proneness (with
Y = 1 for FP modules and Y = 0 for NFP modules) is obtained from a repository (e.g. open
repository, open source projects, industrial projects).
(2) Data preprocessing.
      </p>
      <p>(a) Data cleaning, noise elimination, sampling.
(b) Data multiplication. From the sample S obtained in step (1) a training set of size 2=3 the size
of S and a validation set of size 1=3 the size of S are chosen at random k times. In this way
k training samples T1; : : : ; Tk and k validation samples V1; : : : ; Vk are obtained. These samples
are categorized into ` categories with respect to the data imbalance defined as the percentage
F PTi
of the NFP modules in Ti and calculated as: %N F PTi = F PTi +NF PTi .
(c) Feature selection. For each training set Ti a feature selection is performed. As a result some of
the random variables Xj are excluded from the model. The inclusion/exclusion frequencies of
Xj for each of the categories introduced in step (2b) are recorded.
(3) Learning.</p>
      <p>(a) Building a learning model. A learning model is built for each training set Ti using the learning
techniques under consideration.
(b) Evaluating model performance. Using the validation set Vi, the model built in step (3a) is
evaluated using various evaluation metrics. Let M be the random variable measuring the value of
one of these metrics.
(4) Statistical analysis.
(a) Variation analysis. The differences between ` samples of a random variable M obtained from
samples Ti and Vi belonging to different categories introduced in step (2b) are analyzed using
statistical tests. This step is repeated for each evaluation measure used in step (3b).
(b) Cross-dataset validation. The whole process is repeated from step (1) for m datasets from
various application domains and sources. The differences between ` m samples of a random variable
M are analyzed using statistical tests and the results reveal whether general behavior exists.</p>
      <p>To summarize, the conclusions are based on the results of statistical tests comparing the mean values
of performance evaluation metrics (see Table I) across different data imbalances of a training sample.
The stability of performance evaluation metrics obtained with different feature selection procedures is
evaluated in the same way.</p>
    </sec>
    <sec id="sec-4">
      <title>2. DATA IMBALANCE</title>
      <p>
        Data imbalance has received considerable attention within the data mining community during the last
decade. It becomes a central point of this research, since the problem is present in a majority of data
minin
        <xref ref-type="bibr" rid="ref5">g application areas [Weiss 2004</xref>
        ]. In general data imbalance degrades the learning performance.
The problem arises with learning accuracy of the minority class, in which we are usually more
interested. Usually, we are interested to timely predict rare events represented by the minority class, for
which the probability of its occurrence is low, but its occurrence leads to significant costs.
      </p>
      <p>
        For example, suppose that only very low number of system modules is faulty, which is the case with
systems with very low tolerance on failures (e.g. medical systems, aeronautic system,
telecommunications, etc.). Suppose that we did not identify faulty module with the help of a software defect prediction
algorithm, and due to that have developed defect detection strategy not concentrating on that
particular module. Thus, we omit to identify a fault in our defect detection activity, and this fault slips to the
customer site. Failure caused by this fault at customer site would then imply significant costs contained
of several items: paying penalty to customer, losing customer confidence, causing additional expenses
due to corrective maintenance, additional costs in all subsequent system revisions and additional cost
during system evolution. This cost would be considered as misclassification cost of wrongly classified
positive class (note that positive class in the context of defect prediction algorithm is a faulty module).
On the other hand, misclassification cost of wrongly classified negative class would be much lower,
because it would involve just more defect detection activities. Obviously, the misclassification costs are
unequally weighted and this is the main obstacle in applying standard machine learning algorithms,
because they usually assumes the same or similar conditions in learning and app
        <xref ref-type="bibr" rid="ref7">lication environment
[Provost 2000</xref>
        ].
      </p>
      <p>The study [Provost 2000] makes a survey of data imbalance problems and methods addressing these
problems. Although different methods are recommended for data imbalance problems, it does not give
definite answers regarding their applicability in the application context. Some answers are obtained
by other researchers in that field afterwards, and a more recent survey is given in [He and Garcia
2009]. Still no definite guideline exists that could guide practitioners.</p>
    </sec>
    <sec id="sec-5">
      <title>2.1 Dataset considerations</title>
      <p>
        The most popular approach to the class imbalance problem is the usage of artificially obtained balanced
dataset. There are several sampling methods proposed for that purpose. In a recent work [Wang and
Yao 2013] an experiment with some of the sampling methods is conducted. However, it is
        <xref ref-type="bibr" rid="ref1">concluded in
[Kamei et al. 2007</xref>
        ] that sampling did not succeed to improve performance with all the
        <xref ref-type="bibr" rid="ref1">classifiers. In
[Hulse et al. 2007</xref>
        ] it is identified that classifier performance is improved with sampling, but individual
learners respond differently on sampling.
      </p>
      <p>Another problem with datasets is that in practice, the datasets are often very complex, involving a
number of issues like overlapping, lack of representative data, within and between class imbalance,
and often high dimensionality. The effects of these issues were widely analyzed separately sample size
in [Raudys and Jain 1991], dimensionality reduction: [Liu and Yu 2005], noise elimination
[Khoshgoftaar et al. 2005], but not in conjunction with the data imbalance. The study performed in [Batista
et al. 2004] observes that the problem is related to a combination of absolute imbalance and other
complicating factors. Thus, the imbalance problem is just an additional issue in complex datasets such
as datasets for software defect prediction.</p>
      <p>
        Different aspects of feature selection in relation to class imbalance has been studied in
[Khoshgoftaar et al. 2010; Gao and
        <xref ref-type="bibr" rid="ref11">Khoshgoftaar 2011</xref>
        ; Wang et al. 2012]. All these studies were performed on
datasets from the NASA MDP repository. In this work we also used a stepwise feature selection as a
preprocessing step, because the dataset is high dimensional and we experiment with logistic
regression. Hence, we were able to investigate the stability of the performance with and without feature
selection procedure over different levels of imbalance.
      </p>
      <p>Besides the methods explained above for obtaining artificially balanced datasets, another approach
is to adapt standard machine learning algorithms to operate for imbalance datasets. In that case
the learning approach should be adjusted to the imbalanced situation. A complete review of such
approaches and methods can be found in [He and Garcia 2009].</p>
    </sec>
    <sec id="sec-6">
      <title>2.2 Evaluation metrics</title>
      <p>
        Another problem of standard machine learning algorithms for imbalanced data is in usage of
inadequate evaluation metrics during learning procedure or to evaluate final result. Evaluation metrics
are usually derived from the confusion matrix and are given in Table I. They are defined in terms
of the following score values. A true positive (TP) score is counted for every correctly (true) classified
fault-prone module, and a true negative (TN) score for every correctly (true) classified non-fault-prone
module. The other two possibilities are related to false prediction. A false positive (FP) score is counted
for every false classified or misclassified non-fault-prone module (often referred to as Type II error),
and a false negative (FN) score is counted for every false classified or misclassified fault-prone module
(often referred to as Type I error) [Runeson et al. 2001; Khosh
        <xref ref-type="bibr" rid="ref5">goftaar and Seliya 2004</xref>
        ]. For example,
classification accuracy ACC, the most commonly used evaluation metric in standard machine
learning algorithms, is not able to value the minority class appropriately, and leads to poor classification
performance of minority class.
      </p>
      <p>In the case of class imbalance, the precision (PR) and recall (TPR) metrics given in Table I are
recommended in number of studies [He and Garcia 2009], as well as the F –measure and G–mean
which are not used here. The precision and recall in combination give a measure of correctly classified
fault–prone modules. Precision measures exactness, i.e., how many fault–prone modules are classified
correctly, and recall measures completeness, i.e., how many fault–prone are classified correctly.</p>
      <p>Metrics
Accuracy (ACC)
True positive rate (TPR)
(sensitivity, recall)
Precision (PR)
(positive predicted value)</p>
      <p>
        The output of a probabilistic machine learning classifier is the probability for a module to be
faultprone. Therefore, a cutoff percentage has to be defined in order to perform classification. Since choosing
a cutoff value leaves room to bias and possible inconsistencies in a study [Lessmann et al. 2008], there
is another measure that deals with that problem called the area under curve, AUC [Fawce
        <xref ref-type="bibr" rid="ref9">tt 2006</xref>
        ].
It takes into account the dependence of T P R and a similar metric for false positive proportion on the
cutoff value.
      </p>
      <p>All of the aforementioned techniques are not cost sensitive, and in the case of rare cases with very
high misclassification cost of type I error the key performance indicator is cost. The most favorable
evaluation criteria for imbalanced datasets are cost curves and is also recommended in [Jiang et al.
2008] for SDP domain.</p>
    </sec>
    <sec id="sec-7">
      <title>3. PRELIMINARY CASE STUDY</title>
      <p>To illustrate the application of the research strategy proposed in Section 1.2, verify strategy, provide
evidence for the dependence of the machine learning performance on the level of data imbalance, and
indicate our future goals, we have undertaken a preliminary case study.
(1) Dataset KC1 from NASA MDP repository has been acquired. It consists of n = 29 features, i.e.,
independent variables Xj. The dependent variable in this dataset is the number of faults in a
system module. From this variable we derived binary dependent variable Y by setting ten different
thresholds for fault proneness, from 1 to 19 with step of 2 (1, 3, 5,...). In this way we obtained ten
different samples S and we continue the analysis for all of them.
(2) (a) The well known issues with the dataset are eliminated using data cleaning tool [Shepperd et al.</p>
      <p>
        2013].
(b) For each of the ten samples obtained in step (1), we made 50 iterations of the random splitting
into training and validation samples. Thus we obtained k = 500 samples Ti and Vi with the
range of data imbalance from 51% to 96%. The samples are categorized into ` = 5 categories of
equal length (each spanning 9%).
(c) In the case study we also consider the influence of a feature selection procedure, as already
mentioned in 2. We consider the forward and backward stepwise selec
        <xref ref-type="bibr" rid="ref9">tion procedure [Han and
Kambar 2006</xref>
        ]. The decision for inclusion and exclusion of a feature is based on level of
statistical significance, the p value. The common significance levels for inclusion and exclusion of
features are use
        <xref ref-type="bibr" rid="ref3">d as in [Mausa et al. 2012</xref>
        ; Briand et a
        <xref ref-type="bibr" rid="ref7">l. 2000</xref>
        ] with p in = 0:05 and p out = 0:1
respectively. The percentage of inclusion of a feature for both procedures and different
categories of data imbalance are given in Table II. We conclude that feature selection stability of
some features is very tolerant to data imbalance (e.g. Feature 5, 22, 28, 29 is always excluded,
for both forward and backward model). Some features are very stable until certain level of
balance (for example Feature 2 is always included 100% until category with data imbalance of
78%). It is also interesting to observe that some features have similar feature selection
stability in ideal balance case and highly imbalanced case, whereas for moderate imbalance have
opposite feature selection decision.
(3) (a) Learning models are built using multivariate binary logistic regression (LR) [Hastie et al.
2009]. The model incorporates more than just one predicting variable and in fault predicting
case performs according to the equation
(X1; X2; :::Xn) = 1 + eC0+C1X1+:::+CnXn ;
eC0+C1X1+:::+CnXn
(1)
where Cj are the regression coefficients corresponding to Xj, and is the probability that a
fault was found in a class during validation. In order to obtain a binary outgoing variable,
a cutoff value splits the results into two categories. Researchers often set the cutoff value to
0.5 [Zimmermann and Nagappan 2008]. However, the logistic regression is also robust to data
imbalance and this robustness is achieved with setting of cutoff value to optimal value
dependent on misclassification costs [Basili et al. 1996]. Our goal is to explore learning performance
over different imbalance levels. However, in this study, due to space limitation, we provide
preliminary results exploring learning performance stability of standard learning algorithms.
Therefore, we provide results of experiments with cutoff value set to 0.5 (that is how standard
learning algorithms equally weight misclassification costs). We considered there three
different models (with forward feature selection, backward feature selection and without feature
selection) and for each of these models, the coefficients are calculated separately.
(b) For all validation samples from step (2b) we count the TN, TP, FN and FP scores of the
corresponding model, and calculate the learning performance evaluation metrics ACC, TPR (Recall),
AUC and Precision using formulas in Table I.
(4) We made a statistical analysis of the behavior of evaluation metrics measured in step (3b) between
different categories introduced in step (2c). Since the samples are not normally distributed, we used
the non-parametric tests. The Kruskal-Wallis test showed for all metrics that the values depend
on the category. To explore the differences further, we applied multiple comparison test. It reveals
that all considered evaluation metric become unstable at the level of imbalance of 80%. According
to the theory explained in section 2, we expect that we will get significantly different mean values
for all metrics in category of highest data imbalance (90% - 100%).
1:8
      </p>
    </sec>
    <sec id="sec-8">
      <title>4. DICUSSION</title>
      <p>Data imbalance problem has been widely investigated and there were numerous approaches studying
its effects aiming to propose a general solution to that problem. However, from the experiments in
machine learning theory it becomes obvious that this is not only related to proportion of minority over
majority class but there are also other influences present in complex datasets. As the datasets in
software defect prediction (SDP) research area are usually extremely complex, there is a huge unexplored
area of research related to applicability of these techniques in relation to the level of data imbalance.
That is exactly our main motivation for this work.</p>
      <p>There are many approaches, depending on particular dataset, to SDP and development of the
learning model. Since we are interested in the performance stability of machine learners over SDP datasets,
we should rigorously explore the strengths and limitations of these approaches in relation to the level
of data imbalance. Therefore, we present an exploratory research strategy and an example of a case
study performed according to this strategy. Although, we use our experiment to eliminate as much as
possible inconsistencies and threats of applying the strategy, there is still place for improvement.</p>
      <p>In our case study we present how performance stability is significantly degraded at a higher level of
imbalance. This confirms the results obtained by other researchers using different approaches. That
conclusion have proved reliability of our strategy. Moreover, with the help of our research strategy we
confirmed that feature selection becomes instable with higher data imbalance. We have also observed
that the feature selection is consistent across levels of imbalance for some features.</p>
      <p>
        Future work should involve extensive exploration of SDP datasets with the proposed strategy. Our
vision is that at the end we can gain deeper knowledge about imbalanced data in SDP and applicability
of techniques in different levels of imbalance. Finally, we would like to categorize datasets using the
proposed strategy and results of this exhaustive research that would serve as a guideline for
practitioners while developing software defect prediction model.
H. Wang, T. M. Khoshgoftaar, and A. Napolitano. An Empirical Study on the Stability of Feature Selection for Imbalanced
Software Engineering
        <xref ref-type="bibr" rid="ref3">Data. In Proceedings of the 2012</xref>
        11th International Conference on Machine Learning and Applications
- Volume 01, ICMLA ’12, pages 317–323, Washington, DC, USA, 317–323.
      </p>
      <p>S. Wang and X. Yao. Using Class Imbalance Learning for Software Defect Prediction. IEEE Transactions on Reliability, 62(2):434
- 443, 2012.</p>
      <p>G.M. Weiss. Mining with rarity: a unifying framework. In SIGKDD Explor. Newsl., 6(1):7–19, 2004.</p>
      <p>T. Zimmermann and N. Nagappan. Predicting defects using network analysis on dependency graphs. In Proceedings of the 30th
international conference on Software engineering, ICSE ’08, pages 531–540, New York, NY, USA, 2008. ACM.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Andersson</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Runeson</surname>
          </string-name>
          .
          <article-title>A replicated quantitative analysis of fault distributions in complex software systems</article-title>
          .
          <source>IEEE Trans. Softw</source>
          . Eng.,
          <volume>33</volume>
          (
          <issue>5</issue>
          ):
          <fpage>273</fpage>
          -
          <lpage>286</lpage>
          , May
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Andrews</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Stringfellow</surname>
          </string-name>
          .
          <article-title>Quantitative analysis of development defects to guide testing: A case study</article-title>
          .
          <source>Software Quality Control</source>
          ,
          <volume>9</volume>
          :
          <fpage>195</fpage>
          -
          <lpage>214</lpage>
          ,
          <year>November 2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Banthia</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          .
          <article-title>Investigating fault prediction capabilities of five prediction models for software quality</article-title>
          .
          <source>In Proceedings of the 27th Annual ACM Symposium on Applied Computing, SAC '12</source>
          , pages
          <fpage>1259</fpage>
          -
          <lpage>1261</lpage>
          , New York, NY, USA,
          <year>2012</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Basili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Briand</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. L.</given-names>
            <surname>Melo</surname>
          </string-name>
          .
          <article-title>A validation of object-oriented design metrics as quality indicators</article-title>
          .
          <source>IEEE Trans. Software Engineering</source>
          ,
          <volume>22</volume>
          (
          <issue>10</issue>
          ):
          <fpage>751</fpage>
          -
          <lpage>761</lpage>
          ,
          <year>October 1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>G. E. A. P. A.</given-names>
            <surname>Batista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Prati</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.C.</given-names>
            <surname>Monard</surname>
          </string-name>
          .
          <article-title>A study of the behavior of several methods for balancing machine learning training data</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>20</fpage>
          -
          <lpage>29</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>L. C. Briand</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Daly</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Porter</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          <article-title>¨ st. A comprehensive empirical validation of product measures for object-oriented systems</article-title>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>L. C. Briand</surname>
            , J. Wu¨ st,
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Daly</surname>
            , and
            <given-names>D. V.</given-names>
          </string-name>
          <string-name>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>Exploring the relationship between design measures and software quality in object-oriented systems</article-title>
          .
          <source>J. Syst. Softw.</source>
          ,
          <volume>51</volume>
          :
          <fpage>245</fpage>
          -
          <lpage>273</lpage>
          , May
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Brooks</surname>
          </string-name>
          .
          <article-title>Meta Analysis-A Silver Bullet for Meta-Analysts. Empirical Softw</article-title>
          . Engg.,
          <volume>2</volume>
          (
          <issue>4</issue>
          ):
          <fpage>333</fpage>
          -
          <lpage>338</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Fawcett</surname>
          </string-name>
          .
          <article-title>An introduction to ROC analysis</article-title>
          .
          <source>Pattern Recogn. Lett.</source>
          ,
          <volume>27</volume>
          (
          <issue>8</issue>
          ):
          <fpage>861</fpage>
          -
          <lpage>874</lpage>
          , Aug.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>N. E.</given-names>
            <surname>Fenton</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Ohlsson</surname>
          </string-name>
          .
          <article-title>Quantitative analysis of faults and failures in a complex software system</article-title>
          .
          <source>IEEE Trans. Softw</source>
          . Eng.,
          <volume>26</volume>
          (
          <issue>8</issue>
          ):
          <fpage>797</fpage>
          -
          <lpage>814</lpage>
          , Aug.
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Gao</surname>
          </string-name>
          and
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Khoshgoftaar</surname>
          </string-name>
          .
          <article-title>Software defect prediction for high-dimensional and class-imbalanced data</article-title>
          .
          <source>In SEKE</source>
          , pages
          <fpage>89</fpage>
          -
          <lpage>94</lpage>
          .
          <source>Knowledge Systems Institute Graduate School</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Giger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pinzger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. C.</given-names>
            <surname>Gall</surname>
          </string-name>
          .
          <article-title>Comparing fine-grained source code changes and code churn for bug prediction</article-title>
          .
          <source>In Proceedings of the 8th Working Conference on Mining Software Repositories, MSR '11</source>
          , pages
          <fpage>83</fpage>
          -
          <lpage>92</lpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>