<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Three-factor nonlinear regression model of estimating the size of java-software</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleksandr Oriekhov</string-name>
          <email>oleksandr.oriekhov@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetyana Farionova</string-name>
          <email>tetyana.farionova@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liubava Chernova</string-name>
          <email>liubava.chernova@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Makarov National University of Shipbuilding, Ukraine</institution>
          ,
          <addr-line>Heroes avenue, 9, Mykolaiv, 54007</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <fpage>3</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper proposes a three-factor nonlinear regression model for Java-software size estimation. Java is one of the most widely used programming languages in the world and is actively used in the development of various software projects. The software size estimation plays a key planning role at the early stages of software project planning for the successful implementation of software development projects. The software size estimation is valuable information that is used to predict software development effort estimation using parametric models such as COCOMO, COCOMO II, COSYSMO, etc. The aim of the study is to build the three-factor nonlinear regression model for early KLOC size estimation of Java-software on the basis of multivariate Johnson SB family normalizing transformations to increase the reliability and accuracy of the software size estimation at the early stage of software project planning using UML class diagram metrics. The object of the study is the process of size estimation for open-source Java-software. The subject of the study is the nonlinear regression models to estimate the software size. To achieve this goal we collected software code metrics information from 571 open-source Java software projects. The obtained dataset is split up in training and testing samples accordingly Cross-Validation technique. The regression model building is based on an iterative method that includes bijective multivariate normalizing transformation, multivariable outliers detection, confidence and prediction interval building. The obtained nonlinear regression model is compared to the existing models by the regression models quality criteria such as the determination coefficient, mean magnitude of relative error (MMRE) and the percentage of prediction of the relative error level 0.25 (PRED(0.25)). The comparison results confirm increasing the accuracy of the software size estimation using the obtained regression model on both training and testing samples. The prospects for further research may include increasing the samples of Java-software code metrics, usage of different software code metrics, extending the set of independent factors and using different normalizing transformations for early KLOC estimation.</p>
      </abstract>
      <kwd-group>
        <kwd>Software size estimation</kwd>
        <kwd>nonlinear regression model</kwd>
        <kwd>normalizing transformation</kwd>
        <kwd>Java</kwd>
        <kwd>CrossValidation1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>The software size estimation plays a key planning role at the early stages of software project</title>
        <p>planning for the successful implementation of software development projects. The software size
can be represented as functional points or number of code lines (KLOC - kilo lines of code). Both
variants have their own advantages, drawbacks and restrictions. One of the crucial advantages of</p>
      </sec>
      <sec id="sec-1-2">
        <title>KLOC usage is in taking into account such important parameter as environmental factors which</title>
        <p>include programming languages or software category [1] and KLOC parameter is widely used in
software development effort estimation parametric models such as COCOMO, COCOMO II, SLIM
[2], COSYSMO [3], etc. Software development effort estimation is one of the significant indicators
of budget, resources and duration planning of any project for software development business.
Reliable estimates of software development effort provide valuable information for project
planning, and it helps to take into account risks, to recognize planning gaps and to increase the
efficiency of the development process [4].</p>
        <p>Java programming language is one of the most widely used in software development [5]. Java
software varies from utility and micro web-service applications to huge information or automotive
or operation system development. Java is object oriented programming (OOP) language and it
executes in a separated environment - Java Virtual Machine that makes it a universal instrument
for any purpose [6]. Java as an OOP language has compatibility with UML software projection
therefore they both have a shared set of software code metrics such as quantity of classes, total
quantity of methods and fields or depending on access modifiers, inheritance, association or
complexity metrics.</p>
      </sec>
      <sec id="sec-1-3">
        <title>The CHAOS report 2015 [7] of The Standish group research confirms that large and grand</title>
        <p>projects have much higher failure rates in comparison with small, moderate or medium projects.
Software development of most of the large projects is based on traditional project management
methodologies like software development life cycle and waterfall, obviously it makes most of the
parametric models more suitable for software development effort estimation than agile
methodologies. Otherwise, company or team specific calibration of the parametric models that can
work with uncertainty, such as COCOMO II and could provide reliable estimates for the projects
that follow Agile methodol
stages of development because of the complexity of integrating with highly distributed and
different data and flows.</p>
      </sec>
      <sec id="sec-1-4">
        <title>Since Java is one of the most demanded programming languages, the size estimation is an</title>
        <p>important task of the software project management life cycle and the share of failures or problems
in software development is still high, lead us to confirmation the calculating requirement of
reliable and accurate estimates of the size in code lines requires appropriate models for Java
software to provide valuable information for project planning at the early stages.</p>
      </sec>
      <sec id="sec-1-5">
        <title>The aim of the study is to build the three-factor nonlinear regression model for early KLOC</title>
        <p>estimation of Java-software on the basis of multivariate Johnson family normalizing
transformations to increase the reliability and accuracy of the Java-software size estimation at the
early stage of software project planning using UML class diagram metrics. The obtained math
model should achieve higher quality criteria values in comparison with the existing math models of</p>
      </sec>
      <sec id="sec-1-6">
        <title>Java-software size estimation.</title>
      </sec>
      <sec id="sec-1-7">
        <title>The object of the study is the process of size estimation for open-source Java-software.</title>
      </sec>
      <sec id="sec-1-8">
        <title>The subject of the study is the nonlinear regression models to estimate the software size.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Review of the literature</title>
      <sec id="sec-2-1">
        <title>Nowadays, Java-software size estimation is achieved with linear and nonlinear regression</title>
        <p>equations and models for general Java-software or categories such as web-service,
informationalsystems, etc on the metrics basis of the conceptual data model derived from a UML class diagram.</p>
      </sec>
      <sec id="sec-2-2">
        <title>The papers [8,9] propose linear regression equations on the basis of 3 independent factors. The</title>
      </sec>
      <sec id="sec-2-3">
        <title>Java KLOC estimation in [8] is based on total amount of classes (CLASS), coupling between objects</title>
        <p>(CBO) and total quantity of fields per class (TFQ) and the [9] is based on CLASS, CBO and average
value of TFQ per class and proposes better variant of the equation in comparison with [8].</p>
      </sec>
      <sec id="sec-2-4">
        <title>Typically, the software code metrics have non-Gaussian distribution. Therefore it has restrictions in the ability to apply linear models for KLOC estimation. One of the theoretical requirements of linear regression models.</title>
      </sec>
      <sec id="sec-2-5">
        <title>The existing nonlinear one-factor and multifactor regression models [10, 11, 12, 13] propose</title>
        <p>Java-software size estimation depending on different variants of software code metrics. A good
benefit of the models is a proposed prediction intervals for possibility to make expected, pessimistic
and optimistic KLOC estimates. The model [10] is based on dataset from researches [8] and [9], and
it can not guarantee robustness and accuracy of KLOC estimation due to Java language evolution
and changing in modern software development practices that makes significant impact on software
metrics previously and nowadays. The [10, 11, 12] regression models are based on restricted
samples of Java applications that also have an impact on reliability and robustness of the nonlinear
regression models. The latest research [13] proves the models [10, 11, 12] are restricted in
estimation of code lines on different samples of code metrics from Java-software therefore the
accurate result is not guaranteed. The [13] research offers one-factors and two-factor nonlinear
regression models which were built using Cross-Validation technique. The regression models
quality criteria such as the coefficient of determination R2, a mean magnitude of relative error
MMRE and percentage of prediction for magnitude of relative error (MRE) level 0.25 PRED(0.25)
reveal good quality level of the obtained two-factor model but the values of some quality criteria
thresholds are not exceeded. Thus, study [14] confirms that with increasing sample sizes, the
values of the quality criteria for the accuracy of predicting mathematical models for training and
testing samples approach each other, which indicates the benefits of using large samples to build
mathematical models in combination with Cross-Validation technique and adding more factors to
nonlinear regression models allows to achieve higher accuracy accordingly quality criteria [13].</p>
      </sec>
      <sec id="sec-2-6">
        <title>The review and analysis of the literature has shown that the linear equations and nonlinear regression models exist for estimating Java-software code size and confirms the necessity of increasing the robustness, reliability and accuracy of Java-software size estimation on early stages of project planning.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Problem statement</title>
      <p>
        To achieve the aim of the research it is necessary to build the three-factor nonlinear regression
model on the basis of multivariate normalizing transformation using the original training and
testing non-Gaussian four-dimension samples of Java-software metrics to build and verify the
regression model. The training and testing sample includes the thousand lines of code (KLOC) Y,
the total number of classes X1, total quantity of unique methods invocations in the code (RFC
response for class) X2 and average value of public and protected methods per class (VMQ - visual
methods quantity) X3 from code metrics information of Java open-source software. The code
metrics information of projects does not intersect between training and testing samples. Suppose
there are bijective multivariate normalizing transformation of non-Gaussian random vector  =
{ ,  1,  2, . . . ,   } into Gaussian random vector  = {  ,  1,  2, . . . ,   } is given by
 =  ( ), (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
where k is number of factors (regressors or independent variables) and the inverse
transformation of (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) is given by
 =  −1( ),
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
where 
is
a
vector
of
bijective
normalizing
transformation
functions,  =
{  ,  1,  2, . . . ,   } .
      </p>
      <sec id="sec-3-1">
        <title>On the bases of the transformations (1) and the required nonlinear regression model has the</title>
        <p>form  =  ( ,  1,  2,  3,  ), where  is a Gaussian random variable that defines residuals, F is
three-factor nonlinear regression models and Y is the dependent variable KLOC.
4. Materials, research methods and three-factor nonlinear regression
model constructing
4.1. Three-factor nonlinear regression model constructing</p>
      </sec>
      <sec id="sec-3-2">
        <title>The authors collected code metrics dataset of 571 open-source Java software hosted on the GitHub</title>
        <p>platform (https://github.com) using the CK tool (https://github.com/mauricioaniche/ck). The
following metrics were acquired: the actual total of Java-software size measured in number of
thousand lines of code (KLOC), total quantity of classes (CLASS) X1, total number of unique method
calls in classes (RFC) X2 and average value of visible methods per class of the software (aVMQ) X3.</p>
      </sec>
      <sec id="sec-3-3">
        <title>CLASS, RFC, aVMQ, metrics can be obtained at an early stage of software planning from the</title>
        <p>conceptual model. The obtained dataset was randomly divided into training and testing samples
with sizes of 286 and 285 rows of Java-software project data, respectively. The distribution of
metrics relative to KLOC is shown in Figure 1 for training dataset and in Figure 2 for testing
dataset.</p>
      </sec>
      <sec id="sec-3-4">
        <title>The variance inflation factors (VIFs) technique is used to check the predictors multicollinearity</title>
        <p>level to avoid a high correlation between regression factors and sensitivity of the model to random
changes in the data.</p>
        <p>For a multiple regression model with  factors Xi ,  = 1,2, . . . , k, the VIFs are represented by the
diagonal elements of the inverse correlation  ×  matrix. If the value of the VIF coefficient
exceeds 10 (the threshold value), a high level of multicollinearity exists between the independent
variables [15].</p>
      </sec>
      <sec id="sec-3-5">
        <title>For factors X1, X2, and X3, the VIFs are equal to 6.0889, 6.0890 and 1.0 respectively, thus it</title>
        <p>indicates the absence of multicollinearity between the factors of the regression model.</p>
      </sec>
      <sec id="sec-3-6">
        <title>The nonlinear regression models building technique is based on statistical analysis methods [16,</title>
        <p>17] and it is based on detecting and discarding outliers in nonlinear regression analysis of
non</p>
      </sec>
      <sec id="sec-3-7">
        <title>Gaussian data and includes bijective normalizing transformations, Mahalanobis distance outliers</title>
        <p>detection, regression residuals distribution verification and prediction interval detection.</p>
        <p>According to the technique it is recommended to detect and discard only one outlier in an
iterative way once it is detected. In case if an outlier is detected, the technique starts from the first
step using the modified sample without the detected outlier from the previous iteration. Otherwise
the nonlinear regression model is successfully constructed. The first step of the technique is based
on normalizing transformation.</p>
        <p>=  ̂ +  ̂
(</p>
        <p>−  ̂
 ̂ +  ̂ − 
),
where
respectively.</p>
        <p>parameters of Johnson SB family normalizing transformation,  &lt; Х &lt;
 +  ,  &gt; 0,  &gt; 0;  - non-Gaussian random variable which X equals to Y, X1, X2, X3,</p>
        <p>
          In the first step for normalizing the multivariate non-Gaussian data of the training sample, we
use the Johnson univariate and multivariate transformation for the SB family. It suggests a rich set
of parameters for accurate calibration of the transformation. In our case the Johnson normalizing
translation of SB family is given by [18]:
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>And the inverse transformation is given by</title>
        <p>=  ̂ +  ̂ [1 +  −( −̂ )/̂ ]−1,
where Z is a Gaussian random variable which Z equals to ZY, ZX1, ZX2, ZX3, respectively.</p>
        <p>
          Parameters of Johnson univariate and multivariate transformations (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) are estimated by the
finding solution of the equation [19] for computer programming task given by
 ̂ =
        </p>
        <p>{ 2 + (ℇ − 3)2 + ̅Z̅2̅ + (  2 − 1)2 },
where  is a vector of estimators,</p>
        <p>= { ,  ,  ,  };  is skewness of a random variable, ℇ is
kurtosis of a random variable;  ̅= 1

∑
 =1   ;   2 = 1</p>
        <p>∑
 =1(  −  ̅)2;   - i-th value of random
variable Z of the sample size N. For the multidimensional data, we use measurement of multivariate
skewness (β1, k) and kurtosis (β2, k) of the sample proposed in a Mardia [20] multivariate
normality test.</p>
        <p>
          Estimators for the parameters of multivariate transformation (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) vector  ̂ for metrics Y are:  ̂  =
 ̂ 3 = 952.137743 ⋅ 109 for the latest
 ̂ = 82315.076609; for X1:  ̂ 1 =
iteration.
        </p>
        <p>The second step includes outlier detection using squared Mahalanobis distance (MD) technique
in the sample of normalized data. The squared Mahalanobis distance are elements on the main
diagonal of the  2 matrix of size N x N. The MD formula is given by
and   is given by
 2 = (  − Z̅)  −1(  − Z̅)</p>
        <p>1

 =1
  =</p>
        <p>∑(  − Z̅)(  − Z̅) ,
freedom.
dimensional sample is given by</p>
        <p>where Z is a normally distributed random variable   = ( 1 ,  2 , … ,   
variance matrix and  is a means vector of independent variable of the sample, Z̅ =

) , SN is a biased sample
(̅Z̅1̅, ̅Z̅2̅, . . . , ̅Z̅k̅) .</p>
        <p>The elements of the main diagonal  2,  = 1,2, . . . 
matrix are detected as outliers if the values
exceed the threshold value of the Chi-Square  2 distribution quantile for the significant level
For the training sample 17 data rows were detected as outliers because their  2 values greater then
threshold value 18.55 of the Chi-Square  2 for the significant level  = 0.005 and 20 degrees of</p>
      </sec>
      <sec id="sec-3-9">
        <title>The linear regression model is built in the third step on the basis of the normalized multi</title>
        <p>
          =  ̂ +  =  ̂ 0 +  ̂ 1 1 +  ̂ 2 2 +  ̂ 3 3 +  ,
where is Gaussian random variable, ∼ 
the linear regression model (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ).
        </p>
        <p>The estimators are calculated by the least square method:  ̂ 0 = −23.784722,  ̂ 1 = −0.005724,
 ̂ 2 = −0.005062,  ̂ 3 = −0.008718 .</p>
      </sec>
      <sec id="sec-3-10">
        <title>In the fourth step the normality distribution of the linear regression residuals are tested with</title>
      </sec>
      <sec id="sec-3-11">
        <title>Pearson Chi-Square criteria for significant level</title>
      </sec>
      <sec id="sec-3-12">
        <title>The actual values of the Chi-Square for the residuals are 116.5 and 19.24 for the 1th and 8th</title>
        <p>
          iterations respectively and the values are bigger than quantile 16.81 of the Chi-Square distribution
2 );  ̂ 0,  ̂ 1,  ̂ 2,  ̂ 3 - estimators for parameters of
for 6 degrees of freedom.
iteratively.
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) to the linear regression models (
          <xref ref-type="bibr" rid="ref8">8</xref>
          ):
        </p>
      </sec>
      <sec id="sec-3-13">
        <title>The 2 data rows with the highest values of residuals are removed from the training sample</title>
      </sec>
      <sec id="sec-3-14">
        <title>In the fifth step, the nonlinear regression models is built by applying inverse transformation to</title>
        <p>
          =  −1( ̂ +  ) =  −1( ̂ 0 +  ̂ 1  1( 1)+  ̂ 2  2( 2)+  ̂ 3  3( 3)+  ),
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
where  −1 is inverse Johnson SB transformation (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ).
        </p>
        <p>
          In the sixth step we constructed prediction interval  ̂
 of the nonlinear regression model (
          <xref ref-type="bibr" rid="ref8">8</xref>
          ).
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )

1
̂


=  −1 ( ̂ ±   /2,    {1 +
+ (  +)  −1(  +)}

),
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
given by
1/2
  = [
        </p>
        <p>],
where</p>
        <p>For the training sample, 1 outlier is detected and discarded. For the obtained prediction interval, the
values of normalized sample means Z̅1, Z̅</p>
        <p>
          2 and Z̅3 are −1321.94070, −1447.061944 and
−1008.068369, respectively. The   /2, = 2.594580 for significant level  = 0.01 and 264
degrees of freedom;    = 0.089394. The inverse matrix of (
          <xref ref-type="bibr" rid="ref11">11</xref>
          ) is
outliers from the training sample.
        </p>
        <p>The Gaussian distribution of the four-dimension training data set ( = 268) is verified by a
multivariate normality test proposed by Mardia [20] at the latest iteration. The test is based on
measurement of multivariate skewness (β1, k) and kurtosis(β2, k) of the sample.</p>
      </sec>
      <sec id="sec-3-15">
        <title>Accordingly the test, distribution of four-dimension data X1 (CLASS), X2 (RFC), X3 (aVMQ) and Y (KLOC) is Gaussian, because multivariate skewness estimate  1/6 = 39.85 is not exceeded Chi</title>
      </sec>
      <sec id="sec-3-16">
        <title>Square quantile value estimate of multivariate kurtosis  2 = 23.73 does not exceed the value Gaussian distribution quantile which is equal to 26.18 for mean 24 and standard deviation 0.85.</title>
        <p>
          4.2. Quality criteria of the obtained nonlinear regression model
The obtained nonlinear regression model (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) is tested with regression models quality criteria [21]
such as the coefficient of determination  2, a mean magnitude of relative error MMRE and
percentage of prediction for magnitude of relative error (MRE) level 0.25 
(0.25). The MMRE
criterion is defined as
where   /2, is a quantile of T-Student distribution with  =  −  − 1 degrees of freedom and
  2 =
        </p>
        <p>=1(   −  ̂  ) ;   + is a vector of central moments of independent
where N - sample size and    is the value of the magnitude of relative error for the i-th
datapoint of the sample.</p>
        <p>The calculation of prediction percentage (PRED) for the magnitude of relative error level 0.25 is
given by
 
(0.25) =

1

∑ {
 =1
1  
0
  ≤ 0.25
 ℎ
.</p>
      </sec>
      <sec id="sec-3-17">
        <title>The acceptable conditions for the quality criteria is</title>
        <p>≤ 0.25 and 
(0.25) ≥ 0.75
for the measurement of the regression models accuracy of prediction results. The coefficient of
determination ( 2) value is acceptable if it is more or equals to 0.75 [21].</p>
        <p>
          The obtained three-factor nonlinear regression model is tested with the quality criteria to assess
the predictive reliability and accuracy on the training and the testing samples and compared with
(
          <xref ref-type="bibr" rid="ref11">11</xref>
          )
(
          <xref ref-type="bibr" rid="ref12">12</xref>
          )
(
          <xref ref-type="bibr" rid="ref13">13</xref>
          )
(
          <xref ref-type="bibr" rid="ref14">14</xref>
          )
existing one-factor and two-factor nonlinear regression models from the paper [13]. The estimates
of  2, 
prediction intervals using the training sample for the one-factor nonlinear regression model on the
basis of Johnson SB transformation [13]. Similar to the obtained three-factor nonlinear regression
model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ), the one-factor nonlinear model is built using the same techniques [16, 17] that is based
on statistical analysis of multidimensional data and multivariate reciprocal transformations. The
estimators for the Johnson SB transformation parameters of the one-factor model are  ̂  =
4311.992036,  ̂ = 3555.710081,  ̂ = 30479.194612,  ̂ = 29712.602140,  ̂  = −0.345327,
 ̂  = −3.150988,  ̂ = 2.471319⋅ 1014,  ̂ = 22965.385684 and the estimators of the linear
regression are  ̂ 0 = −799519.0395,  ̂ 1 = 0.999920. For the obtained prediction interval, the
values of normalized sample means  ̅ = −104.702321⋅ 103. The   /2, = 2.594161 for
significant level  = 0.01 and 270 degrees of freedom;    = 10577.32936. The inverse matrix of
(
          <xref ref-type="bibr" rid="ref11">11</xref>
          ) consists of one element   −1 = 2.46245 ⋅ 10−12. We have randomly chosen 30 of 286 data
rows from the training sample to compare actual ( ) and estimated ( ̂ ) KLOC values by the
regression models and their lower bound (LB) and upper bound (UB) prediction intervals (Table 2).
        </p>
      </sec>
      <sec id="sec-3-18">
        <title>The prediction results and prediction intervals of nonlinear regressions The one-factor nonlinear regression model on the basis on Johnson SB</title>
        <p>#
1
2
3
4
19.818
̂

40.095
208.967</p>
        <p>Prediction intervals of the one-factor and three-factor nonlinear regression models are
compared by given formula</p>
        <p>∑ =1[  1 −   1 ]</p>
        <p>
          ∑ =1[  3 −   3 ]) ⋅ 100%,
(
          <xref ref-type="bibr" rid="ref15">15</xref>
          )
where   3 ,   3 - i-th prediction intervals of the three-factor nonlinear regression model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
and   1 ,   1 - i-th prediction intervals of the one-factor nonlinear regression model. The 
value indicates the three-factor nonlinear regression interval is 52.78% smaller than the one-factor
nonlinear regression interval on the basis of the training sample.
        </p>
        <p>To make visualization of the comparison of the one-factor nonlinear model and the three-factor
nonlinear model accuracy we sorted all 571 data rows in ascending order by actual KLOC values of
the multidimensional dataset with KLOC estimates and prediction intervals of the target models
and first 400 records were chosen by x-axis and 150 KLOC max value was chosen by y-axis due to
scaling. The KLOC estimates and prediction intervals were smoothed by moving averages
technique on a basis of 5 values regarding the centric value of them. The obtained charts are
demonstrated on Figure 3.</p>
      </sec>
      <sec id="sec-3-19">
        <title>The graphic representation displays that the estimates of the three-factor nonlinear model (9)</title>
        <p>are closer to the actual KLOC values than the estimates of the one-factor nonlinear model.</p>
      </sec>
      <sec id="sec-3-20">
        <title>The prediction interval width is smaller in comparison with the interval of the one-factor</title>
        <p>nonlinear model on the basis of the same normalizing transformation.</p>
      </sec>
      <sec id="sec-3-21">
        <title>The values of the quality criteria indicate good model accuracy, the high values of for the</title>
        <p>testing sample confirms the model is not overfitted that confirms robustness and reliability.  2,
 and  (0.25) values of training and testing samples are close to each other thus
confirm the samples have a high level of representativity of the population of the open-source</p>
      </sec>
      <sec id="sec-3-22">
        <title>Java-software [13].</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Discussion</title>
      <sec id="sec-4-1">
        <title>We obtained the three-factor nonlinear regression model for early estimation of KLOC in Java</title>
        <p>software using appropriate techniques for constructing non-linear regression models on the basis
of multivariate normalizing Johnson SB family transformation. The technique is chosen due to the
non-Gaussian distribution of four-dimensional data of the code metrics according to Mardia
criteria and the regression residuals according to the Chi-Square test.</p>
        <p>
          Table 1 and models comparison in [13] confirms that the results of this research demonstrate
the effectiveness and accuracy of the proposed three-factor nonlinear regression model for early
KLOC estimation in comparison with existing linear and non-linear regression models by quality
criteria  2,  and  (0.25). The criteria for the model are exceeded expectation  2 ≥
0.75,  ≤ 0.25 and  (0.25) ≥ 0.75 for both training and testing samples. The values
are  2 = 0.9073,  = 0.1645,  (0.25)0.7692 for training sample and  2 = 0.9016,
 = 0.1617,  (0.25) = 0.8175 for the testing sample. Cross-Validation technique
increases guarantees of robustness of the model. Moreover, the width of prediction interval (
          <xref ref-type="bibr" rid="ref10">10</xref>
          ) of
the model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) is 52.78% smaller (Table 2 and Figure 3) than the interval of one-factor nonlinear
regression model which allows to increase accuracy of KLOC estimation taking in account positive
and negative cases.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>The advantages of the proposed model (9) includes the possibility of early KLOC estimation of</title>
        <p>
          Java-software using three code metrics such as total quantity of classes (CLASS), total number of
unique method calls in classes (RFC) and average value of visible methods per class (aVMQ) that
could be obtained on the early stages of the project planning from UML class diagrams. The
disadvantages of the proposed model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) include the following restrictions for the predictors: the
interval for CLASS is from 25 to 11147, the interval for RFC is from 45 to 117847 and the interval
for aVMQ is from 1.5893 to 85.92. Despites on the usage of the big sample of code metrics, the
proposed model is built only on a basis of open-source Java-software projects from GitHub.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>However, the model achieved high quality level according to quality criteria [21] and Table 2</title>
        <p>
          confirms that the KLOC estimates by the three factor model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) is closer to the actual KLOC values
than the one-factor model, there are more directions for further research. The model (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) could be
improved by adding more metrics to reveal characteristics of inheritance relation between classes
and quantity of visible fields of the classes. But the adding of more predictors could cause
multicollinearity problems that lead us to converting or avoiding usage of the predictors.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>The obtained results confirm that the constructed three-factor nonlinear model (9) improves confidence of KLOC estimation of the Java-software.</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <sec id="sec-5-1">
        <title>The obtained three-factor nonlinear regression model and prediction intervals solves important</title>
        <p>problem of Java-software size estimation at the early stages of software development project
planning using UML class diagram metrics: total quantity of classes (CLASS), total quantity of
responses for class (RFC) and average value of public and protected methods per class (aVMQ).</p>
        <p>The scientific novelty of the obtained results is that the three-factor nonlinear regression
model is improved in comparison with existing models and equation for KLOC estimation of
Javasoftware; Multivariate Johnson SB family normalizing transformation is firstly used for model
construction using CLASS, RFC and aVMQ multidimensional dataset; the three-factor nonlinear
regression model is firstly built using large sample size ( = 286) and the robustness and
reliability verified by using the testing sample with the similar size. The model, in comparison with
other nonlinear regression models, has higher value of coefficient of determination  2, a lower
value of the mean relative error  , higher value of the percentage of prediction of the relative
error level  (0.25) for both training and testing samples and the prediction interval is smaller
in comparison with one-factor nonlinear regression models. The obtainer quality criteria values
confirm the training and the testing samples are close to each other thus verifies the samples have
a high level of representativity of the population of the open-source Java-software</p>
        <p>The practical significance of the obtained results allows us to recommend the built model
for use in practice. The proposed model is implemented as a software product that can be used by
project managers for Java software development effort estimation at early stages of project
planning to reduce risks and cost.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Prospects for further research may include extending a set of independent factors, the usage of other multivariate normalizing transformations and extending wider data sets to build a nonlinear regression model.</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Acknowledgements</title>
      <p>The authors would like to express their sincere gratitude to the Doctor of Technical Sciences,
Professor Sergiy B. Prykhodko, NUOS, Mykolaiv, Ukraine for his support in applying the
mathematical apparatus for the research and to PhD. Maurício Aniche, Assistant Professor in
Software Engineering at TU Delft, for providing valuable Java code static analysis tools CK.</p>
      <sec id="sec-6-1">
        <title>Computer Engineering (TCSET), Lviv-Slavske, Ukraine, 2020, pp. 407-410.</title>
        <p>doi:10.1109/TCSET49122.2020.235464.
[18] P. M. Stanfield, J. R. Wilson, G. A. Mirka, N. F. Glasscock, J. P. Psihogios, J. R. Davis,</p>
      </sec>
      <sec id="sec-6-2">
        <title>Multivariate input modeling with Johnson distributions, in: Proceedings of the 28th Winter</title>
        <p>simulation conference WSC 96, Coronado, CA, USA, December 8-11, 1996, ed. S. Andradóttir,</p>
      </sec>
      <sec id="sec-6-3">
        <title>K.J. Healy, D.H. Withers and B.L. Nelson, IEEE Computer Society Washington, DC, USA, pp.</title>
        <p>1457-1464.
[19] S. B. Prykhodko, Simulation of gaussian random variables using Johnson SU transform,</p>
      </sec>
      <sec id="sec-6-4">
        <title>Informatics and mathematical methods in simulation 1 5 (2015) 92-97.</title>
        <p>[20] K. V. Mardia, Measures of multivariate skewness and kurtosis with applications, Biometrika 57
(1970) 519 530. doi:10.1093/biomet/57.3.519.
[21] D. Port, M. Korte, Comparative studies of the model evaluation criterions MMRE and PRED in
software cost estimation research, in: Proceedings of the 2nd ACM-IEEE International</p>
      </sec>
      <sec id="sec-6-5">
        <title>Symposium on Empirical Software Engineering and Measurement. ACM, New York, USA,</title>
        <p>2008, pp. 51 60. doi:10.1145/1414004.1414015</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trendowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jeffery</surname>
          </string-name>
          ,
          <article-title>Software Project Effort Estimation, Foundations and Best Practice Guidelines for Success</article-title>
          , in: Springer International Publishing,
          <year>2014</year>
          . doi: https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -03629-8
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Munialo</surname>
          </string-name>
          ,
          <article-title>A Review of Agile Software Effort Estimation Methods</article-title>
          ,
          <source>International Journal of Computer Applications Technology and Research. Association of Technology and Science</source>
          <volume>5</volume>
          (
          <year>2016</year>
          )
          <fpage>612</fpage>
          618. doi:
          <volume>10</volume>
          .7753/IJCATR0509.1009
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Valerdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. W.</given-names>
            <surname>Boehm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Reifer</surname>
          </string-name>
          ,
          <article-title>COSYSMO: A constructive systems engineering cost model coming of age</article-title>
          ,
          <source>in: Proc. INCOSE Int. Symp., volume. 13 of No. 1</source>
          , Wiley, Hoboken, NJ, USA,
          <year>2003</year>
          , pp.
          <fpage>70</fpage>
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>S. McConnel</surname>
          </string-name>
          ,
          <article-title>Software Estimation: Demystifying the Black Art</article-title>
          , Microsoft Press, Redmond, Washington, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>TIOBE</surname>
          </string-name>
          ,
          <string-name>
            <given-names>TIOBE</given-names>
            <surname>Index</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://www.tiobe.com/tiobe-index/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Oracle</surname>
          </string-name>
          , Java,
          <year>2024</year>
          . URL: https://www.oracle.com/my/java/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>The</given-names>
            <surname>Standish</surname>
          </string-name>
          <string-name>
            <surname>Group</surname>
          </string-name>
          ,
          <source>Chaos report</source>
          <year>2015</year>
          ,
          <year>2015</year>
          . URL: https://standishgroup.com/sample_research_files/CHAOSReport2015-Final.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H. B. K.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Estimating LOC for information systems from their conceptual data models</article-title>
          ,
          <source>in: Proceedings - International Conference on Software Engineering</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>321</fpage>
          -
          <lpage>330</lpage>
          . doi:
          <volume>10</volume>
          .1145/1134285.1134331.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H. B. K.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Conceptual Data Model-Based Software Size Estimation for Information Systems, ACM Transactions of Software Engineering and Methodology</source>
          <volume>19</volume>
          (
          <year>2009</year>
          ). doi:
          <volume>10</volume>
          .1145/1571629.1571630.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.B.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <article-title>A nonlinear regression model for estimation of the size of Java enterprise information systems software</article-title>
          ,
          <source>Modeling and Information Technologies</source>
          <volume>85</volume>
          (
          <year>2018</year>
          )
          <fpage>81</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Makarova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. O.</given-names>
            <surname>Kudin</surname>
          </string-name>
          ,
          <article-title>Constructing the non-linear regression model for size estimation of web-applications implemented in Java</article-title>
          ,
          <source>Herald (Kherson National Technical University)</source>
          ,
          <volume>69</volume>
          (
          <year>2019</year>
          )
          <fpage>145</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          , T. G. Smykodub,
          <article-title>Four-factor non-linear regression model to estimate the size of open source Java-based applications</article-title>
          , Scientific Notes of Taurida National V.I. Vernadsky University. Series: Technical Sciences,
          <volume>70</volume>
          (
          <year>2020</year>
          )
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          . doi:
          <volume>10</volume>
          .32838/
          <fpage>2663</fpage>
          -
          <lpage>5941</lpage>
          /
          <year>2020</year>
          .2-
          <issue>1</issue>
          /
          <fpage>25</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O. S.</given-names>
            <surname>Oriekhov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Farionova</surname>
          </string-name>
          ,
          <article-title>Mathematical models for the size estimating of JAVA applications</article-title>
          ,
          <source>Visnyk of Kherson National Technical University 89</source>
          <volume>2</volume>
          (
          <year>2024</year>
          )
          <fpage>196</fpage>
          -
          <lpage>203</lpage>
          . doi:
          <volume>10</volume>
          .35546/kntu2078-
          <fpage>4481</fpage>
          .
          <year>2024</year>
          .
          <volume>2</volume>
          .28.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Jia</surname>
          </string-name>
          , W. Qiu,
          <source>Research on an Ensemble Classification Algorithm Based on Differential Privacy, in: IEEE Access</source>
          <volume>99</volume>
          (
          <year>2020</year>
          ). DOI:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2020</year>
          .2995058
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>I.</given-names>
            <surname>Olkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Sampson</surname>
          </string-name>
          ,
          <article-title>Multivariate Analysis: Overview</article-title>
          , in N. J.
          <string-name>
            <surname>Smelser</surname>
            ,
            <given-names>P. B.</given-names>
          </string-name>
          <string-name>
            <surname>Baltes</surname>
          </string-name>
          ,
          <source>International encyclopedia of social &amp; behavioral sciences (eds.)</source>
          1st edn.,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          , Pergamon,
          <year>2001</year>
          , pp.
          <fpage>10240</fpage>
          <lpage>10247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          , N. Prykhodko,
          <article-title>Mathematical Modeling of Non-Gaussian Dependent Random Variables by Nonlinear Regression Models Based on the Multivariate Normalizing Transformations</article-title>
          , in: S. Shkarlet,
          <string-name>
            <given-names>A.</given-names>
            <surname>Morozov</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Palagin,
          <source>Mathematical Modeling and Simulation of Systems (MODS'</source>
          <year>2020</year>
          ).
          <source>Advances in Intelligent Systems and Computing</source>
          , volume
          <volume>1265</volume>
          <source>of MODS</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>166</fpage>
          -
          <lpage>174</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -58124-4_
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Makarova</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Pukhalevych</surname>
          </string-name>
          ,
          <article-title>Outlier Detection in NonLinear Regression Analysis Based on the Normalizing Transformations</article-title>
          ,
          <source>in: 2020 IEEE 15th International Conference on Advanced Trends in Radioelectronics, Telecommunications and</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>