<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Nonlinear regression models for software size estimation of Data Science and Machine Learning Java-applications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleksandr Oriekhov</string-name>
          <email>oleksandr.oriekhov@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetyana Farionova</string-name>
          <email>tetyana.farionova@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liubava Chernova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lyudmila Chernova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mykhaіlo Vorona</string-name>
          <email>mykhailo.vorona@nuos.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Admiral Makarov National University of Shipbuilding, Ukraine</institution>
          ,
          <addr-line>Heroes avenue, 9, Mykolaiv, 54007</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper introduces the usage of regression models and equations for Data Science and Machine Learning Java applications size estimation. Size estimation of applications plays one of the key planning tasks at the early stages of project planning for the successful implementation of software development projects. Application size estimation is used to predict software development effort estimation using parametric models such as COCOMO, COCOMO II, etc. The aim of the study is to increase the reliability and accuracy of size estimation of Data Science and Machine Learning Java applications at the early stage of software project planning using class diagram metrics by building a nonlinear regression model. The object of research is the process of size estimation for open-source Data Science and Machine Learning Java applications. The subject of the study is the regression equations and nonlinear regression models to estimate the software size. To achieve this goal, we analyzed and compared the existing mathematical regression models and equations for Java applications size estimating on the sample of code metrics information from open-source Java applications of Data Science and Machine Learning. Proven the necessity of building the the three-factor nonlinear regression model for estimating the software size of Data Science and Machine Learning Java applications on the basis of the decimal logarithm normalizing transformation using the software code metrics such as the total quantity of classes, the total visible methods quantity, and the average fields quantity per class. The obtained nonlinear regression model is compared with the existing models by the regression models quality criteria such as the determination coefficient, mean magnitude of relative error and the percentage of prediction of the relative error level 0.25. The comparison confirms increasing the accuracy of software size estimation using the given sample by the obtained nonlinear regression model.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Software size estimation</kwd>
        <kwd>nonlinear regression model</kwd>
        <kwd>normalizing transformation</kwd>
        <kwd>Java</kwd>
        <kwd>Data Science</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>non-Gaussian data</kwd>
        <kwd>decimal logarithm normalizing1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Software development effort estimation is one of the significant indicators of budget,
resources and duration planning of any project for software development business.
Reliable estimates of software development effort provide valuable information at the
early stages of project planning, and it helps to take into account risks, to recognize
planning gaps and to increase the efficiency of the development process. If the estimates
are close to the targets, then the plans can assume less risks. It is impossible to create
modern software products without integrating Data Science and Machine Learning
technologies, because the information technologies market requires them to integrate
these technologies to obtain a competitive advantage and offer better software solutions
and services to their users. The capitalization and active growth of the Machine Learning
market indicates the prospects for the development of this area of the IT industry. Thus,
the expected capitalization growth for 2030 is 528.1B USD, compared to 72.17B in 2022.
The appearance of large language models in 2022 and 2023 has caused a real revolution in
the direction of Data Science and Machine Learning in the IT and other industries [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. As a
result, it has led to the active implementation and integration of Data Science and Machine
Learning technologies and developments in the IT infrastructures of various industries,
such as healthcare, education, public sector, finance and economics, e-commerce,
aerospace, bioinformatics, etc [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Java programming language is one of the most popular
in the world [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and it is widely used for the development of software projects in various
areas, ranging from web applications and util application software to automotive or
information systems. The infrastructure of various companies is built using this
programming language. Since Java is one of the most demanded programming languages,
the size estimation is an important task of the software project management life cycle for
Java applications that use or implement Data Science and Machine Learning including such
sub-category as informational system software.
      </p>
      <p>
        The information about the software size for Data Science and Machine Learning Java
applications allows us to predict the software development effort estimation at the early
stages of project planning using well-known parametric models COCOMO, COCOMO II,
SLIM, SEER-SEM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The parametric models use total value lines of code to estimate the
effort. The accuracy of software effort estimation allows optimizing the resources
management in a rational way and, as result, allows to reduce cost of the software
development. It is valuable in the case of large projects development, like informational
systems, because these kinds of projects have higher risk to meet failures and issues.
Modern concurrency competitions lead software companies to work on integrating Data
Science and Machine Learning technologies into their software infrastructure or build
brand new complex solutions. From one side, large projects development is based on
traditional project management methodologies like software development life cycle and
waterfall, that makes most of the parametric models more suitable for software
development effort estimation than agile methodologies. From the other side, it’s expected
that large projects have a bigger code base on final stages of development because of the
complexity of working with complex and different data. The chaos report 2015 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] of The
Standish group proves that large and grand projects have much higher failure rates in
comparison with small, moderate or medium projects. It confirms obtaining a reliable
estimate of the code lines requires appropriate models for Java applications in the areas of
Data Science and Machine Learning.
      </p>
      <p>
        Both the linear [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ] and nonlinear regression equations and models [
        <xref ref-type="bibr" rid="ref10 ref11 ref8 ref9">8,9,10,11</xref>
        ] have
been built to estimate the size of open-source Java applications. The models depend on
certain metrics from conceptual data models based on a class diagram. The models are
based on quantitative metrics, such as the total number of classes, the number of methods
(visual, public, static, etc), the total number of class fields (private, public, protected,
visual), etc. and qualitative metrics, such as LCOM (lack of cohesion), RFC (response for
class), etc. Different combinations of quantitative and quality metrics can have an effect on
the reliability and accuracy of software size estimation. In addition, the models were built
on the sample variation that may have an insufficient representation level of code metrics
population for the Data Science and Machine Learning Java applications.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Review of the literature</title>
      <p>Nowadays, both the linear and nonlinear regression equations and models have been
developed to estimate the number code lines of Java applications and information
systems, depending on the metrics of the conceptual data model based on a class diagram.</p>
      <p>
        Thus, in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the three-factor linear regression equation for estimating the lines of code
of Java applications. It is based on the methods of multiple linear regression analysis using
the metrics of the total number of classes (CLASS), the total number of relationships
between classes (CBO) and the total quantity of class fields (TFQ) in the source code. The
paper [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] offers improved three-factor linear regression equations for estimating the lines
of code for large industrial information Java systems and open-source Java applications.
The model uses variables of the total number of classes (CLASS), the total number of
couplings between objects (CBO), and the average quantity of fields per class (aTFQ).
      </p>
      <p>
        The paper [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is devoted to improving the estimation of the lines of code for industrial
Java information systems and proposes a three-factor nonlinear regression model on the
basis of multivariate Johnson normalizing transformation for the SB family. It is
constructed on the basis of a four-dimension non-Gaussian sample dataset of software
code metrics. The dataset includes the number of classes (CLASS) X1, the total number of
couplings between classes (CBO) X2 and the average quantity of fields per class (aTFQ) X3
from the conceptual model of the application.
      </p>
      <p>
        In the paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the one-factor nonlinear regression model was built to estimate the
software size of Java web-applications. The mathematical model is constructed on the
basis of the Johnson SB family normalizing transformation. It proposes estimation of code
lines using the total quantity of classes (CLASS) metric as an independent factor X.
      </p>
      <p>
        In the paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the authors built the four-factor nonlinear regression model for
estimating the size of open source Java applications. The model was constructed on the
basis of the multivariate Johnson normalizing transform of the SB family. To estimate the
value of lines of code, it uses the metrics of the quantity of classes (CLASS) X1, the total
quantity of static methods (SMQ) X2, the total values sum of lack of cohesion of methods
(LCOM) X3 and the total quantity of unique method invocations in a class (RFC - Response
for a Class) X4.
      </p>
      <p>
        In the paper [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], the three-factor nonlinear regression model was constructed on the
basis of the Box-Cox normalization transformation to estimate the size of Data Science and
Machine Learning Java applications. It estimates lines of code variable using metrics of
kilo-quantity of classes (kCLASS) X1, kilo-quantity of visible class methods (kVMQ) X2 and
kilo-quantity of public class fields (kPFQ) X3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Formulation of the problem</title>
      <p>The analysis of the literature has shown that the linear equations and nonlinear regression
models exist for estimating Java-application code size. For the existing models, there is a
necessity to confirm the possibility of using to estimate the size of Data Science and
Machine Learning Java applications and, otherwise, there is a necessity to build a
regression model based on the metrics of Data Science and Machine Learning Java
applications for the size estimation.</p>
      <p>
        Let's make the assumption that the models [
        <xref ref-type="bibr" rid="ref10 ref11 ref6 ref7 ref8 ref9">6,7,8,9,10,11</xref>
        ] can be used to estimate the
size of Data Science and Machine Learning Java applications. To do this, it is necessary to
confirm the possibility of using these regression equations and models by comparing them
according to quality criteria for the sample of metrics of Data Science and Machine
Learning Java applications.
      </p>
      <p>
        In case of refutation of the assumption, it is necessary to build an appropriate
multivariate nonlinear regression model using proper normalization transformations and
compare it with the existing models [
        <xref ref-type="bibr" rid="ref10 ref11 ref6 ref7 ref8 ref9">6,7,8,9,10,11</xref>
        ] by quality criteria to improve the
estimation of the size of Data Science and Machine Learning Java applications.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Objectives of the study</title>
      <p>The aim of the study is to increase the reliability and accuracy of size estimation of Data
Science and Machine Learning Java applications at the early stages of software
development project planning using the metrics from a conceptual data model by building
a nonlinear regression model.</p>
      <p>The object of the study is the process of software size estimation of open-source Data
Science and Machine Learning Java applications.</p>
      <p>The subject of the study is regression equations and nonlinear regression models for
estimating the size of open-source Data Science and Machine Learning Java applications.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Materials and research methods</title>
      <p>
        To achieve the aim of the paper, it is necessary to analyze and compare the existing
mathematical models and equations for estimating the size of Java applications on the
sample of code metrics of open-source Data Science and Machine Learning Java
applications; to justify the necessity of a nonlinear regression model building for
estimating the size of Data Science and Machine Learning Java applications; to build a
three-factor nonlinear regression model for estimating the size of Data Science and
Machine Learning Java applications using quantitative code metrics, such as the total
quantity of classes, the total quantity of visible methods, the average fields quantity per
class; compare the obtained nonlinear regression model with the existing models
[
        <xref ref-type="bibr" rid="ref10 ref11 ref6 ref7 ref8 ref9">6,7,8,9,10,11</xref>
        ] using the quality criteria of regression models.
      </p>
      <p>There are different approaches for estimating the quantity of lines of code variable
using both linear and non-linear regression models. Typically, the software code metrics
have non-Gaussian distribution. It limits the ability to use linear models to estimate the
variable of lines of code. One of the theoretical conditions for the usage of linear
regression models says that the regression residuals ε should have Gaussian distribution.</p>
      <p>The following approaches can be used to normalize data: decimal logarithm, square
root,</p>
      <p>Box-Cox
transformation,</p>
      <sec id="sec-5-1">
        <title>Johnson</title>
        <p>transformation,
etc.</p>
        <p>The
normalizing
transformations allow the construction of linear regression models based on normalized
data with their further inverse transformation into nonlinear regression models.</p>
        <p>The following quality criteria are used for forecasting quality assessment of regression
models: the coefficient of determination  2, a mean magnitude of relative error MMRE and
percentage of prediction for magnitude of relative error (MRE) level 0.25 
(0.25).</p>
        <p>The MMRE criterion is defined as
where N - sample size and</p>
        <p>is the value of the magnitude of relative error for the
i-th line of data of a random variable.</p>
        <p>= 1 ∑</p>
        <p>=1    ,
   = |(  − ̂ )|.</p>
        <p />
        <p>
          The calculation of prediction percentage (PRED) for the magnitude of relative error
level 0.25

(0.25) = 1 ∑
  =1{0
1  
The
acceptable
values of 
≤ 0.25 and
(0.25) ≥ 0.75
for the
measurement of the regression models accuracy of prediction results. The coefficient of
determination (R2) value is acceptable if it is more or equals to 0.75 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          The nonlinear regression models construction techniques are proposed in the papers
[
          <xref ref-type="bibr" rid="ref13 ref14">13,14</xref>
          ]. It uses statistical techniques for detecting and discarding outliers in non-linear
regression
analysis of non-Gaussian
data
based
on the reciprocal normalizing
transformations. There are following steps to construct the nonlinear regression model.
        </p>
        <p>There is bijective normalizing transformation of a non-Gaussian random vector of the
the form
(1)
(2)
(3)
(4)
(5)
 ℎ</p>
        <p>≤ 0.25 .
 =  ( ),
 =  −1( ),
where k is number of factors (regressors or independent variables) and the inverse
where</p>
        <p>is a vector of bijective normalizing transformation functions,  =
The linear regression model for multivariate normalized data according to (4) will have
  =  ̂ +  =  ̂ 0 +  ̂1 1 +  ̂ 2 2+. . . + ̂    +  .</p>
        <p>where ε is Gaussian random variable, ε ∼  (0, σε2 );  ̂ 0,  ̂1,  ̂ 2, . . . ,  ̂  - estimators for
parameters of the linear regression model (6). The estimators are calculated by the least
square method. To build a nonlinear regression model, the inverse transformation (5) is
applied to the linear regression model (6). The nonlinear regression model will have the
form</p>
        <p>=  −1( ̂ +  ) =  −1(  ̂ 0 +  ̂1 1 +  ̂ 2 2+. . . + ̂    +  ).</p>
        <p>
          The choice of factors in regression models should take into account the level of
multicollinearity, because a high level of correlation between regression factors increases
the sensitivity of the model to random changes in the data. The multicollinearity level is
determined by the variance inflation factors (VIFs) of the independent variables. For a
multiple linear regression model with k factors Xi , i=1 2,..., k, the VIFs are represented by
the diagonal elements of the inverse correlation k×k matrix. If the value of the VIF
coefficient exceeds 10 (the threshold value), a high level of
multicollinearity exists
between the independent variables [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>normalization transformation is chosen to normalize
multivariate data
 =</p>
        <p>10( ),
where Х takes the values  ,  1,  2, . . . ,   respectively.</p>
        <p>
          Normality distribution of multivariate data is checked with Mardia test [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The test is
based on measurement of multivariate skewness (β1,k) and kurtosis(β2,k) of the sample.
sample has normal distribution
        </p>
        <p>The test statistics for  1, has the form

6
 1, ≤  2 ,
(6)
(7)
(8)
(9)
(10)
(11)</p>
        <p>(12)
(13)
 1, = 1 ∑</p>
        <p>2
 2, = 1 ∑</p>
        <p />
        <p>=1 ∑ =1[(  −  )  −1(  −  )] ,</p>
        <p>3
 =1[(  −  )  −1(  −  )] ,</p>
        <p>2

  = 1 ∑
 =1(  −  )(  −  ) ,

where X is k-dimensional vector of random variable, 
= ( 1,  2, . . . ,   ) and SN is a
biased sample variance matrix of multivariate random variable X. it is given by
where  is a means vector of independent variable of the sample,  = ( 1,  2, . . . ,   ) .
Accordingly to Mardia test, the following conditions should be verified to confirm if the
where  2 is approximated Chi-Square distribution with  ( + 1)( + 2)/6 degrees of
freedom and α is a significant level (Accordingly to the proposed technique, α=0.005).</p>
        <p>For  2, , the test statistics is 1-α quantile of the normal distribution 
with the
mathematical expectation  =  ( + 2) and variance  2 = 8 ( + 2)/</p>
        <p>2, ≤  1− ( ,  2) .</p>
        <p>
          Outliers of the sample are detected using the techniques proposed in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. It is based on
outliers detection using the squared Mahalanobis distance and prediction intervals of the
regression model. According to the methods the outlier should be removed from the
sample. The squared Mahalanobis distance is elements on the main diagonal of the  2
matrix of size N x N
        </p>
        <p>where SN is a biased sample variance matrix (11), Z is a normally distributed random
variable. According to the technique the elements of the main diagonal of  2,  = 1,2, . . . 
matrix are detected as outliers if the values are exceeded the threshold value of the
ChiSquare  2 distribution quantile for the significant level - α. These outliers should be
excluded from the sample.</p>
        <p>The prediction interval outlier detection technique is based on prediction intervals of
the nonlinear regression model. If the outliers are detected, ones are discarded from the
sample. For the nonlinear regression model (7), the prediction interval is given by
̂

where  ̂ is the upper and lower bounds of the prediction interval (PI) of the linear
regression (6),  −1 is the inverse normalizing transformation function of the linear
regression estimators. The linear regression prediction interval is given by
 ̂
1/2
,
where   /2, is a quantile of T-Student distribution with  =  −  − 1 degrees of
freedom and α/2 significant level;   2 =

1 
∑
 =1(   −  ̂  )2;   + is a vector of central
moments of independent variables of the sample which is given by { 1 −  1,  2 −
 2, . . . ,    −   } ;   ⬚ is k x k matrix
 2 = (  −  )  −1(  −  ),

(14)
(15)
(16)
(17)


= [  

  ],
where</p>
        <p>= ∑ −1(   −   )(   −   ), q,r = 1,2, …, k.</p>
        <p>It is recommended to detect and discard the outliers in an iterative way. Only one
outlier should be discarded per iteration and the nonlinear regression model building
should be started from the start for the new sample.</p>
        <p>The authors have collected the sample dataset of the code metrics of 74 Data Science
and</p>
        <p>Machine</p>
        <p>Learning</p>
      </sec>
      <sec id="sec-5-2">
        <title>Java</title>
        <p>applications
hosted
on
the</p>
        <p>
          GitHub
platform
(https://github.com). The following metrics were obtained using the CK static code
analysis tool (https://github.com/mauricioaniche/ck). The sample includes lines of code
of projects (KLOC), total quantity of application classes (CLASS), total quantity of unique
method calls in a class (RFC), total quantity of class method cohesion values (LCOM), total
quantity of static methods (SMQ), total quantity of visible methods (VMQ), total quantity
of class fields (TFQ), total quantity of public class fields (PFQ), and total coupling between
objects per project (CBO). The resulting metrics are shown in Table 1. These metrics
(excepting KLOC) can be obtained at an early stage of project planning from the
conceptual data model of the application.
6. Experiment
The regression models and equations [
          <xref ref-type="bibr" rid="ref10 ref11 ref6 ref7 ref8 ref9">6,7,8,9,10,11</xref>
          ] were tested by the multiple
coefficient of determination R2, the mean magnitude of relative error MMRE, and the
percentage of prediction for magnitude of relative error threshold is less than 0.25,
PRED(0.25).
        </p>
        <p>
          The result of the linear regression models [
          <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
          ] didn’t achieve reliable estimation of the
number of lines of code, which confirmed that the values of the regression models quality
criteria are exceeded the permissible range of values.
        </p>
        <p>
          The three-factor nonlinear regression model [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] for estimating the number of lines of
code (KLOC) of industrial information systems cannot be applied to the sample from Table
1, since the values of 53 data points do not meet the permissible range for applying the
normalizing transformation of this model.
        </p>
        <p>
          The one-factor nonlinear regression model [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and the four-factor nonlinear regression
model [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] have unsatisfactory values of MMRE, PRED(0.25) quality criterias. Thus, for
model [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] MMRE=0.9806 and PRED(0.25)=0.0, and for model [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] MMRE=0.8854 and
PRED(0.25)=0.0725. This indicates that these regression models cannot be used to
estimate the number of lines of code for Data Science and Machine Learning Java
applications.
        </p>
        <p>
          The test of the model [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] according to Table 1 has the following quality values
R2=0.8985, MMRE=0.2014, PRED(0.25)=0.66. As we can see, the value of PRED(0.25) is
less than 0.75, which indicates its insufficient quality level.
        </p>
        <p>Thus, it is required to build the three-factor nonlinear regression model (7) to estimate
the software size (in KLOC) for Data Science and Machine Learning Java applications using
decimal logarithm normalizing transformation.
6.2. Building the three-factor nonlinear regression model
To increase the reliability and accuracy of size estimation of Data Science and Machine
Learning Java applications, the three-factor nonlinear regression model is built on the
basis of decimal logarithm normalizing transformation (8) using the data sample from
Table 1. The model is based on the parameters of total quantity of classes (CLASS) X1, the
total quantity of visible methods (VMQ) X2, and the average quantity of fields per class
(aTFQ) X3.</p>
        <p>
          The model factors X1, X2 and X3 were tested for multicollinearity by the VIFs [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. For
factors X1, X2, and X3, the VIFs are equal to 6.1194, 6.4795, and 1.2016, respectively. This
indicates the absence of multicollinearity between the factors of the regression model.
        </p>
        <p>Let’s check the null hypothesis if the linear regression model can be applied for
software size estimation in the following form</p>
        <p>=  ̂ +  =  ̂ 0 +  ̂1 1 +  ̂ 2 2 +  ̂ 3 3 +  . (18)</p>
        <p>The factors of the three-factor linear regression model (18) are estimated using the
least squares method:  ̂ 0=-13.863468,  ̂1= 0.001468,  ̂ 2= 0.010861,  ̂ 3= 8.149336.</p>
        <p>The regression residuals ε are tested for normal distribution converging by Chi-Square
(χ2) test with the significance level α = 0.05 and degrees of freedom v = 4. The hypothesis
is rejected, because the test value χ2 = 163938.42 is greater than the quantile of
ChiSquare distribution, which equals 9.4877. Thus, the building of the nonlinear regression
model is justified.</p>
        <p>
          To build the nonlinear regression model, the sample from Table 1 is normalized using
the decimal logarithm normalization transformation (8). Following [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] multivariate
skewness (9) and kurtosis (10) are estimated for the normalized sample from Table 1. In
this case the values of skewness and kurtosis equal to 33.11 and 63.44 respectively, which
are higher than threshold values of Mardia test conditions (12) and (13). The test does not
confirm the normal distribution of the normalized sample. Using the square Mahalanobis
distance technique, 2 data rows are detected in the normalized data at row positions 37
and 68 which have squared Mahalanobis distance values equal to 62.42 and 17.78
respectively. The values exceed the threshold of the Chi-Square distribution quantile χ2 =
14.86 for the multivariate normalized data for the significance level α = 0.005. Only one
row with the highest squared Mahalanobis distance value is discarded per iteration,
because the technique detects more than one outlier per iteration but the rows may not be
outliers on further iterations.
        </p>
        <p>Next, a linear regression model is built for the normalized data</p>
        <p>=  ̂ +  =  ̂ 0 +  ̂1 1 +  ̂ 2 2 +  ̂ 3 3 +  ,</p>
        <p>The estimators for the parameters  ̂ 0,  ̂1,  ̂ 2,  ̂ 3 of the linear regression model (19) for
the normalized data are calculated using the least square method.</p>
        <p>Applying the inverse normalizing transformation to (8) the linear regression model
(19), the nonlinear regression model has the form</p>
        <p>=  −1(̂  +  ) =  −1(  ̂ 0 +  ̂1 1 +  ̂ 2 2 +  ̂ 3 3 +  ),</p>
        <p>
          For the model (20), the upper and lower bounds of the prediction interval (15) are
determined for the significance level α = 0.05. The data rows 8, 44, 72 are outside of the
prediction interval. According to the technique [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], the data rows are detected as outliers
and removed from the sample.
        </p>
        <p>The outliers are discarded iteratively with only one data row per iteration.</p>
        <p>The estimators for the parameters of the regression model (19) at the last iteration are:
 ̂ 0 = −1.855315,  ̂1 = 0.149619,  ̂ 2 = 0.822320,  ̂ 3 = 0.349549. The normal distribution
was tested and confirmed for the regression residual random variable ε using Chi-Square
test. The test value χ2 = 4.1411 is smaller than the value of the quantile of Chi-Square
distribution 9.4877 for the significance level α = 0.05 and degrees of freedom v = 4.
According to Maridia’s test, the normalized sample of 69 data rows has normal
distribution since multivariate skewness (12), which is 29.39, is less than the quantile of
Chi-Square distribution, which is 40.00 for 20 degrees of freedom and significance level α
= 0.005. Analogically kurtosis (13), which is 22.10, is less than quantile of normal
distribution, which is 28.30 for the mean of 24.</p>
        <p>By
applying
the
inverse
transformation
to
decimal logarithm
normalized
transformation (7), the three-factor nonlinear regression model (20) has the form
 ̂ = 10 −1.855315 10.149619 20.822320 30.349549.
(21)</p>
        <p>The quality criteria of the nonlinear regression model (21) have the following values of
indicators: R2=0.8937, MMRE=0.1867, PRED(0.25)=0.7703. This confirms the high quality
of the obtained model in comparison with the existing models. The regression model (21)
leads to increase the reliability and accuracy of the number of code lines estimating for</p>
      </sec>
      <sec id="sec-5-3">
        <title>Data Science and Machine Learning Java applications. (19) (20)</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Conclusion</title>
      <p>An important problem of increasing the reliability and accuracy of estimating the size of
Data Science and Machine Learning Java applications is solved.</p>
      <p>The scientific novelty of the obtained results is that the nonlinear regression model
for multivariate non-Gaussian data is improved by constructing a three-factor nonlinear
regression model for estimating the size of the code lines of Data Science and Machine
Learning Java applications based on the decimal logarithm normalizing transformation.
This model has a higher value of the multiple coefficient of determination R2, a lower value
of the mean relative error MMRE, and a higher value of the percentage of prediction of the
relative error level PRED(0.25) in comparison to the existing regression models and
equations.</p>
      <p>The practical significance of the obtained results is that the software that implements
the built model was developed using the Kotlin programming language and the Apache
Math3 mathematical package. The experimental results allow us to recommend the built
model for use in practice. The model can reduce risks and cost of software development
for Data Science and Machine Learning Java-applications or information systems.</p>
      <p>Prospects for further research may include the usage of other multivariate
normalizing transformations and data sets to build a nonlinear regression model.</p>
    </sec>
    <sec id="sec-7">
      <title>8. Acknowledgements</title>
      <p>The authors would like to express their sincere gratitude to the Doctor of Technical
Sciences, Professor Sergiy B. Prikhodko, NUOS, Mykolaiv, Ukraine for his support in
applying the mathematical apparatus for the research and Doctor of Technical Sciences,
Professor Serhii K. Chernov, NUOS, Mykolaiv, Ukraine for the offer and motivation to
participate in the conference.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] G. Press, Top Machine Learning Statistics to know [2024], What'stheBigData.com,
          <year>2023</year>
          . URL: https://whatsthebigdata.com/top-machine-learning-statistics/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chinthamu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karukuri</surname>
          </string-name>
          ,
          <article-title>Data Science and Applications</article-title>
          ,
          <source>Journal of Data Science and Intelligent Systems</source>
          , vol.
          <volume>00</volume>
          ,
          <year>2023</year>
          , doi:10.47852/bonviewJDSIS3202837.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>TIOBE</surname>
          </string-name>
          ,
          <string-name>
            <given-names>TIOBE</given-names>
            <surname>Index</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://www.tiobe.com/tiobe-index/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Munialo</surname>
          </string-name>
          ,
          <source>A Review of Agile Software Effort Estimation Methods</source>
          , volume
          <volume>5</volume>
          of
          <source>International Journal of Computer Applications Technology and Research. Association of Technology and Science</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>612</fpage>
          -
          <lpage>618</lpage>
          . doi:
          <volume>10</volume>
          .7753/IJCATR0509.1009
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>The</given-names>
            <surname>Standish</surname>
          </string-name>
          <string-name>
            <surname>Group</surname>
          </string-name>
          ,
          <source>Chaos report</source>
          <year>2015</year>
          ,
          <year>2015</year>
          . URL: https://standishgroup.com/sample_research_files/CHAOSReport2015-Final.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H. B. K.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Estimating LOC for information systems from their conceptual data models</article-title>
          ,
          <source>Proceedings - International Conference on Software Engineering</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>321</fpage>
          -
          <lpage>330</lpage>
          . doi:
          <volume>10</volume>
          .1145/1134285.1134331.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H. B. K.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <source>Conceptual Data Model-Based Software Size Estimation for Information Systems</source>
          , volume
          <volume>19</volume>
          <source>of ACM Transactions of Software Engineering and Methodology</source>
          ,
          <year>2009</year>
          , doi:10.1145/1571629.1571630.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.B.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <article-title>A nonlinear regression model for estimation of the size of Java enterprise information systems software</article-title>
          , volume
          <volume>85</volume>
          <source>of Modeling and Information Technologies</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>81</fpage>
          -
          <lpage>88</lpage>
          . URL: http://nbuv.gov.ua/UJRN/Mtit_
          <year>2018</year>
          _
          <volume>85</volume>
          _
          <fpage>14</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Makarova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. O.</given-names>
            <surname>Kudin</surname>
          </string-name>
          ,
          <article-title>Constructing the non-linear regression model for size estimation of web-applications implemented in Java</article-title>
          , volume
          <volume>69</volume>
          of
          <string-name>
            <surname>Herald</surname>
          </string-name>
          (Kherson National Technical University),
          <year>2019</year>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>153</lpage>
          . URL: http://eir.nuos.edu.ua/handle/123456789/4443
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          , T. G. Smykodub,
          <article-title>Four-factor non-linear regression model to estimate the size of open source Java-based applications</article-title>
          , volume
          <volume>70</volume>
          of Scientific Notes of Taurida National V.I. Vernadsky University. Series: Technical Sciences,
          <year>2020</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          . doi:https://doi.org/10.32838/
          <fpage>2663</fpage>
          -
          <lpage>5941</lpage>
          /
          <year>2020</year>
          .2-
          <issue>1</issue>
          /
          <fpage>25</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>O. S.</given-names>
            <surname>Oriekhov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Farionova</surname>
          </string-name>
          ,
          <article-title>Three-factor nonlinear regression model for estimating the size of Data Science and Machine Learning projects created using the JAVA programming language</article-title>
          , volume
          <volume>4</volume>
          of ITMAS - 2023: Information Technologies: Models, Algorithms, Systems, Mykolaiv: NUOS, Ukraine,
          <year>2023</year>
          , pp.
          <fpage>45</fpage>
          -
          <lpage>47</lpage>
          . URL: https://itconf.nuos.edu.ua/2023/proceedings/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Port</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Korte</surname>
          </string-name>
          ,
          <article-title>Comparative studies of the model evaluation criterions MMRE and PRED in software cost estimation research</article-title>
          ,
          <source>Proceedings of the 2nd ACM-IEEE International Symposium on Empirical Software Engineering and Measurement</source>
          . ACM, New York,
          <year>2008</year>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .1145/1414004.1414015
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          , N. Prykhodko,
          <article-title>Mathematical Modeling of Non-Gaussian Dependent Random Variables by Nonlinear Regression Models Based on the Multivariate Normalizing Transformations</article-title>
          , in S. Shkarlet,
          <string-name>
            <given-names>A.</given-names>
            <surname>Morozov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Palagin</surname>
          </string-name>
          , volume
          <volume>1265</volume>
          <source>of Mathematical Modeling and Simulation of Systems (MODS'</source>
          <year>2020</year>
          ).
          <source>Advances in Intelligent Systems and Computing</source>
          , volume
          <volume>1265</volume>
          <source>of MODS</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>166</fpage>
          -
          <lpage>174</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -58124-4_
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Prykhodko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Makarova</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Pukhalevych</surname>
          </string-name>
          ,
          <article-title>Outlier Detection in Non-Linear Regression Analysis Based on the Normalizing Transformations</article-title>
          , in
          <source>2020 IEEE 15th International Conference on Advanced Trends in Radioelectronics</source>
          , Telecommunications and Computer Engineering (TCSET),
          <source>Lviv-Slavske, Ukraine</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>407</fpage>
          -
          <lpage>410</lpage>
          . doi:
          <volume>10</volume>
          .1109/TCSET49122.
          <year>2020</year>
          .
          <volume>235464</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>I.</given-names>
            <surname>Olkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Sampson</surname>
          </string-name>
          ,
          <article-title>Multivariate Analysis: Overview</article-title>
          , in N. J.
          <string-name>
            <surname>Smelser</surname>
            ,
            <given-names>P. B.</given-names>
          </string-name>
          <string-name>
            <surname>Baltes</surname>
          </string-name>
          ,
          <source>International encyclopedia of social &amp; behavioral sciences (eds.)</source>
          1st edn.,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          , Pergamon,
          <year>2001</year>
          , pp.
          <fpage>10240</fpage>
          -
          <lpage>10247</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K. V.</given-names>
            <surname>Mardia</surname>
          </string-name>
          ,
          <article-title>Measures of multivariate skewness and kurtosis with applications</article-title>
          , volume
          <volume>57</volume>
          of Biometrika,
          <year>1970</year>
          , pp.
          <fpage>519</fpage>
          -
          <lpage>530</lpage>
          . doi:
          <volume>10</volume>
          .1093/biomet/57.3.519.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>