<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Productivity of Software Enhancement Projects: an Empirical Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Università degli Studi dell'Insubria</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Varese</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy luigi.lavazza@uninsubria.it</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DPO Srl.</institution>
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Hangzhou Dianzi University.</institution>
          <addr-line>Hangzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Background. Having a correct, although approximate, knowledge of software development productivity is clearly important. In some environments, the belief that software enhancement projects are characterized by higher productivity than new software development has emerged. Aim. We want to understand whether the mentioned belief is rooted on solid bases or is due to some cognitive biases. Method. An empirical study was performed, analyzing the data from a large dataset that collects data from real-life projects. Several statistical methods were used to evaluate the unitary cost (i.e., the cost per Function Point) of enhancement projects and new developments. Results. Our analyses show that-contrary to some popular beliefs-software enhancement costs more than new software development, at least for projects greater than 300 Function Points. Conclusions. Project managers and other stakeholders interested in the actual cost of software should reject ill-based evaluations that the productivity of software enhancement is greater than new software development. More generally, objective evaluations based on the analysis of representative data should be preferred to evaluations affected by cognitive biases.</p>
      </abstract>
      <kwd-group>
        <kwd>Function Point Analysis</kwd>
        <kwd>Functional Size Measurement</kwd>
        <kwd>Software measurement</kwd>
        <kwd>Software Development Productivity</kwd>
        <kwd>Software Maintenance</kwd>
        <kwd>Software Enhancement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Software cost models are relevant to the market since they influence budget
allocation on individual projects, tenders, contracts and finally they affect the quality of
customer-supplier relationships. Since the number of factors that may influence, at various
levels of impact, the productivity in delivering software systems is huge, it is
unavoidable to make some assumptions and to focus on a subset of variables considered as
significant. Assumptions may be based on direct personal experiences, on common
sense or on empirical evidence derived from collected data. Sometimes the experts that
build software cost models are induced to introduce in models some biases that may
affect the quality of correlation among the considered variables. This is done in absence
of awareness due to the cognitive biases effect.</p>
      <p>Specifically, we have observed the tendency to consider the enhancement of existing
software as less demanding–in terms of total effort–than the development of new
software. Since such beliefs can have quite relevant consequences (e.g., setting unrealistic
prices for software enhancement contracts), it is of great importance that we show if
the belief is rooted on solid bases, or it is affected by cognitive biases.</p>
      <p>In this paper, we apply the empirical method of testing the assumptions using
adequate data to address the problem of evaluating the actual effort of software
enhancement, especially compared to the effort of developing new software. In this way, we
contribute to reduce or inhibit the impact of cognitive biases on software cost models.</p>
      <p>The paper is organized as follows. Section 2 presents the motivations of the work
and highlights the importance of getting objective evaluations of the productivity of
software enhancement. Section 3 describes the empirical study through which we
derive objective effort evaluations. In Section 4 we discuss the threats to the validity of
the empirical study. Section 5 accounts for related work. Finally, Section 6 draws some
conclusions and outlines future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Motivations and Goals</title>
      <p>
        Cognitive sciences have shown that even the finest expert may be affected by
cognitive biases [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Intuition, in software engineering, is often blurred by prejudice,
confirmation bias, overconfidence, group-thinking, availability bias, framing [1817].
      </p>
      <p>Prejudice is a pre-defined cause-effect relationship, based on hyper-generalizations,
that does not need a test, validation, adaptation in order to be considered true by its
“owner”.</p>
      <p>Confirmation bias is the tendency to pay undue attention to sources that confirm our
existing beliefs while ignoring sources that challenge our beliefs. Once we have
elaborated a theory, we tend to look for confirmations instead of unexplained facts.</p>
      <p>Overconfidence bias is the tendency to overestimate one’s skills and abilities. This
may lead to ignore some factors in the model only because we do not think that we may
personally be affected by that factor.</p>
      <p>Group-thinking is a psychological phenomenon that occurs within a group of people
in which the desire for harmony or conformity in the group results in an irrational or
dysfunctional decision-making outcome, by inhibiting critical thinking to avoid
conflicts. This may happen even to people that do not know one each other, but that belong
to a community ruled by opinion leaders and recognized experts. It is very difficult for
anybody to swim against the flow.</p>
      <p>Availability bias is a tendency to allow information that is easier to recall unduly
influence preconceptions or judgments. We use in models the parameters that are the
easiest to be measured, regardless of the relevance they really have.</p>
      <p>Finally, the framing effect is the tendency to react differently to situations that are
fundamentally identical but presented (or framed) differently.</p>
      <p>All these cognitive biases may induce to build a model that may be poorly
representative of reality. The only way to know if this has happened is to “de-bias” the
reasoning with specific approaches possibly supported by empirical data.</p>
      <p>The wrong definition of a software cost model may be favored by situation in which
the outcome of a production process is influenced by two or more factors that may
affect the result in opposite ways and we are wrong in the identification of the resultant
of the vector composition.</p>
      <p>In this paper, we will focus on the productivity associated to a functional
enhancement project compared to a new development project. If we believe that adding,
changing and deleting a certain amount of functionalities in an existent application is much
easier than adding the same amount of functionalities in a new application (due to reuse,
fundamentally), we must expect a higher productivity in enhancement maintenance. On
the contrary, if we consider the difficulty of adding, changing and deleting
functionalities that are not well documented, were written with old technologies, by not
particularly skilled people in an unmanaged environment, then we will expect a lower
productivity with respect to new development. Which one is true or more often true in the
market? The right thing to do is to collect data and derive statistical driven inference.
The most frequent thing that has been done in the past is the application of the cognitive
biases illustrated before.</p>
      <p>
        The initial idea that changing and deleting functionalities should be associated to a
lower functional measure, if compared with the operation of adding new functions, was
supported by a NESMA (Netherlands Software Metrics Users Association) document
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This work proposed a way to consider the impact of change in enhancement
maintenance projects that is mainly associated to a reduction in size. The source of the
document is highly reliable and the approach was also adopted for a while by IFPUG
(International Function Point User’s Group), who gave it an “institutional” benediction.
This favored the group-thinking bias in the community of practitioners and subsequent
confirmation bias that led to prejudice. So, at least in Italy–which is one of the most
advanced country in using functional size in contracts–the sizing proposal was
transformed into a pricing proposal that assigned to a changed function a relevant discount
(50%) and to a deleted function a huge discount (90%). A framing bias was set up
presenting the situation as a case of pure profitable reuse. The idea was reasonable and
consequently “attractive” for practitioners. This approach, in turn, became a
“precedent” for all succeeding contracts (again a framing bias). No empirical evidences have
been given to support this approach but it became steady as a rock.
      </p>
      <p>The specific situation may have, and actually had, a huge impact on markets and
outsourcing contracts (in the order of millions of euro). Consider that the “discount
approach” for enhancement initiatives generates less than half the revenues with respect
to the neutral assumption (same cost as development by scratch) and less than 35% of
the opposite assumption (cost of enhancement is 150% of development by scratch).</p>
      <p>
        A cost model for contractual goals should be built according to sound articulated
approaches and empirically derived knowledge [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>In this paper we present an empirical study–based on data from real-life projects–to
determine objectively (although possibly approximatively) the real relation between the
cost of enhancing existing software and the cost of developing new software.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>The Empirical Study</title>
      <sec id="sec-3-1">
        <title>The dataset</title>
        <p>
          We analyzed data from the ISBSG dataset [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. However, not all the data included in
the dataset were used. First of all, we selected the data concerning projects whose size
was measured in IFPUG Function Points. In addition, following the recommendations
by ISBSG [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], only records having “Data Quality Rating” equal to B or better were
used; that is, only projects with the highest data integrity were considered. Similarly,
we selected records having “UFP rating” equal to B or better, i.e., the UFP counting
was evaluated as sound by the ISBSG quality reviewers. Projects too big or too small
were also removed. Specifically, projects having size smaller than 50 UFP were ignore,
because their development effort is so small that effort estimation is not even
convenient, since estimation cost would be a significant fraction of the development cost.
Projects having size greater than 800 UFP we discarded according to the criteria described
below, when analogy-based estimation is introduced.
        </p>
        <p>The descriptive statistics of the analyzed datasets are given in Table 1. Specifically,
Table 1 illustrates the characteristics of the New developments and Enhancements
datasets. For each dataset, statistics concerning the size expressed in UFP, the effort
expressed in PH (person hours) and the number of person hours required per function
point are given.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Description of the study</title>
        <p>First, we performed a very simple comparison of the effort per FP required by New
development and enhancement projects.</p>
        <p>
          Fig. 1 shows that the quantity of effort per FP required by enhancement projects is
generally greater: both the mean and the median are greater than new developments'.
Similarly, several enhancement projects required a very large amount of effort per FP,
as shown in the left part of Fig. 1. The main statistics concerning the required effort per
FP are given in Table 2. It appears that Enhancement projects require more effort per
FP than new development projects. This fact was tested by means of the Wilcoxon rank
sum test [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], which confirmed that the probability that a randomly selected New
development effort per FP is less than a randomly selected Enhancement effort per FP is
significantly greater than the probability of picking a greater or equal effort per FP
value.
        </p>
        <p>Another simple evaluation of the effort per FP required for New developments and
enhancement projects can be obtained via lowess curves of effort vs. size. Lowess
(locally weighted scatterplot smoothing) is a nonparametric method for fitting a smooth
curve between two variables; in this nonparametric method, the linearity assumptions
of conventional regression methods are relaxed.</p>
        <p>Several project managers could consider the observations reported above to be
sufficient to conclude that–except for small projects–the productivity of New
developments is greater than the productivity of Enhancement projects. However, we would be
more comfortable if we could provide evidence based on proper model, which provide
a statistically sound synthesis of the relationship that links effort and size. To this end,
we proceeded to build model of effort as a function of size.</p>
        <p>
          All effort models were derived via ordinary least square (OLS) regression. We used
Cook’s distance to identify possible influential observations. Data points with Cook’s
distance greater than 4/n (n being the cardinality of the training set) were considered
for removal as suggested by Kitchenham and Mendes [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. All the models illustrated
below were checked for the usual characteristics of OLS regression [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. All the results
reported are statistically significant at the α = 0.05 level, as is common in Empirical
Software Engineering and many other disciplines.
        </p>
        <p>New Development</p>
        <p>Effort = 9 Size
292 (34%)</p>
        <p>Enhancement
Effort = 11.48 Size</p>
        <p>1172 (40%)</p>
        <p>We started building linear models. The obtained models and their characteristics are
given in Table 3.</p>
        <p>The OLS linear regression models described in Table 3 provide some interesting
indications, in that 1) they confirm that Enhancement projects are more expensive than
New development projects; 2) the obtained coefficients are very close to the median
Effort/Size values given in Table 2. However, both models do not conform to the OLS
regression constraints, having not normally distributed residuals. In addition, they were
built by discarding as outliers a large fraction of the projects.</p>
        <p>
          Therefore, we looked for non-linear OLS regression models. Specifically, we
performed logarithmic transformations of both the independent and dependent variables.
This choice was suggested by both the desire to be compliant with earlier research (for
instance, Boehm’s COCOMO [
          <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
          ]) and to deal with the characteristics of the data
distributions (as suggested by Kitchenham and Mendes [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], among others).
NO
The log-log model for Enhancements has not normal residuals; the log-log model for
New developments has normal residuals, but a rather low coefficient of determination.
However, they tend to conform that Enhancement projects have greater cost per FP than
New developments. The difference in terms of required effort can be observed by
looking at the models' curves, shown in Fig. 3.
        </p>
        <p>Since no really adequate OLS regression models could be achieved, we adopted
analogy-based estimation (AbE). With AbE, the effort required by a project is estimated
based on the effort that was required by "similar" projects.</p>
        <p>Given a project P, we selected the projects that contribute to estimate the
development or maintenance effort for P as follows:
1) Let sp=0.02.
2) Let NP be the set of projects such that p ∈ NP iff (1-sp) size(P) ≤ size(p) ≤ (1+sp)
size(P)
3) If |NP| ≥ 7, let the estimated effort for P be the median of the efforts of the
projects belonging to NP.</p>
        <p>4) Otherwise, increase sp by 0.01 and go back to step 2).</p>
        <p>Before proceeding to apply AbE, we determined the size range in which there are
enough data points to support AbE. The histograms in Fig. 4 show that–as could be
expected–there are few large Enhancement projects. Specifically, above 800 UFP
projects are few and sparse, hence it is difficult to find "similar" projects for AbE.
Therefore, in the rest of this study we consider only projects having size not greater than 800
UFP. Noticeably, for New development there are no problems of the type discussed
above: there are many projects having size greater than 800 UFP, actually, there are
many projects up to 2000 UFP. Nonetheless, since we are mainly interested in
comparison between Enhancement and New development projects, we need to study both types
of projects in the same size range.</p>
        <p>
          To evaluate whether AbE estimates are worth considering, we have to verify that
they provide better performance than baseline models. To this end, Shepperd and
MacDonell [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] proposed that an estimation model be taken into consideration only if it
provides better estimates than a baseline model; they also proposed to use random
estimation as a baseline model.
        </p>
        <p>Shepperd and MacDonell also proposed that the accuracy of a given estimation
method be measured via the Mean Absolute Residual (MAR), i.e., the mean of the
absolute values of errors, where errors are computed as actuals minus estimates.</p>
        <p>A random effort estimation for a project is obtained by picking at random the actual
effort of any of the other projects. Of course, in this way there are n−1 possible
estimates for every project; therefore, to compute the MAR of the random model we need
to average all these possible values. Shepperd and MacDonell suggest to make a large
number of random estimates (typically 1000), and then compute the mean MAR.
Shepperd and MacDonell observed also that the value of the 5% quantile of the random
estimate MARs can be interpreted like α for conventional statistical inference.
Accordingly, the MAR of a proposed model should be compared with the 5% quantile of the
random estimate MARs, to make us reasonably sure that the model is actually more
accurate than the random estimation.</p>
        <p>
          We also used “constant” models as baseline, proposed, among others, by Lavazza
and Morasca [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and Di Martino et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. With the mean–respectively, median–
constant model, the estimated effort for a project is given by the mean–respectively,
median–of all other projects' actual efforts.
        </p>
        <p>The models’ estimation errors are given in Table 5 and visualized via a boxplot in Fig.
5, where outliers are not show, for readability. In Fig. 5, the orange diamonds represent
the means, i.e., the MARs. The dashed line is the MAR of the mean constant model,
the dotted line is the MAR of the median constant model, the continuous line is the 5%
quantile of the random estimate MARs. It can be observed that the MARs of AbE
estimates are smaller than the baselines' MARs, hence AbE is an improvement over
baseline estimation methods. This fact is also confirmed by Wilcoxon sign rank test.
1940
2664
2267
3427</p>
        <p>Fig. 6 show Effort estimates, in comparison with actual effort values. Understanding
the trend of estimated effort vs size (hence of effort per FP) from Fig. 6 is not easy.
Thus, we use again the lowess curves to provide a more readable representation of the
estimated effort vs. size. Fig. 7 shows such lowess curves for New developments and
Enhancement, respectively. It can be observed that, just like in Fig. 2, the slope of the
New development curve decreases around 300 FP, while the Enhancement curve
appears approximately straight.</p>
        <p>Fig. 8 compares the lowess curves given in Fig. 7, highlighting that the effort per FP
required by Enhancement projects is greater than the effort per FP required by New
developments. Fig. 8 is remarkably similar to Fig. 2 and Fig. 3: this fact reinforces the
observation given above.
In the previous sections, we evaluated the effort per UFP in several ways: using simple
statistics, building regression models, and via analogy-based estimation. In all cases we
got clear indications that–at least when the project size is larger than 300 UFP, the effort
required by Enhancement projects is larger than the effort required by New
developments.</p>
        <p>However, large variations of the effort per UFP were observed in the ISBSG dataset:
Table 2 shows that the standard deviation is larger than the mean for both New
developments and Enhancements; for the latter it is almost twice the mean. Therefore,
practitioner should be careful in using the productivity values presented in this paper: they
should take into account some variability.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Threats to Validity</title>
      <p>Internal validity of the study could be affected by the way measures were obtained. We
mitigated this threat by carefully selecting the measures that are classified as most
reliable in the ISBSG dataset.</p>
      <p>The models presented in Section 3 do not qualify as proper regression model, since
most of them have not normally distributed residuals (and sometimes other problems
as well). Nonetheless, together with the other findings given in Section 3, they provide
reasonably reliable indications, at least qualitatively.</p>
      <p>Concerning the generalizability of the proposed results, we analyzed the largest
public available dataset, (see the descriptive statistics in Table 1). Although we cannot
claim that our findings are generally valid, the fact that they are based on a large dataset,
which collects data from many different software development organizations, supports
the hypothesis that our findings are representative of many software projects.</p>
      <p>It can be noticed that we built effort models based only on size, while usually effort
models account for multiple factors that are believed to affect development or
maintenance effort. In this paper, we limited the investigation to size based models to get
straightforward indications concerning the amount of effort needed in relation to the
size of the project. In this respect, it is worth noting that currently many public
administrations and private organizations, worldwide, adopt contractual cost models that are
based on the size of the software to be delivered as the only independent variable.
Although such practice has evident limits, it is widely used. This paper provides results
that help applying the mentioned practice based on objective empirical knowledge, thus
avoiding macroscopic mistakes, like assuming that enhancement cost less–on a unitary
basis–than new development.</p>
      <p>Another possible threat comes from the possibility that the observed differences are
due to the managerial choices concerning new development and enhancement projects.
For instance, managers could assign more skilled people to new development projects
and less skilled people to enhancement projects: this would explain the observed
differences, at least partly.</p>
      <p>Finally, we have to notice that the available dataset contains little data concerning
enhancement projects larger than 800 FP. Accordingly, we limited the study to projects
not greater than 800 FP. We cannot make any claim concerning larger projects.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Related Work</title>
      <p>
        Approaches to effort prediction are generally grouped into three general categories:
expert judgement, algorithmic models, and analogy [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Estimation by analogy predicts the effort of the target project using information from
former similar project at the system level and sub-system level. The primary steps in
this method consist of: 1) choosing the right analogy, usually measured via Euclidean
distance, 2) investigating similarities and differences, 3) examining analogy quality, 4)
providing the estimation [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        The key factor of successful EbA method is finding the appropriate right analogy.
The simple method is using a fixed number of analogies starting from k=1 and increase
this number until non further improvement on the accuracy can be obtained [
        <xref ref-type="bibr" rid="ref14 ref15">14,15</xref>
        ].
However, in our case, finding additional analogies to be used in conjunction with size
analogy was not appropriate: having to devise models that support the notion of effort
per FP, we needed to restrict the independent variable to size, specifically, size
measured in Function Points.
      </p>
      <p>In fact, the great majority of software effort estimation methods proposed in the
literature use software size or size dependent elements as a major explanatory variable.
Size measures are considered as the most influential predictors for estimation. Several
functional size measures have been defined, including Function Points, Use Case
Points, NESMA, FiSMA, Mark II FP, COSMIC FP, SiFP and many others. In this
paper we investigate the relationship of effort required between enhancement projects and
new development projects, using only FP as a size measure. Investigations using other
measures is an objective for future work.</p>
      <p>
        Concerning the study of productivity of software enhancement projects, this is
definitely not a new research field. Back in 1993, Abran and Robillard performed an
empirical study based on the analysis of 21 projects [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]: among other results, they found
that the mean effort per function point was 18.96 PH/FP, quite close to 19.67 PH/FP,
the mean value we found in the ISBSG dataset (see Table 2).
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>Having a–possibly approximate–quantitative knowledge of the unitary cost of software
projects in terms of effort per function point is very important for several purposes.
Specifically, when considering the possibility of making a bid, having a reasonably
good idea of the cost of the project is essential.</p>
      <p>Relatively little research was performed concerning the unitary cost of enhancement
projects, at least in comparison with the huge amount of work performed on
investigating the cost of new software development. The relatively small amount of empirical
data is favoring ill-based guessing. In some environment, the false believe that software
enhancement activity cost less than new development–on a unitary basis–is spreading.</p>
      <p>In this paper, we looked at what reliable and objective knowledge we can derive
from real project data.</p>
      <p>Not all the results we presented here are perfectly reliable from a statistical point of
view. Nonetheless, the consistency of results we obtained via different analysis
techniques seems to indicate that the indications we derived are–at least
qualitatively–correct and reliable.</p>
      <p>In summary, we found that:
˗ Enhancement projects have a unitary cost that is generally greater than new
development projects.
˗ Specifically, the unitary cost of enhancements and new developments is similar
for projects up to around 300 FP, while for larger projects the unitary cost of
enhancements is greater.
˗ As shown in Fig. 6, the unitary cost is largely variable, even for projects having
approximately the same size. Therefore, the data presented here must be
regarded as indicating tendencies, but are not necessarily valid for all projects.</p>
      <p>From the cognitive bias point of view, our empirical study showed that, based on the
data available in a very reputed dataset (namely, the ISBSG dataset), the assumption
that productivity is higher for functional enhancement projects than for new
development projects is not supported by evidence. Instead, the opposite is true, except for
fairly small projects. We may thus state that–most likely–many huge contracts have
been undervalued for years because of an apparently reasonable assumption, not
confirmed by empirical data.</p>
      <p>Future work include, among other activities, looking for factors that let us select
project classes characterized by small variations in unitary cost, and experimenting with
different techniques for building effort models.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by the “Fondo di ricerca d’Ateneo” of the
Università degli Studi dell’Insubria.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Boehm</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Software Engineering Economics</article-title>
          . Prentice
          <string-name>
            <surname>Hall</surname>
          </string-name>
          (
          <year>1981</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Boehm</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madachy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steece</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , et al.:
          <article-title>Software cost estimation with Cocomo II. Prentice Hall PTR (</article-title>
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Maxwell</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>Applied statistics for software managers</article-title>
          .
          <source>Prentice Hall PTR Englewood Cliffs</source>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Albrecht</surname>
            ,
            <given-names>A.J.:</given-names>
          </string-name>
          <article-title>Measuring application development productivity</article-title>
          .
          <source>In: Proceedings of the joint SHARE/GUIDE/IBM application development symposium</source>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>92</lpage>
          . IBM (
          <year>1979</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kitchenham</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Further comparison of cross-company and within-company effort estimation models for web applications</article-title>
          .
          <source>In Proceedings of the 10th International Symposium on Software Metrics. IEEE</source>
          ,
          <fpage>348</fpage>
          -
          <lpage>357</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>LNCS</given-names>
            <surname>Homepage</surname>
          </string-name>
          , http://www.springer.com/lncs, last accessed
          <year>2016</year>
          /11/21.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. ISBSG.
          <article-title>Productivity Data Query (PDQ) Tool User Guide</article-title>
          .
          <source>Technical Report. ISBSG</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. ISBSG, International Software Benchmarking Standards Group.
          <article-title>ISBSG repository-release R13 (</article-title>
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shepperd</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>MacDonell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Evaluating prediction systems in software project estimation</article-title>
          .
          <source>Information and Software Technology</source>
          <volume>54</volume>
          (
          <issue>8</issue>
          )
          <fpage>820</fpage>
          -
          <lpage>827</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lavazza</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Morasca</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>On the Evaluation of Effort Estimation Models</article-title>
          .
          <source>In Proceedings of the 21st International Conference on Evaluation and Assessment in Software Engineering. ACM</source>
          ,
          <volume>41</volume>
          -
          <fpage>50</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Di</given-names>
            <surname>Martino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ferrucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Gravino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            and
            <surname>Sarro</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Assessing the Effectiveness of Approximate Functional Sizing Approaches for Effort Estimation</article-title>
          . Information and Software
          <string-name>
            <surname>Technology</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Shepperd</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schofield</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kitchenham</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Effort estimation using analogy</article-title>
          .
          <source>In Proceedings of the 18th International Conference on Software Engineering. IEEE</source>
          ,
          <fpage>170</fpage>
          -
          <lpage>178</lpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Azzeh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Nassif</surname>
          </string-name>
          , A. B.
          <article-title>Analogy-based effort estimation: a new method to discover set of analogies from dataset characteristics</article-title>
          .
          <source>IET Software (2)</source>
          ,
          <fpage>39</fpage>
          -
          <lpage>50</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mosley</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Counsell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>A replicated assessment of the use of adaptation rules to improve Web cost estimation</article-title>
          .
          <source>In Proceedings of the International Symposium on Empirical Software Engineering</source>
          , IEEE,
          <fpage>100</fpage>
          -
          <lpage>109</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Azzeh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Marwan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Value of ranked voting methods for estimation by analogy</article-title>
          ,
          <source>IET Software</source>
          ,
          <volume>7</volume>
          (
          <issue>4</issue>
          ),
          <fpage>195</fpage>
          -
          <lpage>202</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Abran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Robillard</surname>
            ,
            <given-names>P. N.</given-names>
          </string-name>
          <article-title>Reliability of function points productivity model for enhancement projects (a field study)</article-title>
          .
          <source>In Proceedings of the International Conference on Software Maintenance, IEEE 80-87</source>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Gerd</given-names>
            <surname>Gigerenzer</surname>
          </string-name>
          .
          <article-title>Reckoning with Risk: Learning to Live with Uncertainty. Penguin books (</article-title>
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mohanani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salman</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turhan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodríguez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ralph</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>Cognitive biases in software engineering: a systematic mapping study</article-title>
          .
          <source>Transactions on Software Engineering. IEEE</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>NESMA</surname>
          </string-name>
          ,
          <article-title>Function Point Analysis for Software Enhancement</article-title>
          ,
          <source>Guidelines Version 1</source>
          .0,
          <string-name>
            <surname>NESMA</surname>
          </string-name>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Meli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and Iorio,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Software measurement issues in contractual environments</article-title>
          .
          <source>In Proceedings of the Software Measurement European Forum</source>
          . (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Wilcoxon</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>Individual comparisons by ranking methods</article-title>
          . Breakthroughs in statistics. Springer, New York, NY,
          <fpage>196</fpage>
          -
          <lpage>202</lpage>
          (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>