<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>International Conference on Software Process and Product Measurement (MENSURA), September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Software Development Efort Estimation Using Function Points and Simpler Functional Measures: a Comparison</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luigi Lavazza</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angela Locoro</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Meli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Processing Organization Srl</institution>
          ,
          <addr-line>Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Università degli Studi dell'Insubria</institution>
          ,
          <addr-line>Varese</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Università degli Studi di Brescia</institution>
          ,
          <addr-line>Brescia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>4</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Background - Functional Size Measures are widely used for estimating the development efort of software. After the introduction of Function Points, a few “simplified” measures have been proposed, aiming to make measurement simpler and quicker, but also to make measures applicable when fully detailed software specifications are not yet available. It has been shown that, in general, software size measures expressed in Function Points do not support more accurate efort estimation with respect to simplified measures. Objective - Many practitioners believe that when considering “complex” projects, i.e., project that involve many complex transactions and data, traditional Function Points measures support more accurate estimates than simpler functional size measures that do not account for greater-then-average complexity. In this paper, we aim to produce evidence that confirms or disproves such belief. Method - Based on a dataset that contains both efort and size data, an empirical study is performed, to provide some evidence concerning the relations that link functional size (measured in diferent ways) and development efort. Results - Our analysis shows that there is no statistically significant evidence that Function Points are generally better at estimating more complex projects than simpler measures. Function Points appeared better in some specific conditions, but in those conditions they also performed worse than simpler measures when dealing with less complex projects. Conclusions - Traditional Function Points do not seem to efectively account for software complexity. To improve efort estimation, researchers should probably dedicate their efort to devise a way of measuring software complexity that can be used in efort models together with (traditional or simplified) functional size measures.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Unadjusted Function Points (UFP)</kwd>
        <kwd>Simple Function Points (SFP)</kwd>
        <kwd>efort estimation</kwd>
        <kwd>simple functional size measures</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Functional Size Measures (FSM) are widely used for estimating the development efort of
software, mainly because they can be obtained in the early stages of development, when efort
estimates are most needed. Function Point Analysis (FPA) was introduced to yield a measure of
software size based exclusively on specifications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        After the introduction of Function Points (FP), a few “simplified” measures have been proposed,
aiming to make measurement simpler and quicker, but also to make measures applicable when
fully detailed software specifications are not yet available. Among the simplified measures
are Simple Function Points (SFP) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] (formerly known as SiFP [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) and the sheer number of
transaction functions.
      </p>
      <p>
        It has been shown [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ] that, in general, software size measures expressed in Unadjusted
Function Points1 (UFP) do not support more accurate efort estimation with respect to
“simpliifed” measures. However, many practitioners who use UFP for estimation believe that when
considering “complex” projects, i.e., projects that involve many complex transactions and data,
UFP measures support more accurate estimates than SFP or other measures that do not account
for greater-then-average complexity. In this paper, we illustrate an empirical study that is meant
to confirm or disprove the aforementioned belief. The study is based on the analysis of the
ISBSG dataset [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which has been widely used for studies concerning software functional size.
      </p>
      <p>
        The results of the study will likely be helpful for the numerous software development
organizations that use FPA: e.g., organizations that develop software for public administration
and are thus required to provide software size measured via IFPUG FPA by local laws (as in
in Brazil, Italy, Japan, South Korea, and Malaysia). Also, other organizations may need FP
measures because they use efort estimation tools (like Galorath’s Seer-SEM [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], for instance)
that take the size expressed in FP as input (together with some parameters that account for the
development process and technology, non-functional requirements, human factors, etc.).
      </p>
      <p>The results of the study can be interesting also for organizations that use agile development
processes. In fact, traditional functional size measurement is not very popular in agile contexts,
because it is perceived as a “heavy” method, not suitable for agile development. However,
simplified functional size measurement methods could fit easily in agile development practices,
especially when the simplification is pushed to counting only transactions, i.e., functional
elements that can be easily identified from user stories.</p>
      <p>The paper is organized as follows. Section 2 recalls some basic notions concerning Functional
Size Measurement (FSM) methods. Section 3 states the objectives of the work described here,
also by formulating research questions. Section 4 describes the empirical study through which
we addressed the research questions. The achieved results are also illustrated and discussed.
In Section 5 research questions are answered. Section 6 discusses the threats to the validity of
the study. Section 7 accounts for related work. Section 8 draws some conclusions and outlines
future work.</p>
      <sec id="sec-1-1">
        <title>1Following the ISO [6], we consider only unadjusted FP.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>In this section we provide a very brief introduction to Function Points, as well as to simplified
measures, namely SFP and the number of transactions.</p>
      <sec id="sec-2-1">
        <title>2.1. Function Point Analysis</title>
        <p>
          Function Point Analysis was originally introduced by Albrecht to measure the size of software
systems from end-users’ point of view, with the goal of estimating the development efort [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Currently, FPA is oficially documented by the IFPUG (International Function Points User Group)
via the counting practices manual [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ],
        </p>
        <p>The basic idea of FPA is that the “amount of functionality” released to the user can be
evaluated by taking into account 1) the data used by the application to provide the required
functions, and 2) the elementary processes or transactions (i.e., operations that involve data
crossing the boundaries of the application) through which the functionality is delivered to the
user. Both data and transactions are evaluated at the conceptual level, i.e., they represent data
and operations that are relevant to the user. Therefore, IFPUG Function Points are counted on
the basis of functional user requirements (FURs) specifications.</p>
        <p>FURs are modeled as a set of base functional components (BFCs), which are the measurable
elements of FURs: each of the identified BFCs is measured, and the size of the application
is obtained as the sum of the sizes of BFCs. BFCs are data functions (also known as logical
ifles), which are classified into internal logical files (ILF) and external interface files (EIF), and
transaction functions, which are classified into external inputs (EI), external outputs (EO), and
external inquiries (EQ), according to the activities carried out within the considered process
and its main intent.</p>
        <p>
          The size of every BFC is determined based on its type and its “complexity” (see the manual [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]
for details). The functional size of a given application, expressed in unadjusted Function Points,
is given by the sum of the sizes of all its BFCs.
        </p>
        <p>The core of FPA involves the following main activities:
1. Identifying data functions.
2. Identifying transaction functions.
3. Classifying data functions as ILF or EIF.
4. Classifying transactions functions as EI, EO or EQ.
5. Determining the complexity of each data function.
6. Determining the complexity of each transaction function.</p>
        <p>The first four of these activities can be carried out even if the FURs have not yet been fully
detailed. On the contrary, the last two activities require that details are available.</p>
        <p>
          Simplified Functional Size Measurement methods aim at providing estimates of functional
size measures by skipping one or more of the activities listed above. Specifically, simplified
measurement methods tend to skip at least the determination of complexity, since this activity
is time- and efort-consuming [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Simple Function Points</title>
        <p>
          The Simple Function Point measurement method [
          <xref ref-type="bibr" rid="ref11 ref3">3, 11</xref>
          ] has been designed by Meli to be
lightweight and easy to use. Like IFPUG FPA, it is independent of the technologies and of the
technical design principles.
        </p>
        <p>SFP requires only the identification of Elementary Processes (EP) and Logical Files (LF),
based on the assumption that value to a BFC is given as a whole, independently of internal
organization and details.</p>
        <p>SFP assigns a numeric value directly to BFCs, as follows:</p>
        <p>= 7 # + 4.6 #
thus speeding up the functional sizing process at the expense of ignoring the domain data
model and the primary intent of each Elementary Process.</p>
        <p>The weights for each BFC were originally defined to achieve the best possible approximation
of FPA. However, since SFP is a measurement method, those weights are constants, i.e., they
are not subject to update or change for approximation reasons, and are now crystallized for
stability, repeatability, and comparability reasons.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Extremely Simplified Functional Size Measures</title>
        <p>As described in Section 2.2 above, SFP adopts fixed weights for elementary processes and logical
data files. A further simplification consists in not considering at all data in the measurement of
functional size. In this sense, the most straightforward measure is given by the sheer number of
transactions (#TF). Note that applying a weight to the number of transactions would result in a
simple scale transformation, with no improvement in the correlation with development efort.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Research Questions</title>
      <p>
        Some research has already been dedicated to evaluating the possibility of using functional
size measures that are definitely simpler than standard IFPUG UFP for efort estimation [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
Simpler metrics are of great interest for practitioners because they are quicker and less expensive
to collect than traditional FP, and, even more important, simple measures can sometimes be
applied before detailed and complete software requirements are available.
      </p>
      <p>However, previous research proposed empirical studies whose conclusions were based on
the evaluation of estimation accuracy over the entire test set. Such practice, although sound
and informative, does not solve possible doubts about the performance of diference metrics
when dealing with projects having diferent complexity.</p>
      <p>In fact, in some environments, it is believed that traditional UFP are better at accounting
for the complexity of projects, hence, when dealing with relatively complex projects, UFP are
expected to support more accurate efort estimation with respect to simpler FSM methods.
However, as far as we know, hardly any evidence has been produced to support this belief.</p>
      <p>In this paper, we provide some evidence that can be used to either support or disprove the
aforementioned belief. To this end, we formulate the following research questions:
RQ1 If project complexity is not taken into account, is it true that simple functional measures
(namely, SFP and #TF) provide efort estimates that are as accurate as those provided by
standard IFPUG UFP?
RQ2 For projects that have relatively high (respectively, low) complexity, do UFP and simple
functional metrics (namely, SFP and #TF) support efort estimation at significantly diferent
levels of accuracy?</p>
      <p>It is well known that there are multiple ways for i) modeling the dependence of development
efort on software functional size; ii) evaluating (in a statistically sound manner) the accuracy
of the obtained estimates; iii) classifying projects as relatively complex or relatively simple, etc.
Answering the research questions for all the possible ways of addressing the issues mentioned
above is hardly possible. Therefore, in this paper we adopt reasonable models and classification
techniques, preferring simper ones, to avoid the risk of getting results that depend on the
intricacies of the technical instruments being used.</p>
    </sec>
    <sec id="sec-4">
      <title>4. The Study</title>
      <sec id="sec-4-1">
        <title>4.1. The Dataset</title>
        <p>In this section we describe the empirical study that supports our answers to the research
questions.</p>
        <p>
          In our empirical study, we analyzed data from the ISBSG dataset [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which includes data from
real-life software development projects and has been widely used in studies involving Functional
Size Measures.
        </p>
        <p>The ISBSG dataset contains data from both projects addressing the development of new
software products and projects addressing the enhancement of existing projects. In this paper,
we limit our investigation to new developments. Dealing with enhancement projects will be
the objective of future work.</p>
        <p>For each project, many data are provided. Of these, we used the following:
• The type of project, i.e., new development or enhancement.
• The efort spent, expressed in PersonHours (PH).
• The size, expressed in IFPUG Function Points.</p>
        <p>• #ILF, #EIF, #EI, #EO, and #EQ, each split per complexity (high, medium, low).</p>
        <p>It is worth noting that the considered version of the ISBSG dataset contains some measures
(namely, Efort, the size in UFP and the number of transactions) that we use as-is, as well as raw
data (#ILF, #EIF, #EI, #EO, and #EQ) that we used to compute the size in #TF (#EI+#EO+#EQ)
and SFP (7(#ILF+#EIF)+4.6#TF). The type of project was used only to select new development
projects.</p>
        <p>The dataset includes data from 533 new development projects. Descriptive statistics of the
ISBSG dataset are given in Section 4.3.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. The Method</title>
        <p>4.2.1. The efort model
In this paper, we use a very simple method for building efort models. In fact, we assume that
efort can be computed by dividing the size of the software product to be built by the observed
productivity:</p>
        <p>It is clear that formula (1) describes a very simple model of efort, since i) it assumes that
efort depends only on functional size, and ii) it is structurally simple, especially when compared
with models that can be obtained via sophisticated techniques like machine learning, neural
networks, etc. We preferred this extremely simple model to avoid possible confounding efects.</p>
        <p>
          Productivity is defined as [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
        </p>
        <p>Efort =</p>
        <p>Size</p>
        <p>Productivity
Productivity =</p>
        <p>Size
Efort
(1)
(2)
(3)</p>
        <p>However, the value of Productivity to be used in (1) can be obtained in diferent ways. In this
paper, we consider three possible derivations of the Productivity value:
1. For each project in the dataset, we considered its Productivity, as defined in (2). Then
we computed the mean value of the projects’ productivity. In doing this, we used as
a size measure UFP, SFP and #TF, thus obtaining ProductivityUFP, ProductivitySFP and
Productivity#TF.
2. We proceeded as described above, but the productivity was obtained as the median value
of the projects’ productivity.
3. Finally, we considered the case when the productivity is given. This is the case in some
software acquisition markets, where the productivity is assumed to have a conventional
value, so that every UFP (or SFP) has the corresponding value and price.</p>
        <p>Productivity was then used to compute, via formula (1), the estimated efort for each project
in the dataset.</p>
        <p>Estimation errors EstErr were then computed:</p>
        <p>EstErr = ActualEfort − EstimatedEfort = ActualEfort −</p>
        <p>Size
Productivity
Specifically, the computation described by formula (3) was computed for the three considered
Functional size measures, i.e., UFP, SFP and #TF. For instance,</p>
        <p>EstErrUFP = ActualEfort −</p>
        <p>SizeUFP
ProductivityUFP
where SizeUFP is the functional size of the considered project, expressed in UFP. Similarly, we
obtained EstErrSFP and EstErr#TF for each project in the ISBSG dataset.
4.2.2. Evaluation of estimation accuracy
We performed a sign test to evaluate whether any of the considered measures supports more
accurate efort estimates than the other considered FSM methods. For instance, we counted for
how many projects it is |EstErrUFP| &lt; |EstErrSFP|: let UFP be that number; using the binomial
test (with  = 0.05), we evaluated whether we can safely conclude that estimates based on
UFP are more accurate than estimates based on SFP. In practice, we tested if the probability of
achieving UFP (where  is the total number of projects) is greater than 12 .</p>
        <p>Similarly, we evaluated UFP vs #TF and SFP vs #TF.
4.2.3. Classification of projects according to complexity
Research question RQ2 requires identifying projects that are “complex.” To this end, we need to
properly define the notion of complexity. In the context of Function Point Analysis, complexity
is evaluated by weighting base functional components. Therefore, we exploit this practice to
evaluate projects’ complexity. Noticeably, the ISBSG dataset does not provide other thorough
and consistent information about projects’ complexity, hence our choice was actually forced by
the available data.</p>
        <p>Accordingly, we proceeded as follows:
1. For each project, we computed the proportion tf of high complexity transactions over
the total number of transactions.
2. We computed the 20ℎ and 80ℎ percentiles from the distribution of tf, thus obtaining
tf20 = 0.037 and tf80 = 0.493.
3. We selected the projects having tf &lt; tf20 as simple, those having tf &gt; tf80 as
complex, and those with tf20 ≤ tf ≤ tf80 as medium complexity ones.
4.2.4. Dealing with fixed productivity
As mentioned in Section 4.2, we used both productivity values derived from the ISBSG data
and a fixed market value. The latter is the value imposed by CONSIP (Concessionaria Servizi
Informativi Pubblici) for contracts involving the public administration in Italy. Currently, such
productivity value is 1.7 UFP/PD, where a PersonDay (PD) includes 8 PH, hence 1.7 UFP/PD is
approximately 0.283 UFP/PH. Note that 0.283 UFP/PH is much greater than both the mean and
median values from the ISBSG dataset (see Table 1).</p>
        <p>Once the productivity in UFP/PH had been established, we had to devise proper corresponding
values for the productivity in SFP/PH and the productivity in #TF/PH. This task is not easy,
because it involves assuming ratios among the considered size measures; e.g., we have to assume
that  UFP corresponds to  SFP.</p>
        <p>
          In the empirical study described here, we assume that 1 UFP corresponds to 1 SFP, based on
previous research [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ].
        </p>
        <p>To the best of our knowledge, there are no results available concerning the ratio among
the #TF and UFP or SFP; therefore, we computed such ratio using the available ISBSG data. It
turned out that the mean value of #TF/UFP is approximately 8.307; hence, we assumed that the
productivity in #TF corresponding to 0.283 UFP/PH is 0.283 × 8.307 = 1.765 #TF/PH.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Results</title>
        <p>In this section, the results of our analyses are described. Section 4.3.1 illustrates the results
obtained when considering all the new development projects from the ISBSG dataset, while
Section 4.3.2 concerns only those projects that are more efort-consuming.</p>
        <p>In both sections, all combinations of software complexity and productivity are considered:
complexity is either ignored (i.e., all projects are considered together) or it is used to split the
dataset into low, mid and high complexity ones; productivity is obtained as the mean or the
median of ISBSG projects’, or it is a given value.
4.3.1. Results obtained from all new development projects
Table 1 provides descriptive statistics for the new development projects contained in the dataset.</p>
        <p>Figure 1 shows the distribution of all new development projects’ efort in PH.</p>
        <p>When using mean productivity in model (1), we obtained estimation errors whose distribution
is described in Figure 2.</p>
        <p>The boxplots in Figure 2 indicate that the three considered measures yield quite similar error
distributions. This was confirmed by the sign test, whose results are summarized in Table 2.</p>
        <p>Each cell of the table provides a symbol followed by two numbers in parentheses: the symbol
indicates if the measure in row was better (“&gt;”), equivalent (“=”) or worse (“&lt;”) than the
measure in column; the first number indicates how many times the measure in row was better
than the measure in column; the second number indicates how many times the measure in
column was better than the measure in row. For instance, the cell in row 1 and column 3
indicates that UFP supported more accurate estimates for 261 projects, while #TF supported
more accurate estimates for 272 projects; the “=” symbol indicates that according to the binomial
test the observed diference (261 vs 272) is not large enough to conclude that the probability of
UFP being better or worse than #TF is not 0.5.</p>
        <p>In practice, when considering all new development projects and using mean productivity in
the efort estimation formula (1), there is no statistically significant evidence that the considered
functional measures perform diferently.</p>
        <p>When using median productivity in model (1), we obtained the estimation errors described
in Figure 3.</p>
        <p>The results of the sign tests are summarized in Table 3.</p>
        <sec id="sec-4-3-1">
          <title>The results of the sign test are summarized in Table 4.</title>
          <p>We then proceeded to evaluate separately the high-, mid- and low-complexity projects. As
mentioned in Section 4.2, we computed the 20ℎ and 80ℎ percentiles from the distribution of
the proportion tf of high complexity transactions over the total number of transactions,
obtaining tf20 = 0.037 and tf80 = 0.493. The new development projects of the ISBSG dataset
are split by complexity as follows:
• 108 low-complexity (tf &lt; tf20) projects.
• 163 mid-complexity (tf20 ≤ tf ≤ tf80) projects.</p>
          <p>• 107 high-complexity (tf &gt; tf80) projects.</p>
          <p>The estimation errors obtained when using the mean productivity are shown in Figure 5.</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>The results of the sign tests are summarized in Table 5.</title>
          <p>The estimation errors obtained when using the median productivity are shown in Figure 6.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>The results of the sign tests are summarized in Table 6. The estimation errors obtained when using a given fixed productivity are shown in Figure 7. The results of the sign tests are summarized in Table 7.</title>
          <p>4.3.2. Results obtained from selections of new development projects
As shown in Figure 1, the great majority of ISBSG new development projects required a relatively
small efort. Specifically, 30% of the projects required no more than a PersonYear, while more than
50% require less than 2 PersonYear. We can thus conclude that the results reported in Section 4.3.1
are determined mainly by small (in terms of efort) projects. It is thus necessary to reconsider
the research questions in the context of projects that require considerable development efort.
To this end, we repeated the analysis described in Section 4.3.1, considering only projects that
require considerable development efort. For the sake of space, in this section we report only
the results of the sign tests, while estimation error boxplots are omitted.</p>
          <p>As a first step, we had to decide which projects should be involved in the analysis. We decided
to retain the projects that required no less than two PersonYears, i.e., 2 × 210 × 8 = 3360 PH
(assuming 210 working day per year and 8 working hours per day). In this way, we selected 247
projects. The descriptive statistics of this dataset are given in Table 8. In the rest of the paper,
these projects are named conventionally “not too small.”</p>
          <p>When using mean productivity in model (1), the sign tests applied to absolute residuals
yielded the results summarized in Tables 9 (all projects) and 12 (only not too small projects).</p>
          <p>When using median productivity in model (1), the sign tests applied to absolute residuals
yielded the results summarized in Tables 10 (all projects) and 13 (only not too small projects).</p>
          <p>When using the given fixed productivity in model (1), the sign tests applied to absolute
residuals yielded the results summarized in Tables 11 (all projects) and 14 (only not too small
projects).
4.3.3. Results for enhancement projects
In this paper, we focused on projects involving the development of new software (i.e.,
development from scratch). The ISBSG dataset provides also data concerning projects that involved
enhancing existing software. For time and space reasons, we could not extend our analysis to
this type of projects in this paper. However, preliminary analysis of enhancement projects seem
to confirm the results concerning new developments.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>In this section, we answer the research questions enunciated in Section 3. Having considered
two simple functional size measures (SFP and #TF), we answer each question separately for
UFP vs SFP and UFP vs #TF.</p>
      <p>In interpreting the results given below, we recommend some caution when a given
productivity is used: in those cases, the given productivity concerns UFP, while the equivalent
productivities for SFP and #FP are not given; hence we had to devise them (see Section 4.2.4) in
a somewhat subjective manner.</p>
      <sec id="sec-5-1">
        <title>5.1. Answer to RQ1</title>
        <p>Research question RQ1 asks if it is true that SFP and #TF provide efort estimates that are as
accurate as those provided by standard IFPUG UFP, when project complexity is not taken into
account.</p>
        <p>Table 15 summarizes the results illustrated in Section 4.3.1 that are relevant for RQ1.</p>
        <p>Based on the collected results, we can state that there is hardly any diference in estimation
accuracy when using SFP instead of UFP.</p>
        <p>When #TF are used, the answer clearly depends on the method used to compute the
productivity.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Answer to RQ2</title>
        <p>Research question RQ2 asks if UFP and simple functional metrics support efort estimation at
significantly diferent levels of accuracy for projects that have diferent complexity.</p>
        <p>Based on the collected results, it can be observed that in a few cases (depending on how
productivity is computed) UFP appear more accurate when dealing with high-complexity
software, while SFP appear more accurate when dealing with low-complexity software. This
suggests that UFP may be slightly biased in considering complexity.</p>
        <p>When considering #TF, answering RQ2 is very dificult. There is no clear pattern in the
performances of UFP vs #TF: for instance, when using mean productivity for mid-complexity
software, UFP perform better with more efort-consuming (or “not too small”) projects, while
#TF perform better with less efort-consuming projects.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Threats to validity</title>
      <p>A typical concern for considering only empirical data is the lack of theoretical point of view,
for example in defining complexity and complex software projects. However, we started from
some consolidated empirical evidence and practices about the criteria of software functional
size, and we followed the common praxis of the community. One of the reasons why our results
challenge a strong believe in the community (the more the UFP measure is “complex”, the better
it is correlated to efort), probably comes from the lack of too theoretical reflections. However,
this is not a limitation of our paper only, but a more generalized problem.</p>
      <p>Some decisions made while carrying out the study might have influenced the results. However,
such decisions were necessary to perform the analysis. When dealing with the choices that
most obviously could afect our results, we carried out some sensitivity analysis. For instance,
concerning the criteria used to identify “not too small” projects when the median productivity
model is used, we tried increasing (up to doubling) the minimum efort threshold that qualifies
a project as “not to small” and we noticed no diferences.</p>
      <p>Another major concern in these kinds of studies is the generalizability of results outside the
scope and context of the analyzed dataset. The ISBSG dataset is deemed the standard benchmark
among the community, and it includes data from several application domains. Therefore, our
results should be representative of a fairly comprehensive situation. However, additional studies
are needed for confirming the generalizability of the results presented above. This is particularly
true of studies concerning enhancement projects, that we could not treat adequately in this
paper.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Related Work</title>
      <p>
        Since the introduction of Function Point Analysis, many researchers and practitioners strived
to develop simplified versions of the FP measurement process, both to reduce the cost and
duration of the measurement process, and to make it applicable when full-fledged requirements
specifications are not yet available [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref3">13, 14, 15, 16, 17, 18, 19, 3, 20</xref>
        ].
      </p>
      <p>
        These simplified measurement methods were then evaluated with respect to their ability to
support accurate efort estimation [
        <xref ref-type="bibr" rid="ref4">21, 22, 23, 24, 25, 26, 27, 28, 4, 29</xref>
        ].
      </p>
      <p>
        More recently, Lavazza et al. considered using only the number of transaction to estimate
efort [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: it was found that efort models based on the number of transactions appear marginally
less accurate than models based on standard IFPUG Function Points for new development
projects, and marginally more accurate for projects extending previously developed software.
      </p>
      <p>To the best of our knowledge, no studies considered classifying projects according to degrees
of complexity.</p>
      <p>Since the ’90s, the early estimation of software was achieved with diferent methods. Among
them, using regression-like methods [30] or by the “Early &amp; Quick Function Point” (EQFP)
method [31], which uses analogy to discover similarities between new and previously measured
pieces of software, and analysis to provide weights for software objects. Statistical estimation
methods were first introduced by Lavazza et al., who studied the relationships between BFCs
and size measures expressed in FP [32].</p>
      <p>More recently, software efort estimation has been done with machine learning methods.
Case-Based Reasoning and Genetic Algorithm were exploited with benchmark datasets, and
improved accuracy of efort estimation [ 33]. In another study [34], agile development is the
target of the efort estimation research: a mixed method using also qualitative interviews of
software projects teams revealed the necessity to assess 12 hypothesis tests for efort estimation.</p>
      <p>Being those methods far away from the approach presented in this paper, we will not go further
in detail into these researches. It is worth notice that these approaches are complementary to
the one adopted in this paper, which may hence be applied in synergy with the one currently
presented.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusions</title>
      <p>Simplified functional size measures ignore the “complexity” of transactions, which is instead
accounted for by traditional Function Point Analysis. Some believe that this type of omission
makes simplified measures less suitable for efort estimation, when relatively complex software
products are involved. To assess the truth of this belief, an empirical study was conducted,
based on the analysis of the data from the ISBSG dataset.</p>
      <p>Our analysis shows that UFP do not appear to support more accurate efort estimation when
performances over an entire dataset are considered. Instead, when splitting the given dataset
according to transaction complexity, UFP-based estimates appear sometimes more accurate;
however, it also appears that
1. UFP performance depends on how productivity is computed. For instance, when the
median productivity is used, no significant diference between UFP and SFP is observed.
2. When UFP appear more accurate in estimating complex projects, they also appear less
accurate in estimating less complex projects.</p>
      <p>To sum up, the belief that when considering “complex” projects, i.e., projects that involve many
complex transactions and data, traditional Function Points measures support more accurate
estimates than simpler functional size measures that do not account for greater-then-average
complexity, seems not to be confirmed. It seems justified only in specific circumstances, e.g.,
whenever the productivity is imposed and definitely higher than what it is detectable by looking
at the same data. Furthermore, often UFP performs better than SFP to estimate complex projects,
but worse in less complex projects. This may suggest that UFP are not better than SFP, but
rather that they are biased towards more complex projects instead.</p>
      <p>In conclusion, our results show that the complexity of software (even measured using the
rather rough concepts of FPA) can afect efort estimation accuracy; at the same time, embedding
the notion of such complexity in Unadjusted Function Points does not guarantee good results.</p>
      <p>Accordingly, some interesting topics for future work include involving the notion of
complexity in efort models. A first straightforward manner of doing this consists in building a
model</p>
      <p>EstimatedEfort = Productivitycplx × Size
where Size and Productivitycplx are the size and the expected productivity, computed according
to the class of complexity (high, medium or low) of the software to be developed. More complex
models could be achieved, e.g., via machine learning techniques.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by the “Fondo di ricerca d’Ateneo” of the Università
degli Studi dell’Insubria.
[17] L. Bernstein, C. M. Yuhas, Trustworthy systems through quantitative software engineering,
vol. 1, John Wiley &amp; Sons, 2005.
[18] L. Santillo, M. Conte, R. Meli, Early &amp; Quick Function Point: sizing more with less, in: 11th</p>
      <p>IEEE International Software Metrics Symposium (METRICS’05), IEEE, 41–41, 2005.
[19] T. Iorio, R. Meli, F. Perna, Early &amp; Quick Function Points® v3. 0: enhancements for
a Publicly Available Method, in: Proceedings Software Measurement European Forum
(SMEF), 179–198, 2007.
[20] L. Lavazza, A. Locoro, G. Liu, R. Meli, Estimating software functional size via machine
learning, ACM Transactions on Software Engineering and Methodology .
[21] H. van Heeringen, E. van Gorp, T. Prins, Functional size measurement-Accuracy versus
costs–Is it really worth it?, in: Software Measurement European Forum (SMEF 2009), 2009.
[22] F. G. Wilkie, I. R. McChesney, P. Morrow, C. Tuxworth, N. Lester, The value of software
sizing, Information and Software Technology 53 (11) (2011) 1236–1249.
[23] J. Popović, D. Bojić, A comparative evaluation of efort estimation methods in the software
life cycle, Computer Science and Information Systems 9 (1) (2012) 455–484.
[24] P. Morrow, F. G. Wilkie, I. McChesney, Function point analysis using NESMA: simplifying
the sizing without simplifying the size, Software Quality Journal 22 (4) (2014) 611–660.
[25] L. Lavazza, G. Liu, An Empirical Evaluation of the Accuracy of NESMA Function Points
Estimates, in: The 14th International Conference on Software Engineering Advances
(ICSEA 2019), 24–29, 2019.
[26] S. Di Martino, F. Ferrucci, C. Gravino, F. Sarro, Assessing the efectiveness of approximate
functional sizing approaches for efort estimation, Information and Software Technology
123 (106308).
[27] L. Lavazza, G. Liu, An Empirical Evaluation of Simplified Function Point Measurement</p>
      <p>Processes, Journal on Advances in Software 6 (1&amp; 2).
[28] R. Meli, Early &amp; Quick Function Point Method-An empirical validation experiment, in: Int.</p>
      <p>Conf. on Advances and Trends in Software Engineering, Barcelona, Spain, 2015.
[29] F. Ferrucci, C. Gravino, L. Lavazza, Simple function points for efort estimation: a further
assessment, in: Proceedings of the 31st Annual ACM Symposium on Applied Computing,
ACM, 1428–1433, 2016.
[30] D. B. Bock, R. Klepper, FP-S: a simplified function point counting method, Journal of</p>
      <p>Systems and Software 18 (3) (1992) 245–254.
[31] DPO, Early &amp; Quick Function Points Reference Manual - IFPUG version, Tech. Rep.
EQ&amp;FP</p>
      <p>IFPUG-31-RM-11-EN-P, DPO, Roma, Italy, 2012.
[32] L. Lavazza, S. Morasca, G. Robiolo, Towards a simplified definition of Function Points,</p>
      <p>Information and Software Technology 55 (10) (2013) 1796–1809.
[33] S. Hameed, Y. Elsheikh, M. Azzeh, An optimized case-based software project efort
estimation using genetic algorithm, Information and Software Technology 153 (2023) 107088.
[34] S. A. Butt, T. Ercan, M. Binsawad, P.-P. Ariza-Colpas, J. Diaz-Martinez, G. Pineres-Espitia,
E. De-La-Hoz-Franco, M. A. P. Melo, R. M. Ortega, J.-D. De-La-Hoz-Hernandez, Prediction
based cost estimation technique in agile development, Advances in Engineering Software
175 (2023) 103329.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Albrecht</surname>
          </string-name>
          ,
          <article-title>Measuring application development productivity</article-title>
          ,
          <source>in: Proceedings of the joint SHARE/GUIDE/IBM application development symposium</source>
          , vol.
          <volume>10</volume>
          ,
          <fpage>83</fpage>
          -
          <lpage>92</lpage>
          ,
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>International</given-names>
            <surname>Function</surname>
          </string-name>
          Point Users Group (IFPUG),
          <source>Simple Function Point (SFP) Counting Practices Manual Release v2.1</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Meli</surname>
          </string-name>
          ,
          <article-title>Simple function point: a new functional size measurement method fully compliant with IFPUG 4</article-title>
          . x,
          <source>in: Software Measurement European Forum</source>
          ,
          <fpage>145</fpage>
          -
          <lpage>152</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lavazza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meli</surname>
          </string-name>
          ,
          <article-title>An evaluation of simple function point as a replacement of IFPUG function point</article-title>
          ,
          <source>in: 2014 Joint Conference of the International Workshop on Software Measurement and the International Conference on Software Process and Product Measurement (IWSM-MENSURA)</source>
          , IEEE,
          <fpage>196</fpage>
          -
          <lpage>206</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lavazza</surname>
          </string-name>
          , G. Liu,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meli</surname>
          </string-name>
          ,
          <article-title>Using Extremely Simplified Functional Size Measures for Efort Estimation: an Empirical Study</article-title>
          ,
          <source>in: Proceedings of the 14th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM)</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>International</given-names>
            <surname>Standardization</surname>
          </string-name>
          <article-title>Organization (ISO)</article-title>
          , ISO/IEC 20926:
          <year>2003</year>
          ,
          <article-title>Software engineering - IFPUG 4.1 Unadjusted functional size measurement method -</article-title>
          <source>Counting Practices Manual</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>International</given-names>
            <surname>Software</surname>
          </string-name>
          Benchmarking Standards Group, “
          <source>Worldwide Software Development: The Benchmark, release April</source>
          <year>2019</year>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Fischman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McRitchie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. D.</given-names>
            <surname>Galorath</surname>
          </string-name>
          ,
          <string-name>
            <surname>Inside</surname>
            <given-names>SEER</given-names>
          </string-name>
          -SEM, CrossTalk
          <volume>18</volume>
          (
          <issue>4</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>International</given-names>
            <surname>Function Point Users</surname>
          </string-name>
          <article-title>Group (IFPUG)</article-title>
          ,
          <article-title>Function point counting practices manual</article-title>
          ,
          <source>release 4.3.1</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lavazza</surname>
          </string-name>
          ,
          <article-title>On the Efort Required by Function Point Measurement Phases</article-title>
          ,
          <source>International Journal on Advances in Software Volume 10, Number</source>
          <volume>1</volume>
          &amp;
          <fpage>2</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>IFPUG</surname>
          </string-name>
          ,
          <string-name>
            <surname>Simple Function</surname>
          </string-name>
          <article-title>Point (SFP) Counting Practices Manual Release 2</article-title>
          .1,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B. W.</given-names>
            <surname>Boehm</surname>
          </string-name>
          ,
          <article-title>Improving software productivity</article-title>
          ,
          <source>Computer</source>
          <volume>20</volume>
          (
          <issue>09</issue>
          ) (
          <year>1987</year>
          )
          <fpage>43</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Horgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khaddaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Forte</surname>
          </string-name>
          ,
          <article-title>Construction of an FPA-type metric for early lifecycle estimation</article-title>
          ,
          <source>Information and Software Technology</source>
          <volume>40</volume>
          (
          <issue>8</issue>
          ) (
          <year>1998</year>
          )
          <fpage>409</fpage>
          -
          <lpage>415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Meli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Santillo</surname>
          </string-name>
          ,
          <article-title>Function point estimation methods: A comparative overview</article-title>
          ,
          <source>in: FESMA</source>
          , vol.
          <volume>99</volume>
          ,
          <issue>Citeseer</issue>
          ,
          <fpage>6</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <article-title>NESMA-the Netherlands Software Metrics Association, Definitions and counting guidelines for the application of function point analysis</article-title>
          .
          <source>NESMA Functional Size Measurement method compliant to ISO/IEC 24570 version 2</source>
          .1,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>International</given-names>
            <surname>Standards</surname>
          </string-name>
          <string-name>
            <surname>Organisation</surname>
          </string-name>
          , ISO/IEC 24570:
          <fpage>2005</fpage>
          -
          <article-title>Software Engineering - NESMA functional size measurement method version 2.1 - definitions and counting guidelines for the application of Function Point Analysis</article-title>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>