<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Analyzing the Performance of Two COSMIC Sizing Ap- proximation Techniques Using FUR at the Use Case Level</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco Valdés-Souto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Autonomous University of Mexico Science Faculty CDMX</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>For accurate results, standards for the measurement of the functional size of software require that the functionality to be measured be fully known. However, when estimating in the early phases of software development where there is a lack of detail, approximate sizing techniques must be used. An approximation mechanism that has proven useful when there is no historical data is the technique of approximation by EPCU, there are two EPCU contexts with the range of the output variable other than 16.4 CFP and 44 CFP. Previous studies have shown that when functional requirements are at a granularity level of Functional Process, the context recommending being applied is that the output variable has a cut-off at 16.4 CFP, this is done when comparing the distribution of approximation results against the distribution of the REAL sizes. This paper investigates the two EPCU contexts defined in the literature, seeking to identify which technique appears to better represent the distribution of the REAL sizes when the granularity level was Use Cases (UC), the 'Equal Size Bands' (ESB) approximation and fuzzy logic-based approximation technique (EPCU) were also compared to identify which technique appears to represent the distribution of the REAL sizes better, when the granularity level was Use Cases (UC). From the results, it is not clear which approximation technique has the best performance, however carrying out the non-parametric test, it is possible to confirm statistically that the distribution of the EPCU44 approximation technique displays behavior similar to that of the distribution of the COSMIC REAL sizes.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>COSMIC ISO 19761</kwd>
        <kwd>Approximate Sizing</kwd>
        <kwd>Functional Size</kwd>
        <kwd>EPCU Model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>Functional Size Measurement (FSM) methods work best when the information to be
measured – the functional user requirements – is fully known. Santillo [1], for instance,
indicates that the “functional size of software to be developed can be measured
precisely [only] after the functional specification stage: this stage is often completed
relatively late in the development process.” However, when estimating in the early phases
of software development projects, there is often a lack of detailed information, which
hinders the rigorous application of the measurement rules prescribed in international
standards [1, 2, 3].</p>
      <p>As observed by Desharnais et al. [4], when software documentation is lacking, it is
not possible to apply all of the detailed measurement rules as specified in the
international standards for the measurement of the functional size of software. Thus, in such
early phases of the development cycle, to tackle this lack of detail and determine a
relevant range of candidate functional size, measurers must fall back on approximation
techniques for sizing requirements.</p>
      <p>As Vogelezang points out [5], “a rapid size measurement will be acceptable if it can
be produced faster and still can deliver a reliable approximation of the detailed size
measurement.”</p>
      <p>Most currently available approximation techniques for sizing the functional size of
software requiring calibration employ historical data for better results in local contexts,
such as the Equal Size Bands (ESB) approach described in [11]. However, collecting
such data may be both expensive and time-consuming [8], and approximation
techniques based on historical data are of little use without such data. This situation
frequently occurs in the software industry. Additionally, COSMIC size approximation
techniques were initially developed with a small sample of Functional Process (FP)/Use</p>
      <sec id="sec-2-1">
        <title>Cases (UC).</title>
        <p>
          To tackle this situation, a different approximation approach using fuzzy logic,
referred to as the EPCU COSMIC size approximation technique was proposed by Valdés
et al. [
          <xref ref-type="bibr" rid="ref12 ref19 ref23">9, 10, 11</xref>
          ]. This approach does not require local calibration and is useful when
there are no historical data available. Additionally, it is less expensive than the
calibration of the ESB approach or any other approximation approach that requires historical
data [
          <xref ref-type="bibr" rid="ref12 ref19 ref23">8, 9, 10</xref>
          ].
        </p>
        <p>Research on the EPCU size approximation technique has focused on two granularity
levels [11, 12] of the Functional User Requirements (FUR) description: Functional
Process [7] and Use Case [12], with different EPCU context definitions, especially about
changing the domain of its output variable function.</p>
        <p>
          In order to analyze which of both EPCU contexts utilized and previously
documented [
          <xref ref-type="bibr" rid="ref12 ref19 ref23">8, 9, 10</xref>
          ] exhibits, a better performance for each granularity level of the FUR
description, in 2017 Valdés [13] investigated and compared using non-parametric
testing, which of the EPCU contexts (with upper size boundaries at 16.44 CFP1 as defined
in [9] and 44 CFP as defined in [
          <xref ref-type="bibr" rid="ref12 ref19 ref23">10</xref>
          ]) appears to better represent the distribution of the
REAL sizes, when the granularity level description was Functional Process.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>1 In this paper when functional size unit is CFP, the version is v4.0.1.</title>
        <p>This paper presents a case study with a more extensive set of Use Cases aiming to
identify which of the approximation techniques (the ESB technique and the EPCU
technique using two distinct upper size boundaries) perform best, which means, statistically
demonstrating which values distribution from the approximation techniques is more
similar to REAL functional size distribution employing the standard COSMIC method,
when the functional requirements are at the granularity level of Use Cases, a situation
that presents very often in the industry.</p>
        <p>It is known that there is no standard definition for Use Case; however, it has been
observed that frequently, Use Cases correspond to more than one Functional Process,
considering the results in [13], where the EPCU context with upper size boundaries at
16.4 CFP (EPCU16.4) appears to represent the distribution of the REAL sizes better,
and when the granularity level description was Functional Process, the hypothesis for
this work was the following:</p>
        <p>H: The EPCU context with upper size cut-off at 44 CFP (EPCU44) better represents
the distribution of the REAL sizes, when the granularity level of the functional user
requirements description was Use Cases.</p>
        <p>The structure of this paper is organized as follows. Section II presents related work.
Section III presents the experiment. Section IV presents the data including statistical
analysis, while Section V, the conclusions with suggestions for further work.
2
2.1</p>
        <p>Related work on functional size approximation techniques</p>
        <sec id="sec-2-2-1">
          <title>Approximation techniques based on averages</title>
          <p>The IFPUG Function Point Analysis approximation technique for sizing was initially
proposed in 1992 by Bock [16]. In 1997, Meli [14] proposed two variants but did not
report on their performance.</p>
          <p>In 2003, Desharnais et al. [4] analyzed two approximation techniques commonly
used in the industry: Function Points Simplified (FPS) [15], and Backfiring from lines
of code [16]. Using the detailed data from this study (e.g., 90 business information
projects from five organizations), the FPS technique, with average weights for each of
the five function types of the IFPUG Function Points method, exhibited better
performance (MMRE = 10.4%2 and PRED (0.15) = 76.2), while results from the Backfiring
approach were highly inconsistent.</p>
          <p>In 2004, Conte et al. [3] extended the Early &amp; Quick (E&amp;Q) technique to the
COSMIC FSM method and indicated that further tests would be needed to make
adjustments to the proposal, or to confirm it. This E&amp;Q technique is based on (direct)
analogy and (derived) analysis. It is a human-based size approximation technique
impacted by the ability to “recognize” which components of the system belong to the
proposed classes [17].</p>
          <p>Since 2007, in the COSMIC document “Related Topics” [18] that evolved in 2015
into the COSMIC Guideline for Early or Rapid COSMIC Functional Size Measurement</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>2 MMRE, PRED(0.15) calculations using the detailed data from [4]</title>
        <p>[6], two approximation techniques were based on averages where documented: the
average Functional Process approach, and the average Use Case approach.
2.2</p>
        <sec id="sec-2-3-1">
          <title>Approximation techniques based on size bands</title>
          <p>In 2007, in a study of 50 projects, Vogelezang et al. [5] reported on a proposed size
approximation technique based on size bands using the quartile approach. The authors
also investigated the influence of distinct factors in approximate sizing and reported
that, within this sample, the sole factor that exerted a substantial influence on the size
of an average Functional Process in each of the quartiles was the number of Functional
Processes [5]. In their case study, a reference software system with a full set of stable
requirements and stated measured functional size was available.</p>
          <p>In general, an approach to approximate the size of a scaling factor for FUR type(s)
of artifact(s) must be defined locally [18]. This requires, for instance, that an average
size of the artifacts to be measured be established locally.</p>
          <p>This scaling factor represents the size that one can expect to be measured when FUR
are at a level of detail where an accurate measurement can be made because all
necessary details are available [5]. This solution requires historical data to produce an
adequate scaling factor. In 2011, Santillo [1] proposed the Early and Quick COSMIC sizing
approximation, based on earlier work [3] and the Analytic Hierarchy Process [19], a
technique, which provides a means for making choices among sizing alternatives.</p>
          <p>In 2013, Almakadmeh [17] designed a framework to assign scaling factors for
identifying the granularity level of documentation of the functional requirements. Two
variants of criteria for assessing granularity levels were defined: the first considered a
functional component of software, and the second, the elements of a UML use-case
model. To rank the levels of granularity identified, the scaling factors used in [5] were
selected. Next, scaling factor assignment was based on conducting an analogy-based
comparison with similar pieces of software in which the functional size of the software
pieces was accurately measured using the COSMIC measurement method.</p>
          <p>In 2014, De Vito et al. [20] proposed a simplified measurement process
(Quick/Early) that addressed the need for a simplified and rapid COSMIC measurement
avoiding the use of scaling factors, where incorrect calibrations of scaling factors can
lead to inaccurate approximations. The Quick/Early approximation approach can be
applied on Use Case models to reduce measurement time. Quick/Early precision is
directly proportional to the granularity level of the Use Case model analyzed. This means
that Use Cases require stable requirements that, however, do not occur too frequently
in the early stages. Nonetheless, the authors concluded that Quick/Early accuracy is
adequate.
2.3</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Approximation techniques base on fuzzy logic</title>
          <p>In 2012, Valdés et al. [9] proposed a COSMIC size approximation solution using a
fuzzy logic model referred to as the Estimation of Projects in a Context of Uncertainty
(EPCU) [2, 21, 22].</p>
          <p>
            The advantages of the EPCU size approximation technique can be summarized as
follows [
            <xref ref-type="bibr" rid="ref12 ref19 ref23">8, 9, 10</xref>
            ]:
─ Does not require local calibration and is useful when there are no historical data
available.
─ Less expensive to calibrate than the ESB approach, which requires historical data.
─ Exhibits good behavior, even when individuals are not acquainted with the COSMIC
method.
─ Exhibits good behavior, even when requirements are not fully known.
─ Enables systematic replication of the information.
          </p>
          <p>In these studies [2, 21, 22], two EPCU contexts were defined for a continuous range
of possible values with a “natural” upper boundary, or cut-off instead of size bands, and
a mixture of granularity levels (Functional Process and Use Case), simulating the early
phases of the software life cycle:
1. The first EPCU context, defined a cut-off at 16.4 CFP [8, 9] (EPCU16.4), based on
the ESB approach as defined by Vogelezang [5] (Small = 4.8 CFP, Medium =7.7
CFP, Large = 10.7 CFP, and Very Large = 16.4 CFP), and
2. The second context defined a cut-off at 44 CFP [11] (EPCU44), defined after
analyzing the database used by Vogelezang [5], that contains two general analyses over
the functional process measured labeled Q-Size and Q-Number. Considering the
QSize where the total measured size is divided into quartiles and the average FP size
is calculated from each one (Q1=3.7 CFP, Q2=7.7 CFP, Q3=14.6 CFP and Q4=44.1
CFP)</p>
          <p>For this new study, it is considered the integrated analysis, the concept of both is
described below.</p>
          <p>EPCU approach research also focused on the definition of the EPCU context,
selecting several samples from case studies, usually an industry or reference project with
fewer than 12 practitioners, focusing on analyzing the performance of the
approximation technique in the early phases.</p>
          <p>
            For instance, Valdés et al. [
            <xref ref-type="bibr" rid="ref12 ref19 ref23">10</xref>
            ] reported on a case study of a simulation of early
approximation using the EPCU model for an industry project for which only the names
of the Use Cases were made available to participants. This case study confirmed that
the EPCU size approximation approach does not require local calibration and is useful
when there are no historical data available. Besides, it proved less expensive than
calibration of the ESB approach, which requires historical data. In this case study, the
output variable was defined for a continuous range of possible values with an upper
boundary, or cut-off instead of size bands, at 16.4 CFP, as per the ESB approach defined by
Vogelezang et al. [5]. For a case study with a REAL industrial project, the EPCU size
approximation technique yielded better results than the ESB approach, while both
techniques led to lower sizes than the real functional size.
          </p>
          <p>
            In 2015, Valdés et al. [11] proposed another version of their fuzzy logic size
approximation technique. It defined a continuous range of possible values for the output
variable with an upper Q4 (4th Quartile) cut-off of 44 CFP for a Functional Process using
the dataset of Vogelezang et al. [5]. For the study of an industry project that considered
Use Case granularity level, the EPCU cut-off at 44 CFP [11] yielded better results on
comparison with the ESB approach and EPCU cut-off at 16.4 CFP [
            <xref ref-type="bibr" rid="ref12 ref19 ref23">10</xref>
            ]. The Functional
size was underestimated for Functional Process or Use Cases using the EPCU cut-off
at 16.4 CFP. On the other hand, results were above and below the REAL value for Use
Cases using the EPCU cut-off at 44 CFP. More realistic results were obtained using the
          </p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>EPCU44.</title>
        <p>Research on the EPCU size approximation technique has focused on two granularity
levels [11, 12] of the FUR description: Functional Process [7], and Use Case [12],
using two EPCU context definitions; however, it was not clear when to utilize each
EPCU context (EPCU16.4, EPCU44), in order to analyze which of the two has a better
performance for each granularity level of functional requirements. In 2017, Valdés [13]
investigated and compared using a non-parametric test, which of the EPCU contexts
appeared to represent the distribution of the REAL sizes better, when the granularity
level was Functional Process.</p>
        <p>In the study [13], it was statistically demonstrated that distribution for approximation
values using EPCU16.4 was similar to REAL value distribution employing the standard</p>
      </sec>
      <sec id="sec-2-5">
        <title>COSMIC method with 180 Functional Process.</title>
        <p>There is no standard definition for Use Case, and it has been observed that frequently
that Use Cases involve more than one Functional Process, sounds logical that the EPCU
approximation technique with a cut-off of 44 CFP might be more useful if functional
requirements are at the granularity level of Use Cases, a situation that occurs very
frequently in the industry. However, based on the findings of [13], the valid conclusion
is that the EPCU44 approach is not as useful with the Functional Process level of
granularity, as it leads to oversizing, and a similar assessment, but employing Use Cases, is
proposed as further work.
2.4</p>
        <sec id="sec-2-5-1">
          <title>Smmary of COSMIC approximation techniques</title>
          <p>The validity of the majority of approximation techniques is dependent on the
representativeness of the samples with respect to the software being approximated. In other
words, the majority of approximation methods require local calibration, and this
requires local historical data. Even more COSMIC size approximation techniques were
initially developed with a small sample of data. However, as pointed out by
Morgenshtern [8]: “Algorithmic models need historical data, and many organizations do not have
this information. Additionally, collecting such data may be both expensive and
timeconsuming.” Approximation techniques based on historical data are of little use for
organizations without such data. Alternatives must, therefore, be developed for such
contexts of approximation.</p>
          <p>The COSMIC Guideline for Early or Rapid COSMIC Functional Size Measurement
[6] integrates several techniques for the approximate sizing of new, ‘whole’ sets of
requirements. The approximation techniques described in [6] include approximation
techniques based on size bands or based on average.</p>
          <p>The majority of the techniques presented in [6] are based on the existence of
historical data to determine the scaling factor (average, or size bands) or another calibration,
and that there are stable requirements [11].
2.5</p>
        </sec>
        <sec id="sec-2-5-2">
          <title>Impact of approximated size on the estimation of effort</title>
          <p>In 2013, De Marco et al. [23] investigated to what extent some COSMIC-based
approximate sizing could be useful for project managers for early effort estimation for Web
applications. The authors reported an empirical analysis employing data from 25 Web
applications to assess whether two approximate sizes (number of COSMIC Functional
Processes (FP) or the Average Functional Process approach) could be exploited to
acquire accurate effort estimates. These authors concluded that COSMIC-based
approximate sizing was a suitable approach for early effort estimates, while estimates obtained
with approximate sizes were worse than those achieved employing the size obtained
from the application of the standard COSMIC method.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiment with approximation techniques</title>
      <p>This section describes the experiment carried out to evaluate the size approximation
techniques and identify which technique appears to represent the distribution of the</p>
      <sec id="sec-3-1">
        <title>REAL sizes better, when the granularity level was Use Cases (UC).</title>
        <p>3.1</p>
        <sec id="sec-3-1-1">
          <title>Context and participants</title>
          <p>As a part of a consultancy project whose objective was to implement the use of
COSMIC for a Government entity in Mexico carried out in 2016, with the objective of
generating formal estimation models, several projects were measured using the</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>COSMIC method.</title>
        <p>The three main circumstances described in [6], in which only an approximate
COSMIC functional size may be possible were presented in the project:
─ When a size measurement is needed rapidly, and an approximate size measurement
is acceptable if it can be measured much faster than with the standard method. This
is known as ‘rapid sizing’;
─ Early in the life of a project before the actual requirements have been specified in
enough detail for precise size measurement. This is known as ‘early sizing’;
─ In general, when the quality of the documentation of the actual requirements is not
sufficiently good for precise size measurement.</p>
        <p>Considering the information below, the functional size for the projects was gathered
using the approximation approaches as the first step and then, when the required detail
for the requirements was accomplished, the full standard was used to obtain the
functional size.</p>
        <p>To conduct a comparison with the previous study [13] focused on 180 Functional
Process, four projects were selected. These four projects integrated 293 Use Cases that
were approximated using ESB and EPCU techniques.</p>
        <p>The people in the Government entity received 24 hours of training in COSMIC
during the consultancy project, including the EPCU approximation technique and that of
equal size bands. The information required for using the approximation techniques were
required from the technical people, specifically from the project leader for each project,
with a distinct project leader for each project.</p>
        <p>It is important to mention that the techniques related to the Requirements
Engineering used by the Government entity was not affected by the consultancy and was possible
to observe that sometimes the Use Cases include much functionality. Table 1 shows the
number of Use Cases by project.
1. Identify, for each project, the set of Use Cases assigned to be developed.
2. Classify (using expert judgment) by size each of the Use Cases using the following
linguistic values: Small; Medium; Large, and Very Large3.
3. Classify (using expert judgment) the number of objects of interest for each of the</p>
        <p>Use Cases using the following linguistic values: Few; Average, and Many.
4. Assign values (using expert judgment) in the range 0 - 5 ε R for the two previously
classified input variables (points 2 and 3, the Use Cases’ size, the number of objects
of interest related to the Use Cases) defined within the EPCU context, considering
the subjective classification relative to the functional size of the Use Cases (e.g., Step
2), and the subjective classification for the number of objects of interest in each Use
Case (e.g., Step 3).
3 The linguistic values were defined in concordance to the ESB Approach to enabled the
comparison.
5. Measure functional size using the COSMIC method and provide the size for each</p>
      </sec>
      <sec id="sec-3-3">
        <title>Use Case.</title>
        <p>3.3</p>
        <sec id="sec-3-3-1">
          <title>Data collected by participants</title>
          <p>Project leaders identified 293 Use Cases in four projects (Table 1), and the data
provided by the project leaders were the following (see Appendix I for details):
─ A value assigned within the range of 0 - 5 ε R for the size of each Use Case.
─ A value assigned within the range of 0 - 5 ε R for the objects of interest for each Use</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Case.</title>
        <p>─ COSMIC size using the COSMIC method for each Use Case.</p>
        <p>The linguistic classification of the Use Cases and the linguistic classification of the
objects of interest for each Use Cases (data from Steps 2 and 3) were not included in
the table in the Appendix since the input for the EPCU approximation approach were
the values assigned for each variable (data from Step 4).
3.4</p>
        <sec id="sec-3-4-1">
          <title>Researcher steps</title>
          <p>Using the linguistic classification (Small, Medium, Large, and Very Large) assigned
by the participants for the Use Cases the ESB technique was performed.</p>
          <p>Using the values (between 0 and 5) assigned by the participants for the two input
variables of the fuzzy logic based EPCU approximation technique, CFP units were
performed by the researcher using the EPCU approximation technique with distinct EPCU
contexts (EPCU16.4 and EPCU44) defined in [8, 9] and [11].</p>
          <p>The COSMIC size approximated with the data provided by the project leaders was
verified using the COSMIC measurement principles and rules by two consultants with
more than 7,000 CFP measurement experiences at the verification moment.</p>
          <p>
            COSMIC functional size and approximate size for each Use Case are presented in
Appendix II where:
─ Column 1 presents the Project identifier. For confidential purposes, the Projects were
labeled sequentially, from “Proj 1” to “Proj 4.
─ Column 2 presents the Use Case identifier. For confidential purposes, the Use Cases
were labeled sequentially, from “UC 1” to “UC 293.
─ Column 3 presents the functional size obtained utilizing the standard COSMIC
method – in CFP units,
─ Column 4 presents the Equal Size Band approximation approach,
─ Column 5 presents the EPCU size approximation approach using an output variable
domain function from 2 - 16.4 CFP [8] [9], and
─ Column 6 presents the EPCU size approximation approach using an output variable
domain function from 2 - 44 CFP [
            <xref ref-type="bibr" rid="ref12 ref19 ref23">10</xref>
            ].
4
4.1
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Analysis</title>
      <sec id="sec-4-1">
        <title>Quality Criteria</title>
        <p>Three most frequently quoted quality criteria [24] were used to analyze the behavior of
the two approximation techniques :
─ Mean Magnitude of Relative Error (MMRE),
─ Standard Deviation of MRE (SDMRE), and
─ Prediction level, here PRED(25%) was selected.</p>
        <p>The Median Magnitude of Relative Error (MdMRE) is also used. The primary
advantage of the median over the mean is that the median is not sensitive to the outliers.</p>
        <p>Table 2 presents the results for each of these quality criteria for each approximation
approach (top line) for the set of 293 Use Cases:
1. With an MMRE of 61.4%, the ESB presented the best results (in comparison to
MMRE = 65.7% with the EPCU16.4 technique and MMRE = 117.4% with
EPCU44).
2. With an SDMRE of 49.1%, ESB presents the best results, in comparison to SDMRE
of 62.2% for the EPCU16.4 technique and SDMRE = 156.1% for the EPCU44
technique.
3. Within a PRED (25%) at 20.8%, the EPCU with a cut-off at 44 CFP presents the
best results, in comparison to 18.8% with ESB and 17.1% with EPCU with the
cutoff at 16.4 CFP.
4. With a MdMRE of 56.9%, the EPCU16.4 technique presents the best results, in
comparison to 59.5% with ESB and 63.3% with EPCU with a cut-off of 44 CFP. It is
possible to observe that the difference between the maximal and the minimal</p>
        <sec id="sec-4-1-1">
          <title>MdMRE values are less than the other quality criteria.</title>
          <p>Two quality criteria present the best results in the ESB approach (MMRE and
SDMRE); however, the prediction level presents the best results for the EPCU with the
cut-off at 44 CFP, and the MdMRE presents the best results for the EPCU16.4.</p>
          <p>From the quality criteria, it is not clear which approximation technique has the best
performance, this because the central tendency measurements are affected by outliers.</p>
          <p>MMRE has been shown to be a biased estimator of central tendency of the residuals
of a prediction system because it is an asymmetric measure [25], [26], [27], [28].
Shepperd et al. [29] proposed the Mean Absolute Residual (MAR), which, unlike MMRE,
is not biased to compare the accuracy of a given estimation method P against the
accuracy of a reference estimation method P0.
(1)
(2)
(3)
MAR =
1 
 ∑ =1 |
−  ̂ |</p>
          <p>Based on the calculated MARP (the MAR of the proposed method) and MARP0 (the
MAR of a reference method), Shepperd et al. [29] propose to compute a Standardized</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>Accuracy measure (SA) for estimation method P.</title>
          <p>Where values of SA close to 1 indicate that P outperforms P0, values close to zero
indicate that P’s accuracy is similar to P0’s accuracy, and the negative values indicate
that P is worse than P0. The authors [29] suggest to use a referenced model random
based considering the known (actual) values of previously measured projects, however,
Lavazza [30] observed that the comparison with random estimation is not very effective
in supporting the evidence that P is a good estimation model. Instead, proposed to use
a “Constant Model” (CM), where the estimate of the size of the ith project is given by
the average of the sizes of the other projects, then the calculation of the MARCM of
these estimates is realized, and then the compute of SA, comparing method P with a
method CM, generalizing that SA could be used to compare an estimation method P
against any other method P1 used as a reference method.</p>
          <p>SA = 1 −
SA = 1 −</p>
          <p>MAR
SA</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>Calculated using (1), The MAR for ESB = 17.7</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Calculated using (3)</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>Considering ESB as P1</title>
          <p>EPCU 16.4 EPCU 44
to zero (0.05 for EPCU 16,4 and 0.01 for EPCU 44), both EPCU context present similar
accuracy to the reference approximation approach (ESB). Considering the SA measure,
the ESB present a better result, it is not clear which approximation technique has the
best performance.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Graphical Analysis</title>
        <p>Real v.s. ESB</p>
        <p>150
Real</p>
        <p>ESB</p>
        <p>Real v.s. ESB
50
100
200
250
300
400
350
300
250
200
150
100
50
0
100
80
60
40
20
0
0
0</p>
        <p>150
Real</p>
        <p>ESB
50
100
200
250
300</p>
        <p>Note that with the ESB technique, the only four values possible for the approximated
size (in orange) are as follows: 4.8 CFP; 7.7 CFP; 10.7 CFP, and 16.4 CFP
corresponding to the four average size bands of Functional Process (Small, Average, Large, and</p>
        <sec id="sec-4-2-1">
          <title>Very Large).</title>
          <p>From the data (Appendix II, column 3), it is possible to conclude that 230 Use Cases
(78.5%) were underestimated; in consequence, overestimated 63 Uses Cases are
(21.5%). From these overestimated Use Cases, 139 are due to that the upper boundary,
or cut-off, was established at 16.4 CFP and the Use Cases had a functional size higher
than that of the cut-off.</p>
          <p>Fig. 2 depict the graphical comparison with the EPCU16.4 technique. This technique
defines a continuous range of possible values between 2 CFP and an upper boundary
or cut-off at 16.4 CFP; consequently, at least 139 Use Cases were underestimated
because of the upper boundary.</p>
          <p>Looking at the data from (Appendix II, column 3), overestimated Uses Cases
numbered 99 (33.8%), while underestimated Use Cases numbered194 (66.2%). It is possible
to observe that the number of Use Cases underestimated decrease in 36 Use Cases
considering the ESB technique, and the Use Cases overestimated increase.</p>
          <p>Fig. 3 presents the graphical comparison with the EPCU44 technique because this
approach has a cut-off at 44 CFP; naturally, fewer Use Cases were underestimated, 130
(44.4%), while overestimated Use Cases numbered 163 (55.6%), and for the EPCU44
technique, more Uses Cases were overestimated.</p>
          <p>Intuitively from the previous figures, the EPCU44 better represents the distribution
of the REAL sizes; however, it is not easy to infer from Fig.1 to Fig. 3, because there
are several outliers. This confirms the reason regarding the big difference between the
maximal and the minimal values for MdMRE and MMRE from Table 2.</p>
          <p>Considering the difference between MdMRE and MMRE, it is possible to assume
that the distribution is skewed and that the most representative value is the MdMRE,
because central tendency measurements were affected by the outliers.</p>
          <p>In Fig. 4, the boxplots related to the REAL Value of functional size, and ESB,
EPCU16.4, and EPCU44 functional size approximation, are presented. This is a better
approach for analyzing the data without considering the outliers.</p>
          <p>From Fig. 4, it might be easier to infer that EPCU44 better represent the distribution
of the REAL sizes, because both boxplots are very similar.
Real</p>
          <p>EPCU16.4</p>
          <p>Real v.s. EPCU16.44
50
100
150
200
250</p>
          <p>300
Real
400
350
300
250
200
150
100
50
0
100
80
60
40
20
0</p>
          <p>Real v.s. EPCU44</p>
          <p>Real</p>
          <p>Real v.s. EPCU44
0
50
100
150
200
250
300
0
50
100
150
200
250</p>
          <p>300
Real
Considering the quality criteria affected by the central tendency measurements, the
approximation technique that provides better results was ESB. From the plots in Fig.s 1,
2, 3, and 4, the EPCU44 technique appears to better represent the distribution of the
REAL sizes; however, this needs to be confirmed by statistical analysis.</p>
          <p>In non-parametric statistics, a well-known procedure for testing the differences
among more than two related samples is the Friedman test [24, 25] The objective of the
test is to determine whether it can be concluded, from a sample of results, that there is
a difference among treatment effects [32].</p>
          <p>H0: There are NO meaningful differences in the distributions of REAL, ESB,</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>EPCU16.4, and EPCU44 datasets.</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>In consequence, the alternative hypothesis was defined as:</title>
          <p>H1: At least one distribution (REAL, ESB, EPCU16.4, and EPCU44) is significantly
different. A significance level of ɑ (alpha (ɑ)) = 0.05 was assumed.</p>
          <p>SPSS® version 22 software in the Spanish language was utilized to evaluate the
Friedman test for the four distinct treatments (REAL, ESB, EPCU16.4, and EPCU44),
and the results are summarized in Table 4. The full results from SPSS ® are presented
in Appendix III.</p>
          <p>In Table 4, “N” represents the 293 Use Cases, “df” represents the degrees of freedom
(with four distinct treatments; the df is 3 (#treatments -1)). Here, the statistical
significance (“Asymp. Sig.” or p-value) is a very small number at E-101, thus below the
required significance level of ɑ =0.05.</p>
          <p>Therefore, the null hypothesis (e.g., H0: There are NO meaningful differences in the
distribution of REAL, ESB, EPCU 16.4, and EPCU 44) is rejected, and it is possible to
state that at least one treatment has a distinct distribution.</p>
          <p>N</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>Chi-Square df</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Asymp. Sig.</title>
          <p>293</p>
          <p>In order to identify where the difference is, a post-hoc test is needed. In this instance,
a post-hoc test assesses the difference between treatments as follows:
─ REAL and ESB
─ REAL and EPCU16.4
─ REAL and EPCU44
─ ESB and EPCU16.4
─ ESB and EPCU44
─ EPCU-16.4 and EPCU44</p>
          <p>Here, the post-hoc test compared two treatments at a time. The Wilcoxon [32] test
was executed using SPSS® software, and the Bonferroni correction [33] was
considered; thus, the ɑ value (ɑ =0.05) was divided by 4 because four distinct treatments were
used. This means that the ɑ was reset at ɑ =0.0125.</p>
          <p>Considering the latter, the null hypothesis H0 for the post-hoc test was:
H0: There are NO meaningful differences between the distributions for the two
treatments compared (see the previous list).</p>
        </sec>
        <sec id="sec-4-2-6">
          <title>In consequence, the alternative hypothesis was defined as:</title>
          <p>H1: The distribution for the two treatments compared is significantly different,
assuming a significance level of ɑ = 0. 0125.</p>
          <p>Table 5 presents the results of applying the Wilcoxon test for two treatments in
SPSS®. Column 1 indicates the comparison, and column 2, the significance for the
Wilcoxon test. The significance value was compared with ɑ = 0.0125 by accepting (&gt;ɑ
= 0.0125) or rejecting (&lt;ɑ = 0.0125) the null hypothesis; the results are presented in
column 3. The full results from SPSS® are presented in Appendix IV.</p>
          <p>From Table 5, with a p-value of ɑ =0.0125, it is possible to confirm statistically that
only the distribution of the EPCU44 approximation technique (with a cut-off at 44 CFP)
displays a behavior similar to the distribution of the COSMIC REAL sizes (REAL
value), considering the granularity level of Use Cases, which graphically could be
observed in Fig. 4.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS</title>
      <p>In this paper, using a large sample of 293 Use Cases from real projects, two
approximation techniques were evaluated to identify which performs best with this dataset
larger, which is larger than previous sets mentioned in related works. This implies
statistically demonstrating which value distribution from the approximation techniques is
more similar to REAL functional size distribution employing the standard COSMIC
method, when the functional requirements are at the granularity level of Use Cases, a
situation encountered very frequently in the industry.</p>
      <p>From the previous work [13], the EPCU context appears to represent the distribution
of the REAL sizes better; when the granularity level was Functional Process, 180
Functional Process were used.</p>
      <p>From our findings related to quality criteria, it is not clear which approximation
technique executes the best performance, this is because the central tendency measurements
are affected by outliers, and the sample has several outliers, as in reality occurs.</p>
      <p>It is well known that there is no standard definition for Use Case, and this could be
a reason for the outliers. For instance, there are Use Cases with more than 100 or 300
CFP. The presence of outliers can be observed in Fig.s 1 - 4, even though, intuitively
from the previous figures, the EPCU44 better represents the distribution of the REAL
sizes. However, it is not easy to infer.</p>
      <p>On carrying out the non-parametric test, it is possible to confirm statistically that
only the distribution of the EPCU44 approximation technique displays behavior similar
to that of the distribution of the COSMIC REAL sizes (REAL value), considering the
granularity level of Use Cases, accepted the following hypothesis:</p>
      <p>H: The EPCU context with an upper size cut-off at 44 CFP (EPCU44) better
represents the distribution of the REAL sizes, when the granularity level of the FUR
description was Use Cases.</p>
      <p>Considering the findings and the previous work, it is possible to define when the
granularity level of the FUR description was Use Cases, with our recommending the
EPCU44 approximation approach, while when the granularity level of the functional
user requirements description was Functional Process, the EPCU16.4 approximation
approach is recommended.</p>
      <p>The research developed in this paper only includes two of the approximation
techniques mentioned in the Guideline for Early or Rapid COSMIC Functional Size
Measurement [6]; others should be investigated as well, using similar experiments.</p>
      <p>Because the spread of the use of agile practices, a similar assessment to that of this
paper but employing User Histories as the granularity level of the functional user
requirements description should be conducted.
10.
11.
12.
13.
14.
15.
16.
17.
18.
19.
20.
21.
22.
23.
24.
25.
26.
27.
28.
29.
30.
31.
32.
33.</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix I. Data Provided by Participants for Each Use Case Identified</title>
      <p>Table A1 shows the data provided by
participants for each Functional Process identified in the
experiment.
ers)</p>
      <sec id="sec-6-1">
        <title>Column 1 presents the Project identifier. For confidentially purposes, the Projects were labeled sequentially, from “Proj 1” to “Proj 4.</title>
      </sec>
      <sec id="sec-6-2">
        <title>Column 2 presents the Use Case identifier. For confidentially purposes, the Use Cases were labeled sequentially, from “UC 1” to “UC 293.</title>
      </sec>
      <sec id="sec-6-3">
        <title>Column 3 presents the functional size obtained</title>
        <p>utilizing the standard COSMIC method – in</p>
      </sec>
      <sec id="sec-6-4">
        <title>CFP units,</title>
      </sec>
      <sec id="sec-6-5">
        <title>Column 4 presents the value assigned for the input variable “Use Case size” for the EPCU approximation technique.</title>
      </sec>
      <sec id="sec-6-6">
        <title>Column 5 presents the value assigned for the input variable “Presence of objects of interest related to the Use Cases” for the EPCU approximation technique.</title>
        <p>Project
ID</p>
        <p>UC ID</p>
        <p>UC ID</p>
        <p>UC ID
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2
Proj 2</p>
        <p>UC 106
UC 107
UC 108
UC 109
UC 110
UC 111
UC 112
UC 113
UC 114
UC 115
UC 116
UC 117
UC 118
UC 119
UC 120
UC 121
UC 122
UC 123
UC 124
UC 125
UC 126
UC 127
UC 128
UC 129
UC 130
UC 131
UC 132
UC 133
UC 134
UC 135
UC 136</p>
        <p>UC ID</p>
        <p>Project
ID</p>
        <p>UC ID
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 3
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4</p>
        <p>UC 169
UC 170
UC 171
UC 172
UC 173
UC 174
UC 175
UC 176
UC 177
UC 178
UC 179
UC 180
UC 181
UC 182
UC 183
UC 184
UC 185
UC 186
UC 187
UC 188
UC 189
UC 190
UC 191
UC 192
UC 193
UC 194
UC 195
UC 196
UC 197
UC 198
UC 199
UC 200
UC 201
UC 202
UC 203
UC 204
UC 205
UC 206</p>
        <p>UC ID
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4</p>
        <p>UC 207
UC 208
UC 209
UC 210
UC 211
UC 212
UC 213
UC 214
UC 215
UC 216
UC 217
UC 218
UC 219
UC 220
UC 221
UC 222
UC 223
UC 224
UC 225
UC 226
UC 227
UC 228
UC 229
UC 230
UC 231
UC 232
UC 233
UC 234
UC 235
UC 236
UC 237
UC 238
UC 239
UC 240
UC 241
UC 242
UC 243
Presence
(level,
not
number) of
objects
of
interUse Case size est
re(value assign- lated to
ment – range the Use
from 0 - 5) Case
(value
assignment –
range
from 0
5)</p>
        <p>UC ID
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4</p>
        <p>UC 244
UC 245
UC 246
UC 247
UC 248
UC 249
UC 250
UC 251
UC 252
UC 253
UC 254
UC 255
UC 256
UC 257
UC 258
UC 259
UC 260
UC 261
UC 262
UC 263
UC 264
UC 265
UC 266
UC 267
UC 268
UC 269
UC 270
UC 271
UC 272
UC 273
UC 274</p>
        <p>UC ID
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4</p>
        <p>UC 275
UC 276
UC 277
UC 278
UC 279
UC 280</p>
        <p>UC 281
Proj 4</p>
        <p>UC 282
Proj 4</p>
        <p>UC 283
Proj 4</p>
        <p>UC 284
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4
Proj 4</p>
        <p>UC 285
UC 286
UC 287
UC 288
UC 289
UC 290
UC 291
UC 292
UC 293</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Appendix II. COSMIC</title>
    </sec>
    <sec id="sec-8">
      <title>Functional Size and</title>
    </sec>
    <sec id="sec-9">
      <title>Approximation</title>
      <sec id="sec-9-1">
        <title>COSMIC functional size and approximation for each Functional Process are presented in Table A2 II where:</title>
      </sec>
      <sec id="sec-9-2">
        <title>Column 1 presents the Project identifier.</title>
      </sec>
      <sec id="sec-9-3">
        <title>For purposes of confidentiality , the Projects were labeled sequentially, from “Proj 1” to “Proj 4,</title>
      </sec>
      <sec id="sec-9-4">
        <title>Column 2 presents the Use Case identifier.</title>
      </sec>
      <sec id="sec-9-5">
        <title>For purposes of confidentiality, the Use</title>
      </sec>
      <sec id="sec-9-6">
        <title>Cases were labeled sequentially, from “UC 1” to “UC 293,</title>
      </sec>
      <sec id="sec-9-7">
        <title>Column 3 presents the functional size obtained utilizing the standard COSMIC method – in CFP units,</title>
      </sec>
      <sec id="sec-9-8">
        <title>Column 4 presents the Equal Size Bands approximation approach,</title>
      </sec>
      <sec id="sec-9-9">
        <title>Column 5 presents the EPCU size approximation approach using an output variable domain function from 2 - 16.4 CFP [9] [10], and</title>
      </sec>
      <sec id="sec-9-10">
        <title>Column 6 presents the EPCU size approxi</title>
        <p>mation approach using an output variable
domain function from 2 - 44 CFP [11].</p>
        <p>Table A2: Functional size – Real and from 3
approximation techniques
RE
AL</p>
        <p>UC</p>
        <p>ID
UC 30
UC 31
UC 32
UC 33
UC 34
UC 35</p>
        <p>UC</p>
        <p>ID</p>
        <p>RE
AL</p>
        <p>Project
ID
UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB
Project
ID</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
        <p>Project
ID
UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB
UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB
4
4
4
4
4</p>
        <p>UC
286
UC
287
UC
288
UC
289
UC
290
UC
291
UC
292
UC
293</p>
        <p>UC
ID</p>
        <p>RE
AL</p>
        <p>ESB</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Appendix III. Friedman Test Results from SPSS®</title>
      <p>Descriptive Statistics
N
293
293
293
293</p>
      <p>Mean
24.7406</p>
      <p>8.5556
10.8944
21.0759
tion</p>
      <p>Std.
DeviaMean Rank
2.89
1.35
2.20
3.55
293
468.936
3
2.57215388100136E101</p>
    </sec>
    <sec id="sec-11">
      <title>Appendix IV. Wilcoxon Test Results from SPSS®</title>
      <p>4.8089036753386E30</p>
      <sec id="sec-11-1">
        <title>EPCU16 - REAL</title>
        <p>Ranks
EPCU16 - REAL
a. EPCU16 &lt; REAL
Negative Ranks
Positive Ranks
Ties
Total
230a
63b</p>
        <p>0c
293
194a
99b
0c
293</p>
        <p>Mean Rank
165.49
79.49</p>
        <p>Sum of Ranks
38063.00
5008.00
Mean Rank
177.95
86.34</p>
        <p>Sum of Ranks
34523.00
8548.00
b. EPCU16 &gt; REAL
c. EPCU16 = REAL
Test Statisticsa
b. Based in positive ranks.
b. Based in positive ranks.
130a
163b</p>
        <p>0c
293</p>
        <p>Mean Rank
153.62
141.72</p>
        <p>Sum of Ranks
19971.00
23100.00
Ranks
EPCU44 - ESB
a. EPCU44 &lt; ESB
b. EPCU44 &gt; ESB
c. EPCU44 = ESB</p>
      </sec>
      <sec id="sec-11-2">
        <title>EPCU16 – ESB</title>
        <p>Ranks
EPCU16 - ESB
a. EPCU16 &lt; ESB
b. EPCU16 &gt; ESB
c. EPCU16 = ESB
Z
Asymp. Sig. (2-tailed)
a. Wilcoxon Test with sign
b. Based in positive ranks.</p>
        <p>MRE_EPCU44
ESB</p>
        <p>-14.835b
8.7473E-50
1a
292b</p>
        <p>0c
293
39a
254b</p>
        <p>0c
293</p>
        <p>Mean Rank</p>
        <p>4.00
147.49</p>
        <p>Sum of Ranks</p>
        <p>4.00
43067.00
Mean Rank
68.58
159.04</p>
        <p>Sum of Ranks
2674.50
40396.50</p>
        <p>MRE_EPCU16
ESB</p>
        <p>Sum of Ranks</p>
        <p>0.00
43071.00
Z</p>
        <p>Asymp. Sig.
(2-tailed)
a. Wilcoxon Test with sign
b. Based in positive ranks.</p>
        <p>MRE_EPCU16
MRE_EPCU44</p>
        <p>-14.839b</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Santillo</surname>
          </string-name>
          ,
          <article-title>Early and Quick COSMIC FFP Overview</article-title>
          , in: A.
          <string-name>
            <surname>A. Reiner Dumke</surname>
          </string-name>
          (Ed.),
          <source>Cosm. Funct. Points Theory Adv. Pract</source>
          ., CRC Press, Boca Raton, FL, USA,
          <year>2011</year>
          : pp.
          <fpage>176</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Valdés-Souto</surname>
          </string-name>
          ,
          <article-title>Design of a Fuzzy Logic Software Estimation Process</article-title>
          , École De Technologie Supérieure, Université Du Québec,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Conte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Iorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Santillo</surname>
          </string-name>
          ,
          <string-name>
            <surname>E&amp;Q:</surname>
          </string-name>
          <article-title>An Early &amp; Quick Approach to Functional Size Measurement Methods</article-title>
          , in: Istituto di Ricerca Internazionale (Ed.),
          <source>Softw. Meas. Eur. Forum SMEF</source>
          <year>2004</year>
          , Rome, Italy,
          <year>2004</year>
          : p.
          <fpage>416</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>J.M. Desharnais</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Abran</surname>
          </string-name>
          ,
          <article-title>Approximation Techniques for Measuring Function Points</article-title>
          ,
          <source>Proc, in: 13th Inter. Work. Softw. Meas. (IWSM</source>
          <year>2003</year>
          ), Springer, Montreal, Canada,
          <year>2003</year>
          : pp.
          <fpage>270</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Vogelezang</surname>
          </string-name>
          , T. Prins,
          <article-title>Approximate size measurement with the COSMIC method: Factors of influence</article-title>
          ,
          <source>in: Softw. Meas. Eur. Forum SMEF</source>
          <year>2007</year>
          , Rome, Italy,
          <year>2007</year>
          : pp.
          <fpage>167</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Common Software Measurement International</surname>
          </string-name>
          <article-title>Consortium (COSMIC)., Guideline for Early or Rapid COSMIC Functional Size Measurement</article-title>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Common Software Measurement International Consortium</surname>
          </string-name>
          (COSMIC),
          <source>Measurement Manual v4.0.1</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>O.</given-names>
            <surname>Morgenshtern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Raz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dvir</surname>
          </string-name>
          , Factors Affecting Duration and doi:10.1109/SEAA.
          <year>2014</year>
          .
          <volume>30</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>Int. Conf. IWSM-Mensura</source>
          <year>2007</year>
          ,
          <string-name>
            <surname>UIB-Universitat de les Illes</surname>
            <given-names>Baleares</given-names>
          </string-name>
          , Illes Baleares, Spain,
          <year>2007</year>
          : pp.
          <fpage>87</fpage>
          -
          <lpage>101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          2010, Springer-Verlag, Berlin,
          <year>2010</year>
          : pp.
          <fpage>227</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>L. De Marco</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Ferrucci</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Gravino, Approximate COSMIC Size to Early Estimate Web Application Development Effort</article-title>
          ,
          <source>in: 39th Euromicro Conf. Ser. Softw. Eng. Adv. Appl</source>
          . Approx.,
          <source>Conference Publishing Services (CPS)</source>
          , Santander, Spain,
          <year>2013</year>
          : pp.
          <fpage>349</fpage>
          -
          <lpage>356</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>doi:10</source>
          .1109/SEAA.
          <year>2013</year>
          .
          <volume>41</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <article-title>The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance</article-title>
          ,
          <source>J. Am. Stat. Assoc</source>
          .
          <volume>32</volume>
          (
          <year>1937</year>
          )
          <fpage>675</fpage>
          -
          <lpage>701</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Lavazza</surname>
          </string-name>
          ,
          <article-title>Accuracy Evaluation of Model-based COSMIC Functional Size Estimation, in: ICSEA 2017 Twelfth Int</article-title>
          .
          <source>Conf. Softw. Eng. Adv.</source>
          ,
          <year>2017</year>
          : pp.
          <fpage>67</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>B.A.</given-names>
            <surname>Kitchenham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.M.</given-names>
            <surname>Pickard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.G.</given-names>
            <surname>MacDonell</surname>
            , M.J. Shepperd
          </string-name>
          , What accuracy statistics really measure,
          <source>in: IEE Proc. - Softw</source>
          .,
          <year>2001</year>
          : pp.
          <fpage>81</fpage>
          -
          <lpage>85</lpage>
          . doi:
          <volume>10</volume>
          .1049/ip-rsn:
          <fpage>20010506</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Foss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stensrud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kitchenham</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Myrtveit,</surname>
          </string-name>
          <article-title>A simulation study of the model evaluation criterion MMRE, Softw</article-title>
          . Eng.
          <source>IEEE Trans</source>
          .
          <volume>29</volume>
          (
          <year>2003</year>
          )
          <fpage>985</fpage>
          -
          <lpage>995</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>(Bournemouth U. Shepperd</surname>
          </string-name>
          ,
          <article-title>Reliability and validity in comparative studies of software prediction models</article-title>
          ,
          <source>IEEE Trans. Softw. Eng</source>
          .
          <volume>31</volume>
          (
          <year>2005</year>
          )
          <fpage>380</fpage>
          -
          <lpage>391</lpage>
          . doi:
          <volume>10</volume>
          .1109/TSE.
          <year>2005</year>
          .
          <volume>58</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Shepperd</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. MacDonell</surname>
          </string-name>
          ,
          <article-title>Evaluating prediction systems in software project estimation</article-title>
          ,
          <source>Inf. Softw. Technol</source>
          .
          <volume>54</volume>
          (
          <year>2012</year>
          )
          <fpage>820</fpage>
          -
          <lpage>827</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>doi:10</source>
          .1016/j.infsof.
          <year>2011</year>
          .
          <volume>12</volume>
          .008.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Lavazza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Morasca</surname>
          </string-name>
          ,
          <article-title>On the Evaluation of Effort Estimation Models</article-title>
          ,
          <source>in: Proc. 21st Int. Conf. Eval. Assess. Softw. Eng</source>
          . - EASE'
          <fpage>17</fpage>
          ,
          <string-name>
            <surname>Karlskrona</surname>
          </string-name>
          , Sweden,
          <year>2017</year>
          : pp.
          <fpage>41</fpage>
          -
          <lpage>50</lpage>
          . doi:
          <volume>10</volume>
          .1145/3084226.3084260.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Demšar</surname>
          </string-name>
          ,
          <article-title>Statistical Comparisons of Classifiers over Multiple Data Sets</article-title>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mach</surname>
          </string-name>
          .
          <source>Learn. Res</source>
          .
          <volume>7</volume>
          (
          <issue>2006</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fernández</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luengo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <article-title>Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power, Inf</article-title>
          . Sci. (Ny).
          <volume>180</volume>
          (
          <year>2010</year>
          )
          <fpage>2044</fpage>
          -
          <lpage>2064</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>doi:10</source>
          .1016/j.ins.
          <year>2009</year>
          .
          <volume>12</volume>
          .010.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>