<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cases in Source Code to Architecture Mapping using Naive Bayes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tobias Olsson</string-name>
          <email>tobias.olsson@lnu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Morgan Ericsson</string-name>
          <email>morgan.ericsson@lnu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Wingkvist</string-name>
          <email>anna.wingkvist@lnu.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Media Technology, Linnaeus University</institution>
          ,
          <addr-line>Kalmar/Växjö</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Orphan Adoption</institution>
          ,
          <addr-line>Software Architecture, Source Code Clustering, Naive Bayes</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Workshop Proce dings</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Workshop Proceedings</institution>
          ,
          <addr-line>CEUR-WS.org</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The automatic mapping of source code entities to architectural modules is a challenging problem that is necessary to solve if we want to increase the use of Static Architecture Conformance Checking in the industry. We apply the state-of-the-art automatic mapping technique to eight open-source systems and find that there are systematic problems in the automatically created mappings. All of these eight systems have small modules that are very hard to map correctly since only a few source code entities are mapped to these. All systems seem to use some naming strategy, mapping source code to modules; however, naming is often ambiguous. We also find diferences in ground truth mappings performed by experts, which afect mappings based on these, and that architectural refactoring also afects the mapping performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Our previous studies [1, 2] of automated techniques to
map source code entities to high-level software
architectural modules suggest that some entities are much
harder to map correctly than others. Even using the
best algorithm and diferent parameters, certain entities
always seem to fail to map correctly. We conduct an
exploratory study to determine whether our intuition is
correct, i.e., that these hard cases exist, and if they do,
what their properties are, and what makes them hard to
map correctly.</p>
      <p>The software architecture of a system captures major
design decisions at a high level of abstraction and
enportability, reusability, and maintainability [3]. It serves
as a guide for the many decisions that are made during
the implementation of a system. As the system evolves,
the source code must continue to conform to the
architecture or risk accumulating technical debt and no longer
possess the desired qualities.</p>
      <p>Static Architecture Conformance Checking (SACC) is a
collection of methods, such as Reflexion modeling [ 4],
that statically analyze source code to ensure that it does
not introduce architectural violations [5, 6]. These
methods require an architecture model, with modules and
dependencies, and a source code model, with entities
and concrete dependencies, e.g., due to inheritance or
method invocations. They also require a mapping from
CEUR
htp:/ceur-ws.org
ISN1613-073
© 2021 Copyright for this paper by its authors. Use permitted under Creative</p>
      <p>CEUR
the source code model to the architecture model to
determine whether the source code dependencies are
convergent, absent, or divergent compared to the allowed
dependencies specified in the architecture model.</p>
      <p>The need for a mapping between the source code and
architecture models is a significant reason why SACC has
not reached widespread use in the software industry [ 3, 5,</p>
      <sec id="sec-1-1">
        <title>7, 8]; the tools and methods exist, but the mappings do not</title>
        <p>
          or are outdated. Many tools address this by combining
manual mapping and regular expressions to filter file,
module, and package names. Still, such approaches have
proven to be time-consuming and error-prone [
          <xref ref-type="bibr" rid="ref2">5, 7, 8</xref>
          ].
        </p>
        <p>If we want to automate the mapping process using, e.g.,
machine learning, it is vital to understand the hard cases.
map automatically or always maps to the wrong modules,
we need to ensure that these are part of the initial set
that a human expert maps. We perform an exploratory
study using eight systems with ground truth mappings
to determine whether such a class exists. Once we have
established that it exists, we determine its properties to
identify its members automatically. We then investigate
why these properties make the entities dificult to map to
ensure that they will not reduce the efectiveness of the
machine learning approach; we do not want it to learn
the wrong things from the hard cases.</p>
      </sec>
      <sec id="sec-1-2">
        <title>We hypothesize that at least some hard cases would</title>
        <p>be dificult for a human to map and that diferent human
experts would disagree on how they should be mapped.</p>
      </sec>
      <sec id="sec-1-3">
        <title>This can, for example, be due to poor structuring or the</title>
        <p>evolution of the system. We rely on diferent
groundtruth mappings of the same system and metrics to identify
such cases and study how well these correlate to the hard
cases.</p>
      </sec>
      <sec id="sec-1-4">
        <title>Orphan Entity be analyzed to determine its purpose and its similarity</title>
        <p>SOtrrpuhcatunrEalnRtietylattoiotnhsefMroampptehde StringChange to Athesupbu-rpproosbeleomf tohfeomrpohdaunleasd.option is orphan
kidnapGAMrUocdhIuitleectural Entities AMutaop?mpaintegd Logic tppiiivnneggc,alwunshteeenrrtienitgsyo.tftwTozaaerrenpeeowvsomalunotdidoHunloecl,atouirdsieennstoaitfhnyeeare
dfithwfocorrridtrsee,rmciooanrrpe-c</p>
        <p>ChangeScanner AttachFileAction Allowed Module Dependency DOIChek XMLUtil DataBank related to orphan kidnapping, Interface minimization; it
Initially Mapped Set is not a good idea to reassign an entity to another
modFigure 1: An example mapping that shows the initial sets ule if the removal of the entity will cause the module to
of the GUI and Logic modules of JabRef 3.7. A new orphan get a larger public interface, i.e., the entity is an entry
StringChange is about to be mapped. point/facade to the module.</p>
        <p>
          HuGMe [
          <xref ref-type="bibr" rid="ref2 ref6">10, 8</xref>
          ] relies on orphan adoption to map from
the source code to the architecture model. It starts from
an initial set of entities that are manually mapped to the
2. Automated Mapping correct module. The remaining entities are considered
To reason about how well an implementation conforms orphans. HuGMe is applied iteratively, and as the set
to the intended architecture using, e.g., Reflexion mod- of mapped entities can grow for each iteration, more
eling, we need a mapping from the source code to the orphans have the potential to be automatically mapped.
architecture. In this section, we discuss how such a map- In each iteration, there is also the possibility for human
ping can be created semi-automatically, starting from an intervention using the result of the failed automatic
mapinitial set of mapped source code entities. ping attempts as a guideline. The automatic mapping is
        </p>
        <p>The source code model consists of Entities (E) and De- done by calculating the attraction between the orphan
pendencies (ED). The entities are, e.g., classes defined and the mapped entities for each module. Christl et al.
in a programming language, and the ED are due to, e.g., present two attraction functions, CountAttract and
MQAtmethod calls and inheritance, see StringChange, ChangeS- tract, based on dependencies, i.e., the structure criterion.
canner, etc., in Figure 1. Bittencourt et al. evaluate two new attraction functions</p>
        <p>The architecture model consists of Modules (M) and based on information retrieval techniques. They use the
Dependencies (MD) between these. The modules repre- names of modules and entities and the names of
idensent the major parts of the architecture; see, e.g., GUI tifiers in the entities to form vocabulary documents for
and Logic in Figure 1. The directed MD indicates how modules and entities, i.e., the naming and semantic
critethese modules are allowed to interact and depend on each ria. They then use a cosine similarity function, IRAttract,
other. If there, for example, is an MD from GUI to Logic, and latent semantic indexing, LSIAttract, to calculate the
then entities mapped to GUI are expected to call entities attraction values.
mapped to Logic. Our attraction function, NBAttract, combines ideas</p>
        <p>An automated mapping algorithm aims to map each from the previous two and considers the structure,
namentity to the correct module without human assistance. ing, and semantic criteria [2]. The approach is similar
For example, classes in the implementation that deal with to that of Bittencourt et al., but we instead use a Naive
the application’s business rules should be mapped to the Bayes classifier to determine similarity to other entities.
module Logic. Once this mapping exists, we can compare To include the structure criterion, NBAttract uses a novel
the ED of the implementation to the MD allowed by the approach, Concrete Dependency Abstraction (CDA), to
architecture and determine whether they are convergent, encode dependencies as text [2]. NBAttract has
outperabsent, or divergent [4]. formed CountAttract in our previous study [2], and
Coun</p>
        <p>
          We rely on orphan adoption [
          <xref ref-type="bibr" rid="ref4">9</xref>
          ] to map entities to mod- tAttract was not clearly outperformed in [7]. We,
thereules automatically. An unmapped entity is considered an fore, only use NBAttract in the remainder of this paper.
orphan that should be adopted by one of the modules, e.g.,
StringChange in Figure 1. Tzerpos and Holt identify four 3. Method
criteria that can afect the mapping. Naming, naming
standards can reveal what module is suitable. Structure, Based on our experiences with diferent attraction
funcdependencies between an orphan and already mapped tions, we hypothesize that no matter how well the
funcentities can be used as a mapping criterion. Style, mod- tion performs, there is a specific set of entities that are
ules are often created using diferent design principles always misclassified. We seek to investigate this further
(e.g., high cohesion or not). Classifying the orphan based to determine whether our hypothesis is correct or if the
on style can give hints on how to use, for example, the misclassifications happen by chance due to randomness
structure criteria. Semantics, the source code itself can in the composition and size of the initial set.
        </p>
        <p>
          We have previously implemented a tool to evaluate for misclassification based on our own experience and the
diferent mapping approaches, including reporting de- advice from related work, and present exciting findings
tailed mapping results [
          <xref ref-type="bibr" rid="ref8">11</xref>
          ]. We use this tool to create from the data. The ultimate goal is to construct strategies
a new dataset over the mapping results for each source to detect entities with a high risk of being misclassified so
code entity. that a human can intervene and classify these manually.
        </p>
        <p>We run NBAttract, with the following settings. We More specifically, we will investigate:
use an initial set of mapped entities of random size and Is the set of problematic entities a good candidate for
composition. We extract package names, filenames (these the initial set? This set needs human intervention for
correspond to the outer class names in Java), attribute automatic mapping to perform well, efectively removing
identifier names, and variable identifier names from the the problem from the automatic mapping. This can be
source code entities in the initial set and tokenize these assessed by computing the F1 score of the precision and
based on Camel-case and the characters - and _ . The recall, as we did in [2]. We will compare the F1 scores
tokens are then stemmed using a Porter stemmer. Tokens across the entire range of initial set sizes visually.
that are shorter than three characters are removed. We Is the set of problematic entities related to small modules?
use our CDA technique to represent dependencies as text In general, machine learning techniques need good data
strings. We use a binary token frequency (present or not) to perform. In particular, there is a need for a balanced
and 0.9 as the threshold for automatic classification. dataset where there is approximately the same amount</p>
        <p>These settings correspond to the settings used in [2] of data to learn from in each class. If the dataset is
imwith one exception; we do not require the initial set to balanced, there is a high chance that smaller classes will
contain at least one source code entity from each module not be properly handled. An architectural module should
in this study. We are interested in how individual files contain a fair amount of source code entities. Still, there
are mapped to find possible flaws in the technique, which may exist modules that hold source code entities that do
is why we allow for a module to be empty initially. not fit well in other modules, or the system may be under</p>
        <p>
          As we run several experiments with random initial sets, evolution, and intended source code has not been created
we get a dataset that shows the correct mapping of each yet, etc. We need to know if such small modules exist
entity and the number of mappings for each entity and and whether they are common or problematic.
module. Based on this information, we can compute an Is the set of problematic entities related to entities with
error rate for each entity according to Equation 1. If the poor naming? Tzerpos and Holt [
          <xref ref-type="bibr" rid="ref4">9</xref>
          ] define naming as one
attraction function was completely stochastic, the error of the key criteria that influence the mapping. In our
rate for each entity would converge to the stochastic experience, it is also a common strategy for developers
error rate, defined in Equation 2. to create folders, packages, and filenames that reflect the
modular architecture to some degree. It would thus be
|erroneous mappings| interesting to know if the naming of source code entities
errnba = |mappings| (1) iinntceluredsetisntghteo mknoodwulief’tshneraemaereiat misbmigaupitpieesdintot.heItniasmailnsgo,
i.e., if several module names match the name of a source
|modules| − 1 code entity.
        </p>
        <p>errsto = |modules| (2) borIdsetrheofseat mofopdruolbel?emBiabtiiceetnatli.t[ie1s2r]e,lTazteedrptooseanntidtieHsoolnt [t9h]e,
As NBAttract is not a stochastic function, the    for and Bittencourt et al. [7] state that dependencies have
an entity should converge to something less than    if an impact on the mappings. We use a textual
representathere are no systematic problems, i.e., it should systemat- tion of dependencies in NBAttract, but this may not be
ically produce better mappings than a random mapping. good enough. We will investigate the ratio of external
Hence, we can conclude that there are systematic prob- dependencies, e.g., an entity with many external
depenlems if we do not find such a convergence for a source dencies would likely be an entity that lies on the border
code entity after several iterations. If we find systematic of a module. If we find a correlation between the external
errors in a majority of the systems, we will further an- dependency ratio and the error rate, this could suggest
alyze all problematic entities to find common, possible that border entities are problematic.
causes for the misclassification. An entity is considered There are several metrics based on dependencies. We
problematic if its    ≥ 0.5, i.e., it is misclassified in use coupling (the count of all dependencies to or from
50% or more of the mappings. The motivation for this all other entities) and fan (the existence of a dependency
limit is that a non-problematic attraction function should, to or from all other entities). The coupling may be very
on average, produce a correct mapping in at least 50% high between two entities, but the fan can at most be
of the cases for each entity. This part of the research is one between two entities, i.e., fan is a subset of coupling.
highly exploratory. We investigate the possible reasons While coupling captures the absolute number of
dependencies fan focuses on the diversity of diferent entities, Table 1
i.e., a high fan value captures that an entity has many Mapping Data Overview.
dependencies to other diferent entities.</p>
        <p>Is the set of problematic entities related to problems in the System Lines # Mod # Ent err ≥ 0.5 err ≥ errsto
ground truth mapping? We have access to two versions Ant 36 699 16 468 187 39.96% 72 15.38%
of the JabRef system in which the modules and relations A.UML 62 392 19 767 165 21.51% 74 9.65%
between them are the same (same intended architecture), JR 3.7 59 235 6 1 017 107 10.52% 40 3.93%
but the mappings are not the same for all entities. This JR 3.5 51 840 6 733 96 13.1% 51 6.96%
gprroouvinddestraunthomppaoprptuinngistyantod shtuowdytdhiesscereapfeacntctihees ainuttoh-e SLP.urHocMe3nDe 33549 899164247 479 251616147 613089 1213..663.795%%% 11979 1331...433518%%%
matic mapping performance. One complicating factor in T.Mates 54 904 12 450 115 25.56% 49 10.89%
this analysis is that JabRef underwent an architectural
evolution between these two versions. Therefore, we Entity Error Rates per Project
limit our analysis to entities that remain the same (no 1.0
changes to the source code) but are mapped to diferent
modules.</p>
        <p>Is the set of problematic entities related to files that are 0.8
being refactored due to architectural evolution? The two
versions of JabRef provide an opportunity to study enti- 0.6
ties that have changed packages and mapping (a sign of
architectural evolution), have changes to the source code 0.4
(a sign of refactoring), or were recently added.</p>
        <p>We study eight open-source systems implemented in 0.2
Java. Ant1 is an API and command-line tool for process
automation. ArgoUML2 is a desktop application for UML 0.0
bmiboldioelgirnagp.hJiacbarlerfe3feisreandceess,katonpd wapepulisceattihoen3f.o5ramndan3.a7gvinerg- tnA .LAUM .rJv53 .rvJ73 ceneuL roPM .3SDH t.seaTM
sions. Lucene4 is an indexing and search library. ProM5 Figure 2: The entity error rates for each project.
is an extensible framework that supports a variety of
process mining techniques. Sweet Home 3D6 is an interior
design application. TeamMates7 is a web application for
handling student peer reviews and feedback. 4. Results and Analysis</p>
        <p>Table 1 presents the sizes of the systems in lines of
code, number of entities, and number of modules. There We performed the experiment and collected mapping
exist a documented software architecture as well as a data per entity for each system. All systems show several
mapping from the implementation to this architecture entities always being misclassified (an error rate of 1.0)
for each system. Jabref 3.7, TeamMates, and ProM have (cf. Figure 2). Table 1 shows an overview of the data
been the subjects of study at the Software Architecture collected. Note that each entity has a random chance to
Erosion and Architectural Consistency Workshop (SAE- be included in the initial set and not be an orphan in that
roCon) 2016, 2017, and 2019 respectively, where a system particular run of the experiment. There is also a chance
expert has provided both the architecture and the map- an entity will not be mapped (e.g., due to variations in
ping. The architecture documentation and mappings are the initial set). However, each entity has been mapped at
available in the SAEroCon repository8. ArgoUML, Ant, least 500 times.
and Lucene were studied by Brunet et al. and Lenhard We now construct the initial set using entities with
et al., and the architectures and mappings were extracted    ≥ 0.5, i.e., only entities with    &lt; 0.5 are
confrom the replication package of Brunet et al. as well as for sidered orphans, and all the troublesome entities are
inSweet Home 3D. JabRef 3.5 was extracted from Lenhard cluded in the initial set. We compare this with randomly
et al.. selecting from all entities in the initial set. We collected
14 849 and 13 754 data points from the respective groups.</p>
        <p>Figure 3 shows the running median (±100 data points)
and limits of the running 75th and 25th percentiles of
the F1 scores, respectively, for JabRef 3.7. Since the other
systems show similar trends, so we focus on JabRef. We
ifnd that our idea is promising overall, especially when
the initial set size increase.
1https://ant.apache.org
2http://argouml.tigris.org
3https://jabref.org
4https://lucene.apache.org
5http://www.promtools.org
6http://www.sweethome3d.com
7https://teammatesv4.appspot.com
8https://github.com/sebastianherold/SAEroConRepo
JabRef 3.7 f1 Scores</p>
        <p>Relative Miss-Classifications vs Relative Module Size</p>
        <p>All Entities
Error Rate &lt; 0.5</p>
        <p>However, there is also an interval between the initial all entities are misclassified. Also, note that there are
set sizes of 0.1 to 0.18, marked with vertical lines in Fig- small modules with a relatively low number of
misclasure 3, where the F1 score is considerably lower than using sified entities. Another factor to consider is that only
all entities. This indicates that entities with    ≥ 0.5 321 entities of 4377 (7.33%) are mapped to these small
are not good representatives of modules. Upon further modules.
inspection of the actual modules and entities in JabRef 3.7, To measure the extent of using a naming strategy (NS)
we find that there is a set of modules with very few enti- when naming concrete entities in each system, we check
ties in the ground truth mapping, and entities mapped to if the words in the package or class name for an entity
these all have a high error rate. contain the module name. Table 2 shows that most
sys</p>
        <p>The high error rate makes sense in general, as ma- tems use a naming strategy (column NS) to a rather high
chine learning techniques produce better results if there degree. Lower values (e.g., TeamMates) are often due to
is more data. More specifically, for Naive Bayes, the a module naming discrepancy, e.g., TeamMates defines
probability of finding an entity in such a module is very a module view with a corresponding path word named
low, so it does not make sense to map entities to it. We, ui; however, there are also cases where there is no clear
therefore, investigate if all systems have such small mod- naming strategy for an entity. We consider an entity
ules prone to misclassification. If entities are equally to have ambiguous naming if its path or filename
condistributed among the modules of a system, there would tains several diferent module name words. For example,
be 1/|modules|% entities in each module. We regard a net.sf.jabref.logic.net.ProxyPreferences, from JabRef v3.7,
module as small if it has less than half the number of contains both the module names logic and preferences.
entities of 1/|modules|%. Thus we define the limit for a Ambiguity in entity naming strategy (ANS) seems to be
small module as 0.5/|modules|%. It could be argued that quite common in some systems (Ant, ArgoUML, JabRef
the lines of code should be used as a more fine-grained 3.5, Sweet Home 3D, and TeamMates) and not at all in
measure of size, i.e., mapping one huge entity in terms of others (JabRef v3.7 ProM and Lucene). In some systems,
lines of code. However, for example, path and file name the ambiguity is caused by having a parent-level package
information is per entity, and for efectively learning a that is also a module. For example, Ant uses ant as both
pattern based on entity names, more entities are needed. a high-level package and a module. The misclassification</p>
        <p>Table 2 shows the limit, the number of small modules, rate in the ambiguously named entities (ANSM) seems to
and the rate of misclassification of entities in these mod- follow the inverse pattern of the ANS; the lower the ANS,
ules. A surprising result is that all systems have such the higher the ANSM. This makes sense since a higher
small modules, and all systems have small modules where ANS means there is more data to learn the pattern of the
all entities are misclassified. There are 30 (out of a total of ambiguous naming from (if there is one).
73) modules where all entities are misclassified. Figure 4 We now turn our attention to whether entities that lie
shows how the relative number of misclassifications and on the border of a module, i.e., have relatively many
derelative module size are related. Note the cloud of points pendencies to entities in other modules, are problematic.
in the upper left corner. These are the 30 modules where We use the common coupling and fan metrics. Results for
Change in Error for Entities with Changed Mapping</p>
        <p>pare the performance of diferent approaches, but not
.01 JJaabbRReeff vv33..75 cnoodceocdheacnhgaendge to specifically analyze problematic cases. We highlight
the conclusions of prior work made regarding what may
.08 explain the performance.</p>
        <p>
          The orphan adoption criteria naming, structure, style,
.60 and interface minimization are used in an algorithm
evalrrroE uated in three case studies [
          <xref ref-type="bibr" rid="ref4">9</xref>
          ]. We find an evolving
in.40 dustrial system where the architecture was created by
researchers with the help of developers the most
inter.02 esting of these. 939 entities were assigned to modules,
and in 46 cases (4.9%), the algorithm suggested a diferent
.00 mapping than the developers. In 33 of these cases, the
1 2 3 4 5 6 7 8 9 10 11 12 developers agreed with the algorithm’s mapping, i.e., the
        </p>
        <p>Entity algorithm was able to find developer mistakes. In some
Figure 6: The change in the error of 12 entities from large of the 13 cases where the suggested module was not
acmodules in JabRef that have changed mapping but not cepted, the developers mentioned that (code) changes to
changed package. The first five (blue) entities have had the entity were needed for it to conform to the developer
no change in source code and the last seven (orange) have mapping.
changed source code. Bibi et al. compared the structural criteria part of the
algorithm proposed by Tzerpos and Holt with supervised
JabRef 3.7 Large Module Error Rates machine-learning approaches; Bayesian classification,
k-nearest-neighbor, and neural networks. Their study
.10 focuses on using dependencies as features (i.e.,
structural criteria) for incremental clustering. They evaluate
.08 the approaches using two versions of six open-source
software systems and find that dependencies between
.06 entities within the same module are important to avoid
misclassifications, especially when there are few
depen.4 dencies between entities in diferent modules.
0 We previously constructed a structure-based heuristic
for automatic mapping of source code to
Model-View.02 Controller-based architectures [15]. We evaluated the
approach on four products in a product line of games,
.00 all using the same game engine. We compared the
aurefactored new normal tomatic mapping to the manual mapping, and if they
Figure 7: The error rates of entities in large modules that are disagreed, then the type was flagged as containing an
undergoing refactoring, are new, or normal in JabRef 3.7. architectural problem. We compared the mappings of
653 entities and were able to correctly identify 76 out of
101 architectural problems as well as 18 false positives.
detect. The risk is that a module can be quite chaotic dur- The heuristic suggested a diferent mapping in 96 (14.7%)
ing a transition phase with multiple entities in diferent of 653 cases.
stages of the refactoring process. Furthermore, two of the projects were refactored to</p>
        <p>Another interesting observation is that new files tend be fully conformant. This refactoring removed 33 true
to have a lower error rate, indicating that the developers positives and six false positives. The true positives were
have understood the new architecture and that normal remedied by refactoring the source code. In the context
code changes could slowly make an entity harder to clas- of evaluating the performance of a method for automatic
sify. This could be due to some form of design erosion, mapping using the manual mappings as ground truth,
where changes are introduced that make the entity less these true positives would be regarded as erroneous
mapcohesive over time. pings when they, in fact, are pointing to source code with
architectural problems that need to be refactored.</p>
        <p>
          The CountAttract and MQAttract attraction functions
5. Related Work of HuGMe have been evaluated in four case studies [
          <xref ref-type="bibr" rid="ref2 ref6">10, 8</xref>
          ].
The focus is on evaluating the influence of two
configThere is previous work in the area of orphan adoption [9, uration parameters and comparing the performance of
10, 8, 7, 12, 15, 16, 17]. The focus is to evaluate and com- the attraction functions. Both attraction functions
assume a modular design based on the high cohesion low sented a fairly good correlation, and in one system, they
coupling style, and mapping would become problematic could find a repeating pattern of directories. Possibly
for modules designed specifically to not use this style. the ground truth architectures recovered in their study
Christl et al. suggest the incorporating a detection step is more low level than the modular architectures that
to better handle such modules, which would correspond we study. Still, it is likely that there is a variation on
to handling the style criteria. Furthermore, Chen et al. what dimension of an architecture that is expressed in
improves on CountAttract in an evolutionary case, i.e., a the package structure. This is further supported by
Buckpre-existing mapping is used. ley et al. where one system of five studied did not have
        </p>
        <p>Bittencourt et al. present two new attraction functions any clear correlation between packages and modules [19],
based on information retrieval techniques. They use the presenting clear dificulties and significant efort when
semantic information in the source code and calculate at- performing the manual mapping.
tractions based on cosine similarity (IRAttract) and latent
semantic indexing (LSIAttract). They make a quantitative
comparison between the performance of their attraction 6. Discussion and Validity
functions with CountAttract and MQAttract in an
evolutionary setting (where a few new files are to be assigned Our results clearly show that there is a set of entities in
a mapping). They find that a combination of attraction the systems that are systematically hard for the
statefunctions (e.g., if CountAttract fails, then try IRAttract) of-the-art automatic mapping techniques to map. One
performs best. This is explained by their qualitative anal- reason for this is the surprising result that all studied
ysis, where they find that CountAttract usually misplaces systems exhibit some very small modules. An automated
entities on module borders, MQAttract performs better technique would have very little data to use for these
when mapping entities with dependencies to many dif- modules, lowering the chance for successful mapping.
ferent modules, IRAttract and LSIAttract perform better In general, unbalanced data is problematic for machine
when mapping entities in libraries or entities on module learning techniques, and in particular, the distribution of
borders, but perform less well if there are modules that probabilities is important in Naive Bayes. 30% (237 out of
share vocabulary but are not related. 784) problematic entities are mapped to such small
mod</p>
        <p>Sinkala and Herold present InMap, which is not an ules in the ground truth mappings. This is a significant
automated approach to mapping per se; instead, InMap problem that needs to be solved.
suggests mappings to the end-user, who can choose to In essence, small modules need to be flagged (either
accept the suggested mapping (or not). It is an iterative automatically or manually) and handled separately. One
approach where a number of mappings are presented, idea in the context of Naive Bayes would be to manipulate
and the accepted mappings are used to improve the sug- the probability distribution appropriately to not wholly
gested mappings further. The suggested mappings are disregard small modules in the mapping. Schemes that
produced with the help of information retrieval informa- could be tested are a uniform distribution or diferent
tion similar to Bittencourt et al. with the addition of a ifxed settings (large, medium, small). These should be
descriptive text for each architectural module. The enti- reasonably easy for an end-user to assign to a module.
ties are treated as a database of documents, and InMap Still, there is a risk that overall performance will drop
uses Lucene to search this database using module infor- as potentially more entities will be hard to map. Another
mation as a query. As InMap is highly interactive, it will approach is to investigate why a few of the small modules
also use negative evidence to some degree, i.e., a rejected do not contain many problematic entities. We suspect
mapping suggestion will not be suggested again. The that these modules possibly exhibit a unique design, e.g.,
data from [16] suggest that using only the module names being very cohesive or having very clear naming, which
as a search criterion often results in high precision at the is perhaps not easy to address directly in a technique as
expense of the recall. This is most likely due to the fact it may simply be a way a module is designed.
that module names often reflect package names to some Using the naming strategy and possible ambiguity in
degree. Adding more and more module information in naming is an attractive approach to create an initial set
usthe query tends to lower precision, but increase the re- ing a specialized mapper. It should be possible to prompt
call, e.g., source code comments increase recall but lower an end-user with, e.g., keywords from the package or
precision in the mapping suggestions. class name asking for a mapping of the keyword. This</p>
        <p>
          Garcia et al. discuss the use of package and naming could significantly reduce the efort of creating an initial
information in software architecture recovery [
          <xref ref-type="bibr" rid="ref20">18</xref>
          ]. In set that could then be used as a basis for other
mapgeneral, they found that their ground truth components ping techniques. However, a complete approach must be
often spanned or shared several packages. They could prepared to handle subject systems where the naming
not find a correlation between components and single information does not reflect the modular architecture.
package or directory names. One of their four cases pre- The data on finding problematic entities among entities
that lie on the borders of modules is conflicting. On will become less semantically cohesive as the vocabulary
the one hand, we cannot see any correlation between becomes a mix of words from the previous architecture.
the external fan ratio and the error rate. On the other The error rate of entities could then be used as a metric
hand, we observe a higher median external fan ratio in to know if an entity is properly aligned to other entities
problematic entities. We observe very high error rates in in the module.
combination with very low external fan ratios and vice Comparing a human-made mapping to the mapping
versa. This indicates that the external fan ratio is not a made by an automatic technique seems to be a useful
useful metric, and a more refined metric could give better piece of information. The related work [
          <xref ref-type="bibr" rid="ref4">9, 15</xref>
          ] shows that
answers. There is possibly a diference between incoming this often points to cases where (further) refactoring or
and outgoing dependencies that could be a factor. In [
          <xref ref-type="bibr" rid="ref4">9</xref>
          ], discussion is needed and that the automatic technique is
these entities were specifically detected and only used not necessarily wrong per se. However, if no human
mapwhen suggesting a new module (orphan kidnapping). ping exists, is it important for an automated technique to
Such an approach could also be investigated. notify a human user of such issues and not automatically
        </p>
        <p>
          We studied two diferent mappings in two versions assign the entity a mapping.
of JabRef and found six cases where only the mapping Comparing mappings using several diferent techniques
had changed (no change of source code), of which five could be a way forward, similar to what is done in [7] but
mapped to large modules. We found eleven entities with a diferent intent. This also points to a problematic
where the mapping and source code had changed (though situation as we cannot fully trust the ground truth
mapthe entity had not changed package), of which seven were pings; a perfect mapping technique would thus be flawed.
in large modules. For these entities, there was a signifi- There is also a general lack of ground truth mappings
cant diference in error rate between the two mappings. made by human experts and even fewer mappings made
We are relatively confident that the diference in the six by diferent experts on the same system. Four of the
entities with no change is due to disagreement among systems (JabRef v3.5, JabRef v3.7, ProM, and TeamMates)
the developers; in the other eleven, it could also be due have mappings done by experts. The others (ArgoUML,
to the actual change of the entities’ source code. This Ant, Lucene, and Sweet Home 3D) have mappings created
would indicate that between 0.8% and 2.3% of entities are by researchers studying the systems’ documentation and
hard to map correctly, even for JabRef experts. implementations [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The architects or developers of
        </p>
        <p>It should also be noted that JabRef is only one case and these systems would likely not agree to all of these
mapthat it was undergoing architectural refactoring during pings even if it is likely that large parts of the mappings
this time in development. We are reasonably confident are correct.
that this afects the results. We can argue that there may Two limiting factors in this study are that all systems
be more confusion among the developers during refactor- are implemented in Java and that we have only studied
ing, which should increase the chance of disagreements. one set of parameters of the attraction function, i.e., the
There is also the possibility that the process of refactor- one from [2] giving the best mapping performance.
Aning has brought the architecture to everyone’s attention, other set of parameters would likely give diferent error
possibly lowering the chance of disagreements. The low rates; however, we think the main points of the paper
error rate of new entities suggests the latter as more would still hold.
likely.</p>
        <p>The two mappings and versions of JabRef allow us to
study entities under refactoring and new entities. We 7. Conclusions and Future Work
ifnd 61 entities under refactoring and 348 new entities. If
we remove entities from small modules (with confound- We investigate the flaws in the automatic mapping of
ing error rates), we find that entities under refactoring source code to modules in eight open-source software
are considerably harder to map correctly. This is likely systems. We show that the state of the art technique has
because architectural refactoring is a process that can systematic flaws in its suggested mappings that need to
take some time to complete. The functional aspects of the be addressed. We find that a major contributing factor is
entities are likely fixed first, possibly with the removal of that all investigated systems have modules with very few
unwanted dependencies (especially as JabRef has some ground truth mappings. We also find that all systems use
tests for this). a naming strategy, but this strategy is often ambiguous.</p>
        <p>There is, however, a risk that the semantic information We found no clear evidence that entities that have many
(e.g., variable names) will not be changed and correctly dependencies to or from entities in other modules are
reflect the vocabulary of the module. It would be interest- systematically problematic. Our data indicate that such
ing to see if this happens to these entities in future ver- dependencies can be a factor, but the metrics used are
sions of JabRef or if the current state is considered good likely not well suited to clearly show such problems.
enough. If so, there is a considerable risk that modules We studied diferences in expert mappings in one of</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Acknowledgments</title>
      <p>The research was supported by the Centre for Data
Intensive Sciences and Applications at Linnaeus University.
the systems, where we had two diferent versions and two
diferent ground truths. We found that disagreements
exist and that such entities are likely to have a high error
rate in the mappings, although there are not many such
entities. We also studied refactored files and new entities.
Refactored entities tend to have a significantly higher
error rate compared to both new entities and normal
entities. There is a risk that refactoring is considered
done when the entity is moved and the functional aspects
are fixed. Automatic mapping could indicate when the
entity is properly aligned to other entities in the module
or noticeably diferent.</p>
      <p>Our priority for the future is to address the small
modules. We will try diferent approaches to manipulating
the probability distribution of the modules and find the
efect on overall mapping performance. Another area of
interest is the use of naming information to create an
initial set, as this could significantly reduce the mapping
efort.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>verse Engineering</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>163</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Christl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Koschke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Storey</surname>
          </string-name>
          , Automated
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>mation and Software Technology</source>
          <volume>49</volume>
          (
          <year>2007</year>
          )
          <fpage>255</fpage>
          -
          <lpage>274</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>V.</given-names>
            <surname>Tzerpos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Holt</surname>
          </string-name>
          ,
          <article-title>The orphan adoption prob-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>ing Conference on Reverse Engineering</source>
          ,
          <year>1997</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Christl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Koschke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Storey</surname>
          </string-name>
          , Equipping the
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <year>2005</year>
          , pp.
          <fpage>98</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ericsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wingkvist</surname>
          </string-name>
          , s4rdm3x: A
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>techniques</surname>
          </string-name>
          ,
          <source>Journal of Open Source Software 6</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          (
          <year>2021</year>
          )
          <article-title>2791</article-title>
          .
          <source>doi:1 0 . 2 1 1 0 5 / j o s s . 0 2</source>
          <volume>7 9 1 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Maqbool</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kanwal</surname>
          </string-name>
          , Supervised learn-
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>Science</source>
          <volume>29</volume>
          (
          <year>2016</year>
          )
          <fpage>287</fpage>
          -
          <lpage>313</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brunet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bittencourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Serey</surname>
          </string-name>
          , J. Figueiredo,
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Engineering</surname>
          </string-name>
          ,
          <year>2012</year>
          , pp.
          <fpage>257</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lenhard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Herold</surname>
          </string-name>
          , Exploring the suit-
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>tectural inconsistencies</article-title>
          ,
          <source>Software Quality Journal</source>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ericsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wingkvist</surname>
          </string-name>
          , Towards im- (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>proved initial mapping in semi automatic clustering</article-title>
          , [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Toll</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wingkvist</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ericsson</surname>
          </string-name>
          , Evalu-
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>in: Proceedings of the 12th European Conference ation of a static architectural conformance checking</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>on Software Architecture: Companion Proceedings</source>
          ,
          <article-title>method in a line of computer games</article-title>
          , in: 10th in-
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>ECSA '18</source>
          ,
          <year>2018</year>
          , pp.
          <volume>51</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>51</lpage>
          :7. ternational ACM Sigsoft conference on Quality of [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ericsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wingkvist</surname>
          </string-name>
          , Semi- software architectures,
          <source>ACM</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>118</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>automatic mapping of source code using naive</article-title>
          [16]
          <string-name>
            <given-names>Z. T.</given-names>
            <surname>Sinkala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Herold</surname>
          </string-name>
          , Inmap: Automated inter-
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>bayes, in: 13th European Conference on Software active code-to-architecture mapping recommenda-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Architecture</surname>
          </string-name>
          - Volume
          <volume>2</volume>
          ,
          <year>2019</year>
          , p.
          <fpage>209</fpage>
          -
          <lpage>216</lpage>
          . tions, in: IEEE 18th International Conference on [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>De Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Balasubramaniam</surname>
          </string-name>
          , Controlling soft- Software
          <source>Architecture (ICSA)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>183</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <article-title>ware architecture erosion: A survey</article-title>
          ,
          <source>Journal</source>
          <volume>of</volume>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lian</surname>
          </string-name>
          , An improved mapping
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>Systems and Software</source>
          <volume>85</volume>
          (
          <year>2012</year>
          )
          <fpage>132</fpage>
          -
          <lpage>151</lpage>
          . method for
          <source>automated consistency check between</source>
          [4]
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Notkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          ,
          <article-title>Software software architecture and source code</article-title>
          , in: IEEE
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>reflexion models: Bridging the gap between source 20th International Conference on Software Quality,</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>and high-level models</article-title>
          ,
          <source>ACM SIGSOFT Software Reliability and Security (QRS)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>60</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>Engineering Notes</source>
          <volume>20</volume>
          (
          <year>1995</year>
          )
          <fpage>18</fpage>
          -
          <lpage>28</lpage>
          . [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Garcia</surname>
          </string-name>
          , I. Krka,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mattmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Medvidovic</surname>
          </string-name>
          , Ob[5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. O</given-names>
            <surname>'Crowley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Herold</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Buck- taining ground-truth software architectures</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <article-title>ley, Architecture consistency: State of the practice</article-title>
          ,
          <source>35th International Conference on Software Engi-</source>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>challenges and requirements</article-title>
          ,
          <source>Empirical Software neering (ICSE)</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>901</fpage>
          -
          <lpage>910</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>Engineering</source>
          <volume>23</volume>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          . [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Buckley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>English</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rosik</surname>
          </string-name>
          , S. Herold, [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Knodel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Popescu</surname>
          </string-name>
          ,
          <article-title>A comparison of static archi- Real-time reflexion modelling in architecture rec-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <source>IEEE/IFIP Working Conference on Software Archi- Software Technology</source>
          <volume>61</volume>
          (
          <year>2015</year>
          )
          <fpage>107</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>tecture</surname>
          </string-name>
          ,
          <year>2007</year>
          , pp.
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          . [7]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bittencourt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Jansen de Souza Santos</surname>
          </string-name>
          , D. D. S.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>