<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Application Domains in the Research Papers at BENEVOL: A Retrospective</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Capiluppi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nemitari Ajienka</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bilyaminu Auwal Romo</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept of Computer Science, Brunel University London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept of Computer Science, Edge Hill University</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept of Engineering and Digital Technologies, Coventry University</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Research on empirical software engineering has increasingly used the data that is made available in online repositories, speci cally Free/Libre/Open Source Software projects (FLOSS). The latest trends for researchers is to gather \as much data as possible" to (i) prevent bias in the representation of a small sample, (ii) work with a sample as close as the population itself, and (iii) showcase the performance of existing or new tools in treating vast amount of data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The e ects of harvesting enormous amounts
of data have been only marginally considered
so far: data could be corrupted; repositories
could be forked; and developer identities could
be duplicated. In this paper we posit that
there is a fundamental aw in harvesting large
amounts of data, and when generalising the
conclusions: the application domain, or
context, of the analysed systems must be the
primary factor for the cluster sampling of FLOSS
projects.</p>
    </sec>
    <sec id="sec-2">
      <title>This paper presents two contributions: rst, we analyse a collection of 100 BENEVOL papers that appeared showing whether (and how</title>
      <p>much) FLOSS data has been harvested, and
how many times the authors agged an issue
in their di erent application domains. Second,
we discuss the implications of using
`application domain' as the clustering factor in the
sampling of FLOSS data, and the
generalisations within and outside the clusters.</p>
    </sec>
    <sec id="sec-3">
      <title>FLOSS, application domains,</title>
      <p>Index terms|
BENEVOL papers
1</p>
      <sec id="sec-3-1">
        <title>Introduction</title>
        <p>The use of open, available data has been a welcomed
accelerator in the software engineering research eld.
Data on the processes and products available via an
Open Source approach has led to an increasingly large
number of workshops, conferences, papers and
research attempts to describe the phenomenon.
Researchers gathered initially in 2001 around the Open
Source Software (OSS) workshops, held annually in
colocation with the ICSE series of conferences. Before
the OSS workshop spawned into the OSS conference
in 2005, the BENEVOL community started to group
together researchers from the Software Evolution
domain. Its initial focus was `(...) to bring researchers
to identify and discuss important principles, problems,
techniques and results related to software evolution
research and practice'1.</p>
        <p>
          While the goal of a few BENEVOL papers has been
to achieve the generality of the results [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the domain,
context and uniqueness of a software system have not
been considered very often by empirical software
engineering research. As in the example reported in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], the
extensive study of all JSON parsers available would
1https://smartcare.be/events/benevol-04-workshop
nd similarities between them or common patterns.
That type of study would focus on one particular
language (JSON), one speci c domain (parsers) and
inevitably draw limited conclusions. On the other hand,
considering the \parsers" domain (but without
focusing on one single language) would show the common
characteristics of developing that type of systems, and
irrespective of their language.
        </p>
        <p>
          The underlying vision of this paper is to open a
proper debate on the importance of context for any
software system, and the uniqueness of its
application domain. This position paper stems from the
work of several prominent researchers who called the
community to `go deeper, not wider' (Michael
Godfrey at MSR 2017) and `minding the mine, mining the
mind' [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. We posit that past empirical investigations
using FLOSS systems have been mostly blind to these
aspects (i.e., context and domain), establishing
similarities between vastly di erent systems if they shared
a common pattern in one measured attribute. Using
an extreme example, one could establish a similarity
between the coupling of a `hammock` and a `bridge'
due to the fact that both are held at the sides.
        </p>
        <p>
          The purpose of this paper is to share some
ndings about a selection of papers discussed during the
last few years of BENEVOL workshops. The focus is
speci cally based on BENEVOL papers that have used
FLOSS data. The context of our analysis is the
diversity of FLOSS projects under study, and how that
was re ected by researchers in their ndings. Some
100 papers are analysed in terms of whether FLOSS
projects are used, how many, and whether
considerations of application domains have been used to inform
the sampling of FLOSS projects, or the validity of the
conclusions. We assume that domains are relevant as
a fundamental construct for any empirical software
engineering research [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Related Work</title>
        <p>
          The vast literature on FLOSS systems of the last 10
years has been possible also due to a series of guidelines
on how to perform quantitative, empirical analysis on
FLOSS processes and products. When SourceForge2
was considered as the de-facto FLOSS forge, a well
received research paper shared more than one insight
on the most common mistakes to avoid when mining
data and results from the projects hosted there [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
Among other more technical issues of mining this
speci c forge, this paper actively warned against an
inaccurate `screening' of projects into samples: reducing a
population to, say, `FLOSS projects with more than 7
developers' would inevitably reduce the variables for
the analysis, but the `number of developers' variable
cannot be used as a dependent or independent variable
for any model or analysis.
        </p>
        <p>
          The acknowledgement of GitHub as the newly
established central focus for FLOSS development
generated a similar requirement, in terms of shared
guidelines to avoid common mining mistakes [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Di erently
from [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the 2014 paper mostly focused on the
technical aspects of GitHub, and how the collected metrics
could skew the results, due to the inner workings of the
Git toolset, and the di erent approach to FLOSS
development observed on GitHub (forking, non-software
development, inactivity of projects). Neither [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] or [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
warned about the variability of FLOSS projects, the
importance of their context, or the uniqueness of their
domains.
        </p>
        <p>
          Outside of the FLOSS literature, the diversity and
context of software systems have received some
attention in the past [
          <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
          ]. The phrase \large scale" has
been frequently used in empirical software engineering
research to denote the magnitude of the analyzed case
study or studied software sample. Notwithstanding,
Nagappan et al. argue that analyzing a high number
of projects is not always necessary [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. But what is
even more important is the selection of the projects
studied.
        </p>
        <p>Interesting patterns valuable to researchers and
practitioners are often identi ed in domain-based
analysis of software projects. Results from one domain
might not be applicable in another. As such, it is
important for results to be representative.</p>
        <p>
          Software categorization or domain clustering has
gained importance over the years. For example, the
knowledge of software trends in a particular domain
can assist developers in the search for domain-speci c
reusable components [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Tian et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] proposed a
technique based on Latent Dirichlet Allocation for
automatic software categorization in open-source
software repositories.
        </p>
        <p>
          According to Hae iger et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], \domain
analyses, documentation, and quality standards enhance
the ability to reuse software components". However,
our survey of past BENEVOL papers that have
analyzed OSS projects demonstrates that software
domains have not been considered in most of the past
software engineering studies.
3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>A Survey 2012 to 2018 of</title>
      </sec>
      <sec id="sec-3-4">
        <title>BENEVOL</title>
      </sec>
      <sec id="sec-3-5">
        <title>Papers:</title>
        <p>In order to show how FLOSS data has been used and
analysed by the BENEVOL community, we report here
an investigation of the research papers appeared in the
last 5 years of the BENEVOL event. An overall 101
papers have been considered in this study: we share
the raw data in the spreadsheet at https://tinyurl.
com/y69wkadr.</p>
        <p>Each paper was read by one of the co-authors, and
summarised along the following points:
• Use of FLOSS systems (yes/no): at rst we
checked whether FLOSS projects are used in the
paper at all. This served as an indicator of the
pervasiveness of FLOSS projects in the literature
produced by BENEVOL papers.
• Number of FLOSS systems used : in second
instance, we trawled through the paper, annotating
where the authors mentioned how many FLOSS
systems were used. In the case of full papers, the
abstract, introduction, methodology and
conclusion were read for that purpose.
• Analysis of application domains: thirdly, we
considered the methodology, results and conclusion
of each paper, along with the threats to
validity, looking for considerations of application
domains. We checked if the authors considered
this attribute in the sampling of FLOSS projects,
whether they limited their results against this
axis, or whether it was considered a speci c threat
to validity. This attribute was coded as either
fyes | nog.</p>
        <p>The contributions to the 2016 edition of BENEVOL
are not available online, so they had to be excluded
from our analysis. The spreadsheet with the
categorisation of the papers has been made available for
inspection under the following link: https://tinyurl.
com/y69wkadr.
3.1</p>
        <sec id="sec-3-5-1">
          <title>BENEVOL use of FLOSS Systems</title>
          <p>In this section we provide the rst point of our
analysis: `how many BENEVOL papers have used FLOSS
systems in their analyses? '. As visible in the two plots
of Figure 1, researchers (and accepted BENEVOL
papers) have steadily used FLOSS systems for their
papers. The rst plot shows the absolute numbers of
accepted BENEVOL submissions that use one or more
FLOSS projects.</p>
          <p>The bottom plot of Figure 1 shows the ratio of
FLOSS and non-FLOSS papers in the BENEVOL
sample of papers. It is getting increasingly more
common to use one or more commercial software systems,
or a combination of FLOSS and non-FLOSS projects.
3.2</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>Number of FLOSS systems used in</title>
        </sec>
        <sec id="sec-3-5-3">
          <title>BENEVOL</title>
          <p>In this section we report on the number of FLOSS
systems evaluated by BENEVOL papers. For this
purpose, we analysed the methodology description, or the
empirical approach, of each paper to determine how
many FLOSS systems were reported in the study.
Figure 2 displays the cumulative number of FLOSS
systems used in BENEVOL papers, per year. The median
number of systems has increased from one analysed
OSS system in 2012 to 1,127 systems in 2018.</p>
          <p>The exponential number of FLOSS systems being
used by BENEVOL papers has been accelerated by
many factors: (i) availability of open forges
(FreshMeat, SourceForge, Savannah, Apache FSF, GitHub
and many others); (ii) common, shared toolsets to
perform the analyses; (iii) guidelines on how to e ectively
use forges.</p>
          <p>Below we give a summary of ndings to assess the
trends observed in the number of FLOSS systems
analysed by the BENEVOL papers.</p>
        </sec>
        <sec id="sec-3-5-4">
          <title>Growth of sample sizes</title>
          <p>The trend that we observed throughout the subsequent
years of the BENEVOL contributions is,
fundamentally, summarised as `the more the better'. Authors
have started to include larger and larger FLOSS
samples to their papers. We can assume that this pattern
has been followed in order to achieve the generality
of a paper's ndings. At the last edition of available
BENEVOL contributions (BENEVOL 2018), over one
million FLOSS systems were considered for
investigation, jointly by the accepted papers.</p>
        </sec>
        <sec id="sec-3-5-5">
          <title>Uncertainty on sample sizes</title>
          <p>
            Several BENEVOL papers use ecosystems [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ], or
umbrella projects [12], as their cases studies, whereas
other papers either take a subset of those
superprojects, or explicitly declaring the number of
subprojects (e.g., Scala [
            <xref ref-type="bibr" rid="ref12">13</xref>
            ] or Python [14] projects) that they
analysed. This means that our nal gures are mostly
lower bounds of the actual number of FLOSS systems
being used by the BENEVOL community.
          </p>
        </sec>
        <sec id="sec-3-5-6">
          <title>Ecosystems vs time of analysis</title>
          <p>Several BENEVOL papers have used umbrella projects
(for example, Gnome). In most cases we considered
them as single FLOSS systems: depending on the time
of the analysis, these larger projects can contain a
variable number of sub-projects. This makes it di cult to
de ne the status of the super-project, in terms of
number of its sub-projects, as well as their domains. This
also makes it di cult to replicate those studies, as well
as their results and conclusions.</p>
        </sec>
        <sec id="sec-3-5-7">
          <title>Sampling and Pruning</title>
          <p>Throughout the editions of BENEVOL, sampling of a
few systems (see for instance [15] or [16]) has given
way to whole-forge analyses. Also, there seems to be
a general view that `pruning' a sample is a good idea
for removing outliers, or for promoting quality. This
has an e ect on the sample studied, and the
representativeness of the population as a whole.
3.3</p>
        </sec>
        <sec id="sec-3-5-8">
          <title>Application domains and FLOSS projects</title>
          <p>The third analysis was based on the application
domains of the systems considered in the empirical study.
For all the papers (not only for those using FLOSS
projects), we tried to establish whether the authors
considered the results, ndings or discussion as
constrained by the type of system (e.g., its domain). This
included checking how the threats to external validity
(if any) addressed limited the conclusion to the
domain(s) under investigation.</p>
          <p>We grouped the papers into two categories (and
plotted them accordingly per year):
1. papers that directly considered application
domains as drivers in the variability of the results
(stack "YES" in Figure 3);
2. papers that didn't considered application domains
as drivers (stack "NO" in Figure 3).</p>
          <p>The results of this analysis are shown in Figure 3:
a ratio (%) is used to separate the papers in the
two categories. It is clear from the visualisation that
BENEVOL papers do not generally acknowledge the
variability of results as driven by the domains of the
systems involved. Earlier papers (especially from the
2012 batch) have a good cover of domains in the
evaluation of the results, but this is not re ected in the
later editions of BENEVOL.</p>
          <p>As visible in the Figure, the majority of ndings
on FLOSS, as reported by BENEVOL papers, do
not mention application domains. In some cases,
researchers have acknowledged the variability of the
results [17, 18], and hinted that other factors could play
a role in such variability. We considered as a
\limited" acknowledgment of the relevance of the
application domain when authors mentioned the diversity of
the systems under study.
4</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>BENEVOL: FLOSS and Domains</title>
        <p>The birdseye view on the type of BENEVOL
contributions (Sections 3.1, 3.2 and 3.3 above) reveals some
interesting trends when dealing with FLOSS projects.
Below we discuss in more detail whether FLOSS
papers were analysed ("YES" or "NO"), and whether
domains were considered in the analysis ("YES" or
"NO").
So far in the BENEVOL series, few papers explicitly
addressed the importance of domains when analysing
systems, or when discussing ndings. An interesting
perspective is given in [19], since it considers a very
speci c type of systems, the `cross-system packages'.
These systems are likely to show similar characteristics
since they are supposed to act as vectors to an from
the overarching system.</p>
        <p>By drawing on the importance of the application
domains in this paper [20], the authors signify the
importance of domain analysis when creating a
theoretical and practical framework that supports the
development and the evolution of adaptive data-intensive
software systems for ubiquitous environments in their
study. Thus, they focus on data and in particular on
the problem of nding the most suitable portion of
data that have to be provided by the application in
the of context of `self-adaptive system'.</p>
        <p>Likewise in the 11th edition of BENEVOL (2012),
[21] examined the impact and role of social media on
software development. The authors argued that
\social media is poised to bring about a paradigm shift
in software engineering research" particularly in OSS
community.</p>
        <p>
          In the 2014 edition, only one BENEVOL study
focusing on OSS projects implicitly highlighted the
need to investigate projects from various domains [
          <xref ref-type="bibr" rid="ref13">22</xref>
          ].
The authors studied an OSS project called DrJava
and implicitly mentioned domains but did not
investigate multiple projects clustered into several domains.
According to the authors, \we chose an IDE since
they contain elements of multiple domains. The IDE
project was taken from the Qualitas Corpus and it
consists of 3000 revisions since 2000 and the system grew
from 30K SLOC in 2003 to 200K SLOC in 2013.
        </p>
        <p>We concluded that application domains are not well
represented or studied in the papers that use FLOSS
data.
4.2</p>
        <sec id="sec-3-6-1">
          <title>FLOSS: YES, Domains: NO</title>
          <p>
            The vast majority of BENEVOL contributions, based
on FLOSS systems, do not consider domains as one
of the factors to take in consideration. An interesting
example of this approach is given in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], where the
authors pose that `... (to) gather as much as possible
should be the aim of empirical software engineering '.
          </p>
          <p>
            More in general, the approach of researchers is to
focus on speci c languages or source code models (see
for instance the paper in [
            <xref ref-type="bibr" rid="ref14">23</xref>
            ], focused on all available
meta-models from GitHub), hence representing
convenience sampling. For example in the study on control
ow, Landman et al., [
            <xref ref-type="bibr" rid="ref15">24</xref>
            ] focused on the Sourcerer
Corpus which contains 18K (13K non empty) Java
projects. In an empirical analysis of the
maintainability of CRAN packages, Claes et al., [
            <xref ref-type="bibr" rid="ref16">25</xref>
            ] presented early
results on analysing the dependencies of the CRAN R
packages repository.
          </p>
          <p>We concluded that most of the papers studied from
the BENEVOL series do not consider the application
domains as an important factor for software analysis
or evolution.
4.3</p>
        </sec>
        <sec id="sec-3-6-2">
          <title>FLOSS: NO, Domains: YES</title>
          <p>A few of the papers that we analysed are not based on
FLOSS systems, but more in general on commercial,
or in-house software. In a few cases, we observed that
the authors actually considered the limitations of their
case studies to the one domain that was investigated.</p>
          <p>
            As a few of such examples, we noted a paper
based on a banking system [
            <xref ref-type="bibr" rid="ref17">26</xref>
            ]; and one focused on
the speci c features of home-automation system [
            <xref ref-type="bibr" rid="ref18">27</xref>
            ].
Both these papers clearly acknowledged the limitations
given by the chosen application domains that their
systems are based on. In other cases, the authors
specifically focused on one domain (for example, GIS
systems [
            <xref ref-type="bibr" rid="ref19">28</xref>
            ], or the larger business domain [
            <xref ref-type="bibr" rid="ref20">29</xref>
            ]).
          </p>
          <p>In general, the BENEVOL papers using non-OSS
software as their case studies do not use the domains
to aggregate results. Nonetheless a few BENEVOL
contributions have shown a clear pathway into not
generalising the ndings to all domains.
5</p>
        </sec>
      </sec>
      <sec id="sec-3-7">
        <title>Conclusion</title>
        <p>This paper analysed how open source software has
been used by the BENEVOL contributions between
2012 and 2018. We showed the increasing number of
BENEVOL contributions that used FOSS projects for
their analyses.</p>
        <p>Although the majority of contributions do not
acnowledge the importance of domains when discussing
the ndings, there is an increasing number of papers
that limit the results, or the data sampling, to speci c
domains. We believe that one of the major challenges
for empirical software engineering is to better
understand the role of domains, especially in the evolution
of software systems. We propose for papers that
empirically analyse software systems to acknowledge such
challenge in a `threat to domain validity'.
[12] Maelick Claes. Applying biological evolution to
software ecosystems a case study with gnome.
[14] Jose Javier Merchante and Gregorio Robles. From
python to pythonic: Searching for python idioms
in github.
[15] Ward Muylaert and Coen De Roover. Untangling
source code changes using program slicing. In
BENEVOL, pages 36{38, 2017.
[16] Jie Tan, Mircea Lungu, and Paris Avgeriou.
Towards studying the evolution of technical debt
in the python projects from the apache software
ecosystem. In BENEVOL, pages 43{45, 2018.
[17] Zeeger Lubsen, Andy Zaidman, and Martin
Pinzger. Using association rules to study the
coevolution of production &amp; test code. In Mining
Software Repositories, 2009. MSR'09. 6th IEEE
International Working Conference on, pages 151{
154. IEEE, 2009.
[18] Christian Rodr guez-Bustos and Jairo Aponte.</p>
        <p>How distributed version control systems impact
open source software projects. In Mining
Software Repositories (MSR), 2012 9th IEEE
Working Conference on, pages 36{39. IEEE, 2012.
[19] Eleni Constantinou, Alexandre Decan, and Tom
Mens. Breaking the borders: an investigation of
cross-ecosystem software packages. arXiv preprint
arXiv:1812.04868, 2018.
[20] Marco Mori and Anthony Cleve. A framework
to support the development and evolution of
selfadaptive data-intensive systems. In 11th edition
of the BElgian-NEtherlands software eVOLution
symposium (BENEVOL 2012), 01 2012.
[21] Maelick Claes. Applying biological evolution to
software ecosystems a case study with gnome.
In 11th edition of the BElgian-NEtherlands
software eVOLution symposium (BENEVOL 2012),
01 2012.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Pietri</surname>
          </string-name>
          and
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Zacchiroli</surname>
          </string-name>
          .
          <article-title>Towards universal software evolution analysis</article-title>
          .
          <source>In BENEVOL</source>
          , pages
          <volume>6</volume>
          {
          <fpage>10</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Meiyappan</given-names>
            <surname>Nagappan</surname>
          </string-name>
          , Thomas Zimmermann, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bird</surname>
          </string-name>
          .
          <article-title>Diversity in software engineering research</article-title>
          .
          <source>In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering</source>
          , pages
          <volume>466</volume>
          {
          <fpage>476</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Ko</surname>
          </string-name>
          .
          <article-title>Mining the mind, minding the mine: grand challenges in comprehension and mining</article-title>
          . In Andy Zaidman, Yasutaka Kamei, and Emily Hill, editors,
          <source>Proceedings of the 15th International Conference on Mining Software Repositories, MSR</source>
          <year>2018</year>
          , Gothenburg, Sweden, May
          <volume>28</volume>
          - 29,
          <year>2018</year>
          , page 118. ACM,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Steve</given-names>
            <surname>Easterbrook</surname>
          </string-name>
          , Janice Singer,
          <string-name>
            <surname>Margaret-Anne Storey</surname>
            , and
            <given-names>Daniela</given-names>
          </string-name>
          <string-name>
            <surname>Damian</surname>
          </string-name>
          .
          <article-title>Selecting empirical methods for software engineering research</article-title>
          . In Guide to advanced
          <source>empirical software engineering</source>
          , pages
          <volume>285</volume>
          {
          <fpage>311</fpage>
          . Springer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>James</given-names>
            <surname>Howison</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Crowston</surname>
          </string-name>
          .
          <article-title>The perils and pitfalls of mining sourceforge</article-title>
          .
          <source>In Proceedings of the International Workshop on Mining Software Repositories (MSR 2004. Citeseer</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Eirini</given-names>
            <surname>Kalliamvakou</surname>
          </string-name>
          , Georgios Gousios, Kelly Blincoe, Leif Singer, Daniel M German,
          <string-name>
            <given-names>and Daniela</given-names>
            <surname>Damian</surname>
          </string-name>
          .
          <article-title>The promises and perils of mining github</article-title>
          .
          <source>In Proceedings of the 11th working conference on mining software repositories</source>
          , pages
          <volume>92</volume>
          {
          <fpage>101</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Carmine</given-names>
            <surname>Vassallo</surname>
          </string-name>
          , Sebastiano Panichella, Fabio Palomba,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Proksch</surname>
          </string-name>
          , Andy Zaidman, and Harald C Gall.
          <article-title>Context is king: The developer perspective on the usage of static analysis tools</article-title>
          .
          <source>In 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER)</source>
          , pages
          <fpage>38</fpage>
          {
          <fpage>49</fpage>
          . IEEE,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Yunwen</given-names>
            <surname>Ye</surname>
          </string-name>
          and
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Fischer</surname>
          </string-name>
          .
          <article-title>Reuseconducive development environments</article-title>
          .
          <source>Automated Software Engineering</source>
          ,
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <volume>199</volume>
          {
          <fpage>235</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Kai</given-names>
            <surname>Tian</surname>
          </string-name>
          , Meghan Revelle, and
          <string-name>
            <given-names>Denys</given-names>
            <surname>Poshyvanyk</surname>
          </string-name>
          .
          <article-title>Using latent dirichlet allocation for automatic categorization of software</article-title>
          .
          <source>In 6th IEEE International Working Conference on Mining Software Repositories</source>
          ,
          <year>2009</year>
          . MSR'
          <volume>09</volume>
          ., pages
          <volume>163</volume>
          {
          <fpage>166</fpage>
          . IEEE,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Stefan</surname>
            <given-names>Hae iger</given-names>
          </string-name>
          , Georg Von Krogh, and
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Spaeth</surname>
          </string-name>
          .
          <article-title>Code reuse in open source software</article-title>
          .
          <source>Management Science</source>
          ,
          <volume>54</volume>
          (
          <issue>1</issue>
          ):
          <volume>180</volume>
          {
          <fpage>193</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Tom</given-names>
            <surname>Mens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Bram</given-names>
            <surname>Adams</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Josianne</given-names>
            <surname>Marsan</surname>
          </string-name>
          .
          <article-title>Towards an interdisciplinary, socio-technical analysis of software ecosystem health</article-title>
          .
          <source>arXiv preprint arXiv:1711.04532</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Yunior</surname>
            <given-names>Pacheco</given-names>
          </string-name>
          , Jonas De Bleser, Tim Molderez, Dario Di Nucci, Wolfgang De Meuter, and Coen De Roover.
          <article-title>Mining extension point patterns in scala</article-title>
          .
          <source>In BENEVOL</source>
          , pages
          <volume>16</volume>
          {
          <fpage>20</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Davy</surname>
            <given-names>Landman</given-names>
          </string-name>
          , Alexander Serebrenik, and
          <string-name>
            <given-names>Jurgen</given-names>
            <surname>Vinju</surname>
          </string-name>
          .
          <article-title>The relationship between cc and sloc: a preliminary analysis on its evolution. In Benevol 2014 (Seminar on Software Evolution in Belgium and the Netherlands</article-title>
          , Amsterdam, The Netherlands,
          <source>November 27-28</source>
          ,
          <year>2014</year>
          ), pages
          <fpage>29</fpage>
          {
          <fpage>30</fpage>
          .
          <string-name>
            <surname>Centrum</surname>
          </string-name>
          voor Wiskunde en Informatica,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>O</given-names>
            <surname> nder</surname>
          </string-name>
          <string-name>
            <surname>Babur</surname>
          </string-name>
          , Loek Cleophas, and Mark van den Brand.
          <article-title>Metamodel clone detection with samos</article-title>
          .
          <source>BENEVOL</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Davy</surname>
            <given-names>Landman</given-names>
          </string-name>
          , Alexander Serebrenik, and
          <string-name>
            <given-names>Jurgen</given-names>
            <surname>Vinju</surname>
          </string-name>
          .
          <article-title>Control ow in the wild a rst look at 13k java projects</article-title>
          .
          <source>BENEVOL</source>
          <year>2013</year>
          , page
          <volume>35</volume>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [25] Maelick Claes, Tom Mens, and
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Grosjean</surname>
          </string-name>
          .
          <article-title>Towards an empirical analysis of the maintainability of cran packages</article-title>
          .
          <source>BENEVOL</source>
          <year>2013</year>
          , page 42.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Elvan</surname>
            <given-names>Kula</given-names>
          </string-name>
          , Ayushi Rastogi, Hennie Huijgens, and Arie van Deursen.
          <article-title>Characterizing rapid releases in a large banking company: A case study</article-title>
          .
          <source>In BENEVOL</source>
          , pages
          <volume>56</volume>
          {
          <fpage>60</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Tim</surname>
            <given-names>Molderez</given-names>
          </string-name>
          , Coen De Roover, and Wolfgang De Meuter.
          <article-title>Towards a domain-speci c language for automated network management</article-title>
          .
          <source>In BENEVOL</source>
          , pages
          <volume>39</volume>
          {
          <fpage>43</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Cosmin</surname>
            <given-names>Tomozei</given-names>
          </string-name>
          , Iulian Furdu, and
          <string-name>
            <surname>Simona-Elena</surname>
            <given-names>V</given-names>
          </string-name>
          ^
          <article-title>arlan. Gis sdks dynamics echoed by social requirements transformations</article-title>
          .
          <source>In BENEVOL</source>
          , pages
          <volume>22</volume>
          {
          <fpage>25</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Gururaj</given-names>
            <surname>Maddodi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Slinger</given-names>
            <surname>Jansen</surname>
          </string-name>
          .
          <article-title>Responsive software architecture patterns for workload variations: A case-study in a cqrs-based enterprise application</article-title>
          .
          <source>In BENEVOL, page 30</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>