<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Re ning Software Quality Prediction with LOD</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Ceolin</string-name>
          <email>d.ceolin@vu.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Till Dohmen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joost Visser</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Software Improvement Group Rembrandt Toren</institution>
          ,
          <addr-line>15th oor, Amstelplein 1 1096 HA Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>VU University Amsterdam de Boelelaan 1081 1081HV Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The complexity of software systems is growing and the computation of several software quality metrics is challenging. Therefore, being able to use the already estimated quality metrics to predict their evolution is a crucial task. In this paper, we outline our idea to use Linked Open Data to enrich the information available for such prediction. We report our experience so far, and we outline the preliminary results obtained.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Software size and complexity is growing, thus being able to estimate and predict
software quality is crucial to monitor the process of software development and
promptly steer it. In fact, a quality metric provides a value summarizing one
relevant aspect of the software that can be consulted to identify issues or risks
in the development process or in the software itself. Therefore, several di erent
quality dimensions have been de ned, as described, for instance, by Kan [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Estimating software quality is then a crucial but challenging task, for several
reasons including the complexity of the software to be measured and the fact
that these measures are often hard to quantify: some of them depend on
runtime software behavior, some on static software properties. The estimation of
the values of these measures is possible, as demonstrated, for instance, by Alves
and Visser [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Bouwers [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, given the complexity of this task, we
propose to use such estimates to predict the temporal evolution of these values.
      </p>
      <p>Preliminary analyses on a dataset from the Software Improvement Group3
show encouraging results on the use of these estimates as starting point for the
prediction of the evolution over time of software quality ratings.4 We
hypothesize that, by using Linked Open Data (LOD) we can improve and re ne the
accuracy of our predictions. In particular, by enriching the information available
about the projects analyzed, we can categorize these projects (e.g., by
industry sector or programming language), thus increasing the possibility to group
3 http://www.sig.eu
4 For con dentiality reasons, we could not make the dataset publicly available.
together projects showing similar quality evolution over time. We present here
some preliminary encouraging results obtained in this direction, and we discuss
a series of open issues that we need to address in order to extend this research.</p>
      <p>The rest of this paper is structured as follows: Section 2 introduces related
work. Section 3 describes the enrichment of software projects data. Section 4
provides preliminary results, that are discussed in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related</title>
    </sec>
    <sec id="sec-3">
      <title>Work</title>
      <p>
        Software quality prediction is an important issue, that has been tackled from
di erent points of view. As Al-Jamini and Ahmed [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] describe in their review,
several relevant approaches to this problem make use of machine learning.
      </p>
      <p>
        We have also employed machine learning techniques (in particular, Markov
chains [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) to predict software quality based on the starting rating of a project [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The results are promising and we will aim at perfecting them with additional
features, properly selected from external sources, like LOD. The future quality
value of systems shows a strong correlation with the current quality rating, due
to the fact that the rating usually changes very slowly over time. Moreover, a
second trend was discovered which revealed that higher quality systems tend to
deteriorate in quality and low-quality systems tend to improve, both with the
tendency towards the medium quality level. This could be explained as a case of
regression towards the mean [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], i.e., could be due to noise in the extreme quality
ratings that disappears as more accurate estimates are provided. However, this
possible explanation still needs to be evaluated and, anyway, could explain only
the second trend. These two trends, for very high or very low-quality systems,
yield a high uncertainty in the prediction. Using LOD, we expect to obtain more
tailored predictions (e.g., by identifying software quality trends associated to the
programming language adopted) to reduce prediction uncertainty.
      </p>
      <p>
        Misirli et al.[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] propose the use of Bayesian Networks to make software
quality predictions. As the number of potentially useful features grows (consequently
to LOD enrichment), we will consider this approach in the future. Jing et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
use a dictionary learning-approach that represents a more specialized but limited
approach as compared to our use of LOD.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Enriching Software Quality Prediction with LOD</title>
      <p>Our hypothesis is that by enriching the information about the projects we
analyze with LOD, we can obtain features that are useful for improving the software
quality prediction. For instance, software quality could vary in di erent
industrial sectors or the programming language used could a ect quality evolution.</p>
      <p>
        Our focus is on a dataset provided by the Software Improvement Group,
which consists mainly of projects of Dutch companies and of a few additional
European customers. We enriched the dataset using mainly DBpedia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In the
enrichment process, we encountered the following issues:
Missing information DBpedia contains a description of only 209 companies
located in the Netherlands. Additional companies have been identi ed in
the Dutch DBpedia5, which contains the description of 3.883 companies,
but does not provide information about their location.
      </p>
      <p>Disambiguation Some companies have homonyms. To disambiguate resources
and identify the right URI for a given company, we expect to employ
heuristics based on the company website, its location, and industry sector.
Consistency literals vs. URIs Some classi cations are available in an
inconsistent manner. For instance, industry can appear both as http://dbpedia.
org/ontology/industry and http://dbpedia.org/property/industry. In
some cases, the value of one of these two properties is reported only as a
literal value, thus a ecting the possibility to perform ontological reasoning.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Preliminary Results</title>
      <p>We performed a preliminary analysis on a dataset consisting of 1019 snapshots
of maintainability of 112 companies. These snapshots already presented a rst
industry classi cation provided by SIG. In total, 14 industrial sectors are present.</p>
      <p>
        We computed the semantic similarity between each possible combination of
industrial categories using the Wikipedia distance [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and the WU &amp; Palmer
distance [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. On these data, we performed a series of preliminary analyses:
1. We run a Wilcoxon signed-rank test [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] at 95% con dence level to check
if the observations are signi cantly di erent when grouped per industrial
sector. These results show a weak positive Spearman [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] correlation with
both the Wikipedia (0.07) and the Wu &amp; Palmer (0.14) distances.
2. We computed the same procedure as above by using also the
KolmogorovSmirnov test [
        <xref ref-type="bibr" rid="ref14 ref9">9, 14</xref>
        ] . This resulted in a slightly higher correlation, 0.16 for
the Wikipedia distance and 0.24 for the Wu &amp; Palmer distance.
3. We computed the contrast analysis [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] of the linear combinations of the
observations, again grouped per industrial sector. The resulting contrast
estimators showed a weak correlation with the Wikipedia distance (0.15) and
with the Wu &amp; Palmer distance values (0.12).
4. We grouped a small set of observations aligned with DBpedia by industrial
sector of the companies involved (telecommunication and nancial services).
According to a Wilcoxon signed-rank test at 90% signi cance, the two groups
are signi cantly di erent, according to the Kolmogorov-Smirnov test, not.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Discussion and Future Work</title>
      <p>We present an early stage work about the use of LOD to re ne the precision and
accuracy of software quality prediction. We performed a series of exploratory and
preliminary studies which shows a low correlation between the maintainability
5 http://nl.dbpedia.org
and the industry sector of these projects. These results provide the basis for
further exploration because: (1) the existence of a weak correlation is con rmed by
more tests, hence it is possible that we can identify a subset of the data analyzed
that presents a higher correlation; (2) the di erent methods for computing
semantic similarity and di erent statistical signi cance tests provided signi cantly
di erent results, thus indicating the need for exploring di erent computational
techniques; (3) as shown by the last item of Section 4, the industrial sector
seems to be a discriminant for software quality, although this aspect needs to
be evaluated on larger datasets; and (5) our analyses focused on a limited set of
enrichment features, but several others are utilizable. So, we plan to extend this
research to identify the most robust methods to perform these predictions, and
we will extend these analyses including additional LOD features and sources.
Acknowledgements This work is funded by Amsterdam Data Science.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>H.</given-names>
            <surname>Al-Jamimi</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          .
          <article-title>Machine learning-based software quality prediction models: State of the art</article-title>
          .
          <source>In ICISA, pages 1{4</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Alves</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Visser</surname>
          </string-name>
          .
          <article-title>Static estimation of test coverage</article-title>
          .
          <source>In SCAM</source>
          , pages
          <volume>55</volume>
          {
          <fpage>64</fpage>
          . IEEE Computer Society,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          , G. Kobilarov,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          , and
          <string-name>
            <surname>Z. Ives.</surname>
          </string-name>
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>In ISWC</source>
          , volume
          <volume>4825</volume>
          , pages
          <fpage>722</fpage>
          {
          <fpage>735</fpage>
          . Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>E.</given-names>
            <surname>Bouwers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Correia</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. van Deursen</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Visser</surname>
          </string-name>
          .
          <article-title>Quantifying the analyzability of software architectures</article-title>
          .
          <source>In WICSA</source>
          , pages
          <volume>83</volume>
          {
          <fpage>92</fpage>
          . IEEE,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. T. Dohmen,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ceolin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Visser</surname>
          </string-name>
          .
          <article-title>Towards Building a Software Quality Prediction Model</article-title>
          .
          <source>Technical report</source>
          , Software Improvement Group,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>F.</given-names>
            <surname>Galton</surname>
          </string-name>
          .
          <article-title>Regression towards mediocrity in hereditary stature</article-title>
          .
          <source>The Journal of the Anthropological Institute of Great Britain and Ireland</source>
          ,
          <volume>15</volume>
          :
          <fpage>246</fpage>
          {
          <fpage>263</fpage>
          ,
          <year>1886</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>X.-Y.</given-names>
            <surname>Jing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ying</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S.-S. Wu, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Dictionary learning based software defect prediction</article-title>
          .
          <source>In ICSE</source>
          , pages
          <volume>414</volume>
          {
          <fpage>423</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>S.</given-names>
            <surname>Kan</surname>
          </string-name>
          .
          <article-title>Metrics and Models in Software Quality Engineering</article-title>
          . Pearson,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolmogorov</surname>
          </string-name>
          .
          <article-title>Sulla determinazione empirica di una legge di distribuzione</article-title>
          .
          <source>Giornale dell'Istituto Italiano degli Attuari</source>
          ,
          <volume>4</volume>
          :1{
          <fpage>11</fpage>
          ,
          <year>1933</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          and
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>An open-source toolkit for mining wikipedia</article-title>
          .
          <source>Artif</source>
          . Intell.,
          <volume>194</volume>
          :
          <fpage>222</fpage>
          {
          <fpage>239</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Misirli</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Bener</surname>
          </string-name>
          .
          <article-title>A mapping study on bayesian networks for software quality prediction</article-title>
          .
          <source>In RAISE</source>
          , pages
          <volume>7</volume>
          {
          <fpage>11</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Norris</surname>
          </string-name>
          . Markov chains. Cambridge University Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>R.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Rosnow</surname>
          </string-name>
          .
          <article-title>Contrast analysis : focused comparisons in the analysis of variance</article-title>
          . Cambridge University press,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>N.</given-names>
            <surname>Smirnov</surname>
          </string-name>
          .
          <article-title>Table for Estimating the Goodness of Fit of Empirical Distributions</article-title>
          .
          <source>The Annals of Mathematical Statistics</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ):
          <volume>279</volume>
          {
          <fpage>281</fpage>
          ,
          <year>1948</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Spearman. The proof and measurement of association between two things</article-title>
          .
          <source>Amer. J. Psychol.</source>
          ,
          <volume>15</volume>
          :
          <fpage>72101</fpage>
          ,
          <year>1904</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>F.</given-names>
            <surname>Wilcoxon</surname>
          </string-name>
          .
          <article-title>Individual comparisons by ranking methods</article-title>
          .
          <source>Biometrics Bulletin</source>
          ,
          <volume>1</volume>
          :
          <fpage>80</fpage>
          {
          <fpage>83</fpage>
          ,
          <year>1945</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Verb semantics and lexical selection</article-title>
          .
          <source>In ACL. ACL</source>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>