<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sustainability of Evaluations Presented in Research Publications</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Ontology Engineering Group, Departamento de Lenguajes y Sistemas Informaticos e Ingenier a Software Facultad de Informatica, Universidad Politecnica de Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This position paper discusses how research publication would bene t of an infrastructure for evaluation entities that could be used to support documenting research e orts (e.g., in papers or blogs), analysing these e orts, and building upon them. As a concrete example in the domain of semantic technologies, the paper presents the SEALS Platform and discusses how such platform can promote research publication.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The way of publishing evaluation-related information in research papers should
be rethought to facilitate the use of such information.</p>
      <p>The limited extension of research papers does not allow a full description
of all the entities involved in evaluations (evaluation work ows, test data, tools
and results) and of the concrete context in which evaluations were performed.
Therefore, it is di cult if not impossible to reproduce evaluations described
in research papers and to validate them; this forces researchers in most cases
to blindly trust in the paper claims. Besides, both technologies and evaluation
methods evolve over time and the evaluation data included in the paper becomes
rapidly outdated.</p>
      <p>Another perspective to take into account is that of advancing research by
building upon existing one. De ning and performing evaluations is expensive
and prone to errors. This is mainly because, besides the lack of full evaluation
descriptions mentioned above, most lessons learnt during evaluation (both
positive and negative) are not explicit in research papers. Furthermore, performing
complex analyses across research papers (e.g., nding correlations between the
results of di erent evaluations) is currently not possible.</p>
      <p>The goal of this paper is to discuss how research would bene t of an
infrastructure for evaluation entities that could be used to support documenting
research e orts (e.g., in papers or blogs), analysing these e orts, and building
upon them.</p>
      <p>Such infrastructure would allow anyone reading or reviewing a paper to
completely analyse the evaluations presented in the paper and to validate them,
taking advantage of dynamic and enhanced result visualisations.</p>
      <p>Besides, it would permit anyone to reproduce the evaluation presented in the
paper under the same settings or using updated or alternative versions of the
evaluation entities.</p>
      <p>Furthermore, anyone interested in building upon existing evaluations could
reuse the evaluation presented in the paper (fully or parts of it) and even combine
the results from evaluations in di erent papers.</p>
      <p>Clearly, someone could disagree with the above-mentioned claims; the main
stands against them could be the following:
{ Refusal to unveil evaluation details. In research environments, people are
used to having other people review their work in detail, so this should be no
problem. Besides, being against this would incline people to think that the
researcher is hiding something.
{ Refusal to share work with others. This opinion is also not expected since in
research environments people are usually eager to be reused cited.
{ Refusal to devote e ort to share evaluation details. Even if researchers
acknowledge the added value of sharing their evaluations, they will be reluctant
to do so unless the bene ts compensate their spent e orts.
{ Refusal to reuse work from others. Even if the do-it-yourself attitude is
characteristic of computer science researchers, they are also aware of the bene ts
of reuse; therefore, this is something that should not pose rejection.</p>
      <p>
        The SEALS (Semantic Evaluation at Large Scale) European project1 is
developing an infrastructure (the SEALS Platform) that o ers independent
computational and data resources for the evaluation of semantic technologies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Next, the paper presents an overview of the SEALS Platform and then
discusses how such infrastructure could support the publishing and management of
evaluation information in research papers, providing di erent bene ts along the
lines presented above.
2</p>
      <p>
        An Infrastructure for Semantic Technology Evaluation
The idea of software evaluation followed in the SEALS Platform is largely
inspired by the notion of evaluation as de ned by the ISO/IEC 14598 standard
on software product evaluation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In any evaluation a given set of tools are
executed, following a given evaluation work ow and using determined test data.
As an outcome of this process, a set of evaluation results is produced.
      </p>
      <p>
        This high-level classi cation of software evaluation entities can be further
re ned as needed; a detailed description of them and their life cycles can be
found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For example, in accordance with the approach followed in the IEEE
1061 standard for a software quality metrics methodology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], evaluation results
are classi ed according to their provenance, di erentiating raw results (those
evaluation results directly generated by tools) from interpreted results (those
generated from other evaluation results).
      </p>
      <sec id="sec-1-1">
        <title>1 http://www.seals-project.eu/</title>
        <p>Technology</p>
        <p>Providers
Run1me  
Evalua1on  </p>
        <p>Service  
SEALS Repositories</p>
        <p>Evaluation Organisers</p>
        <p>SEALS  Portal  
Evaluation
requests</p>
        <p>Entity
management
requests</p>
        <p>SEALS    </p>
        <p>Service  Manager  
Test  Data    
Repository  </p>
        <p>Service  </p>
        <p>Tools    
Repository  </p>
        <p>Service  </p>
        <p>Results    
Repository  </p>
        <p>Service  </p>
        <p>
          Moreover, our entities include not only the results obtained in the
evaluation but also any contextual information related to such evaluation, a need also
acknowledged by other authors [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. To this end, we also represent the
information required for automating the execution of an evaluation description in the
platform, which, with the rest of the entities presented, yields traceable and
reproducible evaluation results.
        </p>
        <p>The SEALS Platform has been developed around these evaluation entities
and following a service-oriented approach. The architecture of the platform
comprises a number of components, shown in Figure 1, which are described below.
{ SEALS Portal. The SEALS Portal provides a web user interface for
interacting with the SEALS Platform. Thus, the portal will be used by the users
for the management of the entities in the SEALS Platform, as well as for
requesting the execution of evaluations.
{ SEALS Service Manager. The SEALS Service Manager is the core
module of the platform and is responsible for coordinating the other platform
components and for maintaining consistency within the platform. This
component exposes a series of services that provide programmatic interfaces for
the SEALS Platform. Thus, apart from the SEALS Portal, the services
offered may be also used by third party software agents.
{ SEALS Repositories. These repositories manage the entities used in the
platform: test data, tools, results, and evaluation work ows.</p>
        <p>Technology
Adopters</p>
        <p>Software agents,
i.e., technology evaluators</p>
        <p>Evalua1on  </p>
        <p>Descrip1ons  
Repository  Service  
{ Runtime Evaluation Service. The Runtime Evaluation Service is used
to automatically evaluate a certain tool according to a particular evaluation
description and using some speci c test data.</p>
        <p>
          All the evaluation entities stored in the platform are described according to
a set of OWL ontologies2 [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Since the entities presented above share a number
of common properties, we developed an upper ontology to represent them, as
well as di erent ontologies covering each entity domain. During the de nition of
the ontologies we tried, when possible, to reuse current standards and models
(i.e., Dublin Core, FOAF, VCard).
3
        </p>
        <p>Publication and Management of Evaluation Information
This section discusses how the SEALS Platform could support the publication
and management of evaluation information.</p>
        <p>The SEALS Platform o ers manual and programmatic access to the
evaluation entities stored in its repositories. This allows linking the evaluation resources
mentioned in research papers to the actual resources stored in the platform.
Besides, if reverse links were created from the evaluation entities to research papers,
networks of papers around concrete evaluations could be built.</p>
        <p>The SEALS Platform also allows storing di erent versions of tools, evaluation
work ows and test data. This way, it maintains the traceability from the concrete
evaluation used in one paper to those evaluations that include updated versions
of tools, evaluation work ows or test data.</p>
        <p>All the evaluation entities stored in the SEALS Platform are described
using ontologies with the aim of having consensual and interoperable descriptions.
These machine-processable descriptions can be published in the Web or be
embedded in research papers. Furthermore, it provides dynamic and interactive
visualisations of evaluation results that could be used in non-standard research
papers (e.g., multimedia or interactive documents).</p>
        <p>In the SEALS Platform, evaluation reproducibility is a main requirement. To
this end, evaluations are only executed over persistent (i.e., unmodi able) entities
and the whole evaluation execution context is stored. This allows replicating the
concrete evaluation presented in a research paper at any moment and by anyone.</p>
        <p>All the evaluation entities can not only be accessed but also be reused both
inside and outside the SEALS Platform. This reuse can be performed as a whole
(e.g., reusing some test data in another evaluation infrastructure) or partially
(e.g., evaluation work ows are de ned with the BPEL language and new
workows can be de ned from existing work ows and services).</p>
        <p>
          Finally, since evaluation results are represented following common schemas
(i.e., ontologies), researchers could exploit these results in unexpected ways. To
allow this, we have de ned a quality model for semantic technologies that de nes
the main quality characteristics of such technologies and allows the combination
and comparison of results from di erent evaluations [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>2 http://www.seals-project.eu/ontologies/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusions</title>
      <p>This paper proposes to support research publications (or any other type of
research documentation) through an infrastructure for evaluation entities. Having
such infrastructure would allow, on the one hand, connecting research
publications with the actual evaluations used in them and, on the other hand,
interconnecting di erent research e orts.</p>
      <p>The SEALS Platform aims to support these ideas in the domain of semantic
technologies with the ultimate goal of increasing the maturity of the semantic
research community by enriching the body of knowledge on semantic technology
evaluation and by encouraging an experimentation-based research.</p>
      <p>However, the project is still in its way to achieve the approach presented in
this paper since functionalities for linking evaluations with publications are not
planned yet. To this end, future challenges to be faced are not only technological
but also social (e.g., it requires greater commitment since researchers have to
invest more e ort than they are now) or legal (e.g., important issues are the
access and use policies for evaluation data).</p>
      <p>Furthermore, the success of such approach will depend on the existence of
software technologies that are coupled to researchers' working environments and
that leverage the e ort of using an infrastructure such as the SEALS Platform
in day-to-day research.</p>
      <p>Acknowledgements
This work has been supported by the SEALS European project (FP7-238975).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Garc</surname>
            a-Castro,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esteban-Gutierrez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Towards an infrastructure for the evaluation of semantic technologies</article-title>
          .
          <source>In: Proceedings of the eChallenges 2010 Conference</source>
          , Warsaw, Poland (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. ISO/IEC: ISO/IEC 14598-
          <article-title>6: Software product evaluation - Part 6: Documentation of evaluation modules (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Garc</surname>
            a-Castro,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esteban-Gutierrez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kerrigan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grimm</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An ontology model to support the automatic evaluation of software</article-title>
          .
          <source>In: Proceedings of the 22nd International Conference on Software Engineering and Knowledge Engineering (SEKE</source>
          <year>2010</year>
          ), Redwood City, CA, USA (
          <year>2010</year>
          )
          <volume>129</volume>
          {
          <fpage>134</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. IEEE: IEEE 1061
          <article-title>-1998</article-title>
          .
          <article-title>IEEE Standard for a Software Quality Metrics Methodology (</article-title>
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kitchenham</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hughes</surname>
          </string-name>
          , R.T.,
          <string-name>
            <surname>Linkman</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          :
          <article-title>Modeling software measurement data</article-title>
          .
          <source>IEEE Trans. Softw. Eng</source>
          .
          <volume>27</volume>
          (
          <year>2001</year>
          )
          <volume>788</volume>
          {
          <fpage>804</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Radulovic</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a-Castro,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Towards a Quality Model for Semantic Technologies</article-title>
          .
          <source>In: Proceedings of the 11th International Conference on Computational Science and Its Applications (ICCSA</source>
          <year>2011</year>
          ), Santander, Spain, (To be published) (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>