<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interpreting environmental computational spreadsheets</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science, Network Institute, VU University Amsterdam</institution>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Environmental computational spreadsheets are important tools in supporting decision making. However, as the underlying concepts and relations are not made explicit, the transparency and re-use of these spreadsheets is severely limited. The goal of this project is to provide a semi-automatic methodology for constructing the underlying knowledge level model of environmental computational spreadsheets. We develop and test this methodology in a limited number of case studies. Our methodology combines heuristics on spreadsheet layout and formulas, with existing methods from computer science. We evaluate our constructed model with both the original developers and their peers.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Current environmental issues, like climate change and biodiversity loss, are
universal in their scale and long-term in their impact, their mechanisms are complex,
and empirical data are scarce [1{3]. In addition there is an urgent need to nd
strategies to cope with these issues, and political pressure on the research
community is high [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Environmental computer models are considered essential tools
in supporting environmental decision making by exploring the consequences of
alternative policies or management scenarios [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>
        Environmental computer models are mainly developed and used by domain
scientists and typically implemented as spreadsheets, Fortran programs or in
MatLab. These domain scientists have a knowledge level model [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] in their minds
containing the important concepts in their domain, and corresponding de nitions
and interrelations. In the model development process ( gure 1) they inevitably
make choices about which entities and processes they should include to describe
their study area, and how these should be translated and implemented in their
computer model. In this way their knowledge model is implicitly included in the
computer model, as it is re ected in, for example, the used modelling paradigm,
the model structure, the chosen concepts and their interrelations, and the
mathematical equations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        It is hardly possible to obtain the knowledge level model from the domain
scientists themselves. They may give a limited textual explanation about their
ideas and choices in their publications, but they rather focus on the
computational side of modeling [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In fact, they may not even be aware of the knowledge
Research domain
      </p>
      <p>Spreadsheet
Implicit domain
knowledge</p>
      <p>Domain
scientists</p>
      <p>Design
spreadsheets</p>
      <p>Describe
research</p>
      <p>
        Research
results
Publication
level model in their mind [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The knowledge level model is, however, essential to
understand the meaning and context of the results and insights generated with
the computer model. As a consequence, it is hard to make e cient and e ective
use of environmental computer models by other people than the original
developers [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        The focus of this research is on environmental computer models that are
implemented as spreadsheets, from now on called `environmental computational
spreadsheets'. Spreadsheets are widely used by domain scientists to store and
manipulate quantitative data from their research projects [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. A drawback of
current spreadsheets is that their free format leads to both complex layout of
tables, and sloppy or limited speci cation of the semantics of the data and
calculations [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ]. The goal of this project is therefore to provide a methodology for
making the underlying knowledge level model of environmental computational
spreadsheets explicit. Ideally the various elements in the research process, i.e.
observational data, spreadsheet and publications, could be connected to each
other through this explicit knowledge level model.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Relevancy</title>
      <p>Results of this research could enable peers to discuss and assess the scienti c
quality of environmental computational spreadsheets and to reuse corresponding
results and insights. This could contribute to both scienti c cooperation and
progress, and reliable environmental decision making.</p>
      <p>Our research is focused on spreadsheets from the domain of
environmental science. However, scientists from other domains may have a similar way of
designing and using their spreadsheet models as environmental scientists. We
therefore think that the methods and insights from this study might also be
applied to spreadsheets from other domains, provided that these spreadsheets
contain both domain knowledge and quantitative data.</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>
        Many authors in the eld of environmental science advocate standardization
of the modelling process, summarized to as `Good Modelling Practice', to
enhance transparency of environmental computer models [
        <xref ref-type="bibr" rid="ref1 ref12 ref13">12, 1, 13</xref>
        ]. Similarly,
several studies in computer science, especially in the eld of software engineering,
suggest how scienti c software development could bene t from, for example,
clear documentation, relevant training options for scientists and publication of
source code [
        <xref ref-type="bibr" rid="ref14 ref15 ref17">14, 15, 17</xref>
        ]. The suggested procedures and guidelines will likely yield
more reliable software. However, to guarantee more reliable science, the
knowledge included in that software should also be taken into account.
      </p>
      <p>
        In recent years signi cant progress has been made in the semantic
annotation of scienti c models, data sets, and publications. Many tools and techniques
are avaliable to connect measurements and terms to the identity of observable
entities they quantify [18{21]. A higher level of abstraction that is being
investigated is the semantic annotation of scienti c practice as a whole. The open
provenance model, PROV, 1 helps scientists to document and process
provenance information to ensure reproducibility of their analyses [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Furthermore,
in several scienti c disciplines work ow systems [
        <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
        ] are used to integrate
and analyse data in a correct and meaningful way.
      </p>
      <p>
        Several tools and techniques can be used to annotate tabular data. The Data
Cube vocabulary 2, for example, provides a means for publishing statistical data
as linked data with associated metadata in order to support interpretation and
reproducibility. Existing conversion systems like RDF123 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and XLWrap [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]
allow mapping information from spreadsheets to RDF. And some tools, like
Right eld [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Anzo 3, allow the direct annotation of data inside spreadsheet
tables.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Research Question(s)</title>
      <p>In the above described annotation methods the spreadsheets themselves remain
largely black-boxes. As a consequence, we may miss out on valuable information
on the developers' understanding and interpretation of the system of interest.
However, related work also shows that there are plenty solutions to the issue of
representing scienti c tabular data. As such these studies provide useful tools
and information that can be used as a starting point for present study.
The general research question we wish to answer in our study is the following:
To what extent can the underlying knowledge level model of an environmental
computational spreadsheet be made explicit?
We re ne this question into two more speci c subquestions.
1. How can the underlying knowledge level model of an environmental
computational spreadsheet be adequately described?
1 W3C Provenance Working Group, http://www.w3.org/2011/prov/
2 Data Cube, http://www.w3.org/TR/vocab-data-cube/
3 Anzo, http://www.cambridgesemantics.com
5
6</p>
    </sec>
    <sec id="sec-5">
      <title>Hypotheses</title>
    </sec>
    <sec id="sec-6">
      <title>Preliminary results</title>
      <p>An adequate description of the underlying knowledge level model is de ned
as a description that
{ agrees with the views of the original developers of the spreadsheets.
{ can be understood and applied by the original developers of the
spreadsheets and their peers.
{ allows representation of domain concepts, their hierarchical and
property relations, and the computational relations that exist between these
concepts
2. What are the requirements for a methodology for constructing the underlying
knowledge level model of an environmental computational spreadsheet?
When we apply our methodology to an environmental computational
spreadsheet, we expect that the resulting constructed knowledge level model is an
adequate description of the underlying knowledge level model.</p>
      <p>We did two case studies on an existing environmental computer model,i.e., a
spreadsheet model that enables policy analyses concerning the Dutch energy
system .</p>
      <p>
        In the rst case study [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] we manually analyzed the design of the tables and
the formulas in the spreadsheets 4. We semantically characterized the underlying
concepts and their interrelations ( gure 2) and represented these as an
instantiation of an existing ontology, the OM Ontology for units of Measure and related
concepts [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The main concepts and their interrelations as we identi ed them
in our resulting ontology did not con ict with the developer's views. However,
we also discovered that the developers see their models mainly as instruments to
perform simulation studies, and therefore focus on the computational aspects.
4 Spreadsheet Examples, http://semanticweb.cs.vu.nl/edesign/
network of interconnected spreadsheet cells ( gure 3). We used network analysis
to determine which nodes in the graph are the most important, and manually
connected these in a simplifed calculation work ow.
In this project we aim at developing a methodology for semi-automatic
construction of the underlying knowledge model of an environmental computational
spreadsheet. As described above, there are no similar studies on this topic, nor
is it possible to access the knowledge level model in the minds of the original
developers of environmental computational spreadsheets. We therefore consider
it not feasible to set up a study based on quantitative experiments. Instead we
choose an approach based on the analyses of a limited number of case studies,
and as a consequence, our research has an exploratory character.
      </p>
      <p>Our case studies are all scienti c spreadsheet models of existing research
projects from the domain of environmental science. We have access to the actual
spreadsheets and corresponding datasets, as well as to the publications describing
the models and analyses. Furthermore, we have personal contact with the model
developers and users.</p>
      <p>We develop our methodology based on the in-depth, qualitative analysis of
one case study. We will manually analyze the layout of the spreadsheet tables, as
well as the formulas connecting the spreadsheet cells. We determine to what
extent the observed patterns provide insight in the semantics of the content of the
tables, and record our ndings in heuristics. Spreadsheet terms can be matched
automatically with concepts of external vocabularies on domain concepts, and on
quantitative tabular data. We combine this matching with our layout heuristics
to recognize the concepts in the spreadsheets and their interrelations. In
addition, we will automatically trace the dependencies between spreadsheet cells
through formulas and analyze the resulting networks using techniques for
network analysis. We combine these analyses with our heuristics on formulas to
construct the calculation work ow in the spreadsheets.</p>
      <p>Research question 1 is studied by focusing on the performance of our method
in each case study. The di erent steps in our methodology of constructing the
knowledge level model are performed manually by the original developers, and
their results are compared with results from our semi-automatic method. We test
the applicability of the constructed model by using it to connect concepts from
the spreadsheets, with concepts from corresponding publications, or data sets.
In a separate user study we will test to what extent peers are able to understand
and apply the constructed knowledge level model.</p>
      <p>Research question 2 is studied by focusing on the di erent techniques that are
used to describe the knowledge level model. The use of external vocabularies is
evaluated by determining how many of the spreadsheet terms could be matched,
and how relevant these matches are. We also determine which properties of these
vocabularies in uence this matching. The use of network analysis techniques is
evaluated by determining to what extent these techniques are able to recognize
the important variables, as indicated by the original developers, in the
calculation work ow. We determine which properties of the spreadsheets in uence the
performance of our method.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Evaluation plan</title>
      <p>In order to test our hypothesis we will formulate measurable de nitions on what
it means for original developers and peers to understand and apply the
constructed knowledge level model. Possible indicators we could use are, for
example,
{ the number of concepts, relations and variables that occur both in the
constructed model and in the manual analysis of the original developers.
{ the number of connections that can be made from the spreadsheet to
corresponding publications and datasets.
9</p>
    </sec>
    <sec id="sec-8">
      <title>Re ections</title>
      <p>We think our approach is likely to succeed as it is targeted at existing
environmental computational spreadsheets. We expect that studying the patterns
in these spreadsheets will provide us useful insights on environmental modeling.
We also see several promising external developments. Firstly, there is a
growing awareness of both the importance of open source code and data, and the
importance of methods to provide corresponding credits to modelers and data
providers. Besides, there is an increasing availability of external domain
vocabularies.</p>
      <p>
        This PhD research is now at the half way stage. Current work is an extension of
our rst case study (section 6) and involves the development of a semi-automatic
method for de ning the concepts and interrelations in spreadsheets. We use the
external vocabularies AGROVOC [27] and OM[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], to map and categorize the
spreadsheet terms. The plan for the near future is to continue the work of our
second case study by developing a semi-automatic method for the construction
of the calculation work ow.
      </p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>This publication was supported by the Data2Semantics project in the Dutch
national program COMMIT. Guus Schreiber and Jan Top (supervisors), Paul
Groth and Lora Aroyo are acknowledged for providing useful comments and
suggestions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Jakeman, a.,
          <string-name>
            <surname>Letcher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Norton</surname>
          </string-name>
          , J.:
          <article-title>Ten iterative steps in development and evaluation of environmental models</article-title>
          .
          <source>Environmental Modelling &amp; Software</source>
          <volume>21</volume>
          (
          <issue>5</issue>
          ) (May
          <year>2006</year>
          )
          <volume>602</volume>
          {
          <fpage>614</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Schmolke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thorbek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DeAngelis</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grimm</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Ecological models supporting environmental decision making: a strategy for the future</article-title>
          .
          <source>Trends in ecology &amp; evolution</source>
          <volume>25</volume>
          (
          <issue>8</issue>
          ) (
          <year>August 2010</year>
          )
          <volume>479</volume>
          {
          <fpage>86</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>van der Sluijs</surname>
            ,
            <given-names>J.P.:</given-names>
          </string-name>
          <article-title>A way out of the credibility crisis of models used in integrated environmental assessment</article-title>
          .
          <source>Futures</source>
          <volume>34</volume>
          (
          <issue>2</issue>
          ) (
          <year>March 2002</year>
          )
          <volume>133</volume>
          {
          <fpage>146</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Newell</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The knowledge level</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>18</volume>
          (
          <issue>1</issue>
          ) (
          <year>January 1982</year>
          )
          <volume>87</volume>
          {
          <fpage>127</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Villa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Athanasiadis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rizzoli</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          :
          <article-title>Modelling with knowledge: A review of emerging semantic approaches to environmental modelling</article-title>
          .
          <source>Environmental Modelling &amp; Software</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ) (May
          <year>2009</year>
          )
          <volume>577</volume>
          {
          <fpage>587</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>De Vos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janssen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Bussel</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kromdijk</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Van Vliet</surname>
            ,
            <given-names>J.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Top</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Are environmental models transparent and reproducible enough ? In Wongsosaputro</article-title>
          , J.,
          <string-name>
            <surname>Pauwels</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chan</surname>
          </string-name>
          , F., eds.
          <source>: Proceedings of 19th International Congress on Modelling and Simulation</source>
          . (
          <year>2011</year>
          )
          <volume>2954</volume>
          {
          <fpage>2961</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>De Vos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Hage</surname>
            ,
            <given-names>W.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ros</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreiber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Reconstructing Semantics of Scienti c Models : a Case Study</article-title>
          .
          <source>In: Proceedings of the OEDW workshop on Ontology engineering in a data driven world, EKAW</source>
          <year>2012</year>
          , Galway, Ireland (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wolstencroft</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Owen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horridge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krebs</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoep</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>du Preez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>RightField: embedding ontology annotation in spreadsheets</article-title>
          .
          <source>Bioinformatics</source>
          (Oxford, England)
          <volume>27</volume>
          (
          <issue>14</issue>
          ) (
          <year>July 2011</year>
          )
          <year>2021</year>
          {
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Rocha</given-names>
            <surname>Bernardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Mota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Santanche</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Extracting and Semantically Integrating Implicit Schemas from Multiple Spreadsheets of Biology based on the Recognition of their Nature</article-title>
          .
          <source>Journal of Information and Database Management</source>
          <volume>4</volume>
          (
          <issue>2</issue>
          ) (
          <year>2013</year>
          )
          <volume>104</volume>
          {
          <fpage>113</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Rijgersberg</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wigham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Top</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>How semantics can improve engineering processes: A case of units of measure and quantities</article-title>
          .
          <source>Advanced Engineering Informatics</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          ) (
          <year>April 2011</year>
          )
          <volume>276</volume>
          {
          <fpage>287</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parr</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sachs</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>RDF123 : From Spreadsheets to RDF</article-title>
          .
          <source>In: The Semantic Web-ISWC 2008</source>
          , Springer Berlin Heidelberg (
          <year>2008</year>
          )
          <volume>451</volume>
          {
          <fpage>466</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Refsgaard</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Modelling guidelinesterminology and guiding principles</article-title>
          .
          <source>Advances in Water Resources</source>
          <volume>27</volume>
          (
          <issue>1</issue>
          ) (
          <year>January 2004</year>
          )
          <volume>71</volume>
          {
          <fpage>82</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rykiel</surname>
            ,
            <given-names>E.J.J.:</given-names>
          </string-name>
          <article-title>Testing ecological models: the meaning of validation</article-title>
          .
          <source>Ecological Modelling</source>
          <volume>90</volume>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hannay</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macleod</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langtangen</surname>
            ,
            <given-names>H.P.</given-names>
          </string-name>
          , Wilson, G.:
          <article-title>How Do Scientists Develop</article-title>
          and Use Scienti c Software ?
          <source>In: Proceedings of the 2009 ICSE workshop on Software Engineering for Computational Science and Engineering</source>
          , IEEE Computer Society (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Segal</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morris</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Developing scienti c software, Part 2</article-title>
          .
          <source>IEEE software 26(1)</source>
          (
          <year>2009</year>
          )
          <fpage>79</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Merali</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Why scienti c programming doesn't compute</article-title>
          .
          <source>Nature</source>
          <volume>467</volume>
          (
          <year>2010</year>
          ) 6{
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berlin</surname>
            ,
            <given-names>F.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seaborne</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labs</surname>
          </string-name>
          , H.p.
          <article-title>: D2RQ Treating Non-RDF Databases as Virtual RDF Graphs</article-title>
          .
          <source>In: Proceedings of the 3rd International Semantic Web Conference (ISWC2004)</source>
          . (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velardi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cucchiarelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Quantitative and Qualitative Evaluation of the OntoLearn Ontology Learning System</article-title>
          .
          <source>In: Proceedings of the 20th international conference on Computational Linguistics</source>
          . (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.c.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>LIMES - A Time-E cient Approach for Large-Scale Link Discovery on the Web of Data</article-title>
          .
          <source>In: Proceedings of the Twenty-Second international joint conference on Arti cial Intelligence</source>
          , Volume Three., AAAI Press, (
          <year>2011</year>
          )
          <volume>2312</volume>
          {
          <fpage>2317</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Volz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaedke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
          </string-name>
          , G.:
          <article-title>Silk A Link Discovery Framework for the Web of Data. (</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cli</surname>
            <given-names>ord</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Freire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Futrelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Kwasnikowska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Miles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Missier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Myers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Plale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Simmhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Stephan</surname>
          </string-name>
          , E., den Bussche, J.V.:
          <article-title>The Open Provenance Model core speci cation (v1.1)</article-title>
          .
          <source>Future Generation Computer Systems</source>
          <volume>27</volume>
          (
          <issue>6</issue>
          ) (
          <year>June 2011</year>
          )
          <volume>743</volume>
          {
          <fpage>756</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. Ludascher,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Altintas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Berkley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Higgins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Jaeger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.B.</given-names>
            ,
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.a.</given-names>
            ,
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Scienti c work ow management and the Kepler system</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          <volume>18</volume>
          (
          <issue>10</issue>
          ) (
          <year>August 2006</year>
          )
          <volume>1039</volume>
          {
          <fpage>1065</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sroka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hidders</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Missier</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A formal semantics for the Taverna 2 work ow model</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          <volume>76</volume>
          (
          <issue>6</issue>
          ) (
          <year>September 2010</year>
          )
          <volume>490</volume>
          {
          <fpage>508</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Langegger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolfram</surname>
          </string-name>
          , W.:
          <article-title>XLWrap Querying and Integrating Arbitrary Spreadsheets with SPARQL</article-title>
          . (
          <year>2009</year>
          )
          <volume>359</volume>
          {
          <fpage>374</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>De Vos</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>van Hage</surname>
            ,
            <given-names>W.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wielemaker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreiber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Knowledge Representation in Scienti c Models and their Publications : a Case Study</article-title>
          .
          <source>In: Proceedings of K-CAP 2013 Knowledge Capture Conference</source>
          , Ban , Canada (
          <year>2013</year>
          ) 1{
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Soergel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lauser</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fisseha</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keizer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Reengineering Thesauri for New Applications : the AGROVOC Example</article-title>
          .
          <source>Journal of Digital Information</source>
          <volume>4</volume>
          (
          <issue>4</issue>
          ) (
          <year>2004</year>
          )
          <volume>1</volume>
          {
          <fpage>15</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>