<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collaborative platforms for streamlining work ows in Open Science</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konrad U. Forstner</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gregor Hagedorn</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudia Koltzenburg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Fabiana Kubke</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Mietchen</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Claudia Koltzenburg Managing editor of Cellular Therapy and Transplantation, CTT</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Anatomy with Radiology, University of Auckland</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Molecular Infection Biology University of Wurzburg D-97080 Wurzburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Julius Kuhn-Institute Federal Research Center for Cultivated Plants Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Research Centre for Infectious Diseases University of Wurzburg D-97080 Wurzburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Despite the internet's dynamic and collaborative nature, scientists continue to produce grant proposals, lab notebooks, data les, conclusions etc. that stay in static formats or are not published online and therefore not always easily accessible to the interested public. Because of limited adoption of tools that seamlessly integrate all aspects of a research project (conception, data generation, data evaluation, peerreviewing and publishing of conclusions), much e ort is later spent on reproducing or reformatting individual entities before they can be repurposed independently or as parts of articles. We propose that work ows - performed both individually and collaboratively - could potentially become more e cient if all steps of the research cycle were coherently represented online and the underlying data were formatted, annotated and licensed for reuse. Such a system would accelerate the process of taking projects from conception to publication stages and allow for continuous updating of the data sets and their interpretation as well as their integration into other independent projects.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A major advantage of such work ows is the increased transparency, both
with respect to the scienti c process as to the contribution of each
participant. The latter point is important from a perspective of
motivation, as it enables the allocation of reputation, which creates incentives
for scientists to contribute to projects. Such work ow platforms o ering
possibilities to ne-tune the accessibility of their content could gradually
pave the path from the current static mode of research presentation into
a more coherent practice of open science.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>Like most areas of today's life, science has dramatically changed since the
advent of the internet. However, the transformation that has taken place until now
is just the tip of the iceberg. In the following, we want to discuss the mostly
underutilized potential of representing all aspect of science in collaboratively
used online work ow platforms. Since such platforms could help to realize Open
Science, transparency of the funding cycles and access to all data in the research
process, we will shed light on this special aspect and make recommendations
regarding implementations.</p>
      <p>
        While there are numerous projects developing and applying so called Virtual
Research Environments (VRE) - also known as Collaboratories - covering
selected stages of the scienti c process, a platform spanning every phase is missing
so far [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Technically overcoming such gaps and creating a seamless transition
from bench to publication could speed up the research and, with it, the
generation, distribution and reuse of knowledge.
      </p>
    </sec>
    <sec id="sec-3">
      <title>The scienti c work ow in open VREs</title>
      <sec id="sec-3-1">
        <title>Conception and project planning</title>
        <p>Independent of the nature of a research endeavor - hypotheses-driven or
datadriven, performed by a single person or a team - a solid conception phase is the
crucial basis for every project. Despite today's common practice of limiting this
phase to a small group of people, utilizing collective intelligence during the
conception phase could help to avoid redundant research and to improve the design
of the study. As the complexity and scope of scienti c projects are increasing,
the application of project management tools can be useful for managing the
processes and parties involved.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experiments and data generation</title>
        <p>
          Today, data generation in academic research continues to rely strongly on
manual labor. While this is mostly due to the relative low cost of labor force resulting
from the academic system and the limited interdisciplinary education of science
and engineering, the high potential of automation is mostly neglected. Not only
could the e ciency of invested labor be improved by automation, but also
reproducibility could be signi cantly increased. To make this a ordable for the broader
research community, a shift from siloed proprietary devices to well-documented
pieces of standardized, open-source hardware developed by the scienti c
community itself in cooperation with potential vendors is needed. Open hardware
platforms like Arduino [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] could o er starting points for such a development and
rst example of such tools are available (e.g. OpenPCR [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]). The devices could
and should enrich the primary data with further metadata, convert them into
semanti ed formats and directly upload the output into online repositories.
        </p>
        <p>
          One promising example which visualizes the potential of such automation
of otherwise quite labor-intensive research is the robot scientist ADAM [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The
streamlining of mechanical steps and the evaluation of results would bene t from
formal languages that describe the necessary procedures and make the design
and exchange of experimental setups easy [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. As a long term goal, scientists
would mostly engage in programming experiments and engineering the system
to automate those steps that have been performed manually so far. The motto
\work on the system, not in the system" should guide this development.
2.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Data release</title>
        <p>The online release of experimentally generated data should be done shortly after
the generation and can potentially happen in real time. Downstream
analysis within the research project but also the reuse by other parties should be
kept in mind when selecting data formats. These should, as far as possible, be
non-proprietary, machine readable (semantically enriched) and common for the
respective domain of research. If no format ful lls all these requirements, the
conversion into alternative formats should be permitted. Access to the data could
take place via a web interface or domain speci c clients. Especially for large or
highly accessed data sets, the additional distribution via peer-to-peer networks
is recommended.
2.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Data analysis</title>
        <p>
          Since every step in the data analysis should be transparent and easily
reproducible, it should take place preferably in the proposed platform, too. Systems
like the analysis work ow tool Taverna [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] could be used for such processing.
Already today, many research institution o er grid computing infrastructure for
such purposes. Analyses using external tools, especially GUI-tools that do not
o er any possibility to log the performed actions, should be avoided if possible,
as otherwise documentation has to be created manually. For some
computationally intensive analyses the use of shared systems is a more economical usage of
the needed infrastructure, provided the management overhead does not exceed
the computational e ciency gain. As done for the raw experimental data, the
protocols and the result of the data processing should be documented and stored
in repositories to be accessible.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Knowledge generation</title>
        <p>
          The results of analytical processing as well as the raw data can be used by
scientists - or machines [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] - to draw conclusions and to generate knowledge
out of the available information in a well documented way. The platform should
assist to make this happen collaboratively by o ering commenting and rating of
statements. Discussions - text, audio- and/or video-based - should be recorded
to make the path to nding reconstructible.
2.6
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Final publication</title>
        <p>As documentation of every step is an inherent feature of the work ow, the nal
publications resulting from a study can be short reports linking to the major
outcomes and putting them into the scienti c context. The platform should o er
functionalities to perform open peer-review of this nal report.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Implementation</title>
      <sec id="sec-4-1">
        <title>Technology</title>
        <p>As shown above, the many building blocks of a complete scienti c work ow
already exist and only need to be connected seamlessly. The development of open
standards de ning the required interfaces of these parts could enable di erent
parties to assemble the pieces into a consistent work ow and to add further
needed parts. This would o er the possibility to implement a platform either as
one monolithic application or as separate interacting and exchangeable units.
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Funding</title>
        <p>
          Of similar importance as the technical realization is the adaptation of scienti c
culture and funding policies. While research institutions like the National
Institutes of Health (US) or the Welcome Trust (UK) already require open access for
nal peer-review manuscripts that results from research they funded [
          <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
          ], the
regulations are much weaker for the underlying data, and almost nonexistent for
proper annotation. However, the rst attempts to establish such requirements
are on the horizon [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
3.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Licensing</title>
        <p>
          The default copyright restrictions in most jurisdictions hamper the reuse of data.
It is therefore highly desirable that, with very few exceptions, each entity
generated in the research process is explicitly published under a less restrictive license,
e.g., the ones o ered by Creative Commons [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] or is released into the public
domain. As the latter concept may di er or be missing in some countries, release
through the CC0 license [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is recommended.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Reputation</title>
        <p>
          The gain of reputation is the most important incentive for scientists. It is
currently mostly determined on the basis of publications in scienti c journals and
the related measure of success in funding applications. As every contribution to
a research project can be attributed to a distinct person and could be rated by
others, the allocation of reputation is an inherent element of the proposed
platform. The connection to research identi ers like ORCID [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and the analysis of
such microcontributions could assemble a precise image of a scientist's skills and
achievements.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Challenges</title>
      <p>As stated above, considering the allocation of reputation and funding in science
is crucial when redesigning scienti c processes. To bridge a transient phase until
the suggested political changes have taken place, ne granular access control in
the research work ow platform could permit that the technology is adapted by
scientists despite objection regarding the loss of reputation. With such a control
in place, the full process could be opened up after the nal publication or at any
other desired time.</p>
      <p>It is very unlikely that there will be one single platform that can ful ll the
requirements of all scienti c domains. Building and maintaining completely
independent platforms for each domains, on the other hand, may not be sustainable.
A modular and exible system, where possible re-using industry standard
software is therefore called for.</p>
      <p>
        Projects presently exploring this are, e.g.:
{ the eSciDoc platform which builds on the open-source repository software
Fedora Commons [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and is mainly developed for the for the Max Planck
Society has a similar aim and strategy [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
{ the FP7 funded Virtual Biodiversity Research and Access Network for
Taxonomy" (ViBRANT) [16{18]
      </p>
      <p>
        They span data collection, analysis and publishing (in collaboration with
Pensoft Publishers), are based on the established open-source platforms Drupal
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and Mediawiki [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and equipped with speci c extensions.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Annamaria</given-names>
            <surname>Carusi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Reimer</surname>
          </string-name>
          .
          <source>Virtual Research Environment - Collaborative Landscape Study. JISC</source>
          . pp.
          <fpage>72</fpage>
          -
          <lpage>24</lpage>
          2010.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. Arduino, http://www.arduino.cc/</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3. OpenPRC, http://openpcr.org/</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ross D. King</surname>
            , Jem Rowland, Stephen G. Oliver, Michael Young, Wayne Aubrey, Emma Byrne, Maria Liakata, Magdalena Markham, Pinar Pir,
            <given-names>Larisa N.</given-names>
          </string-name>
          <string-name>
            <surname>Soldatova</surname>
            , Andrew Sparkes,
            <given-names>Kenneth E.</given-names>
          </string-name>
          <string-name>
            <surname>Whelan</surname>
            ,
            <given-names>Amanda</given-names>
          </string-name>
          <string-name>
            <surname>Clare</surname>
          </string-name>
          .
          <article-title>The automation of science</article-title>
          .
          <source>Science. 3 April</source>
          <year>2009</year>
          : Vol.
          <volume>324</volume>
          no. 5923 pp.
          <fpage>85</fpage>
          -
          <lpage>89</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Larisa N Soldatova</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ross D King</surname>
          </string-name>
          .
          <article-title>An ontology of scienti c experiments</article-title>
          .
          <source>J R Soc Interface</source>
          .
          <source>2006 Dec</source>
          <volume>22</volume>
          ;
          <issue>3</issue>
          (
          <issue>11</issue>
          ):
          <fpage>795</fpage>
          -
          <lpage>803</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Duncan</given-names>
            <surname>Hull</surname>
          </string-name>
          , Katy Wolstencroft, Robert Stevens, Carole Goble,
          <string-name>
            <surname>Mathew R. Pocock</surname>
            ,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Tom</given-names>
          </string-name>
          <string-name>
            <surname>Oinn</surname>
          </string-name>
          .
          <article-title>Taverna: a tool for building and running work ows of services Nucleic Acids Res</article-title>
          .
          <source>2006 Jul</source>
          <volume>1</volume>
          ;
          <fpage>34</fpage>
          (
          <issue>Web Server issue</issue>
          ):
          <fpage>W729</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Michael</given-names>
            <surname>Schmidt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hod</given-names>
            <surname>Lipson</surname>
          </string-name>
          .
          <article-title>Distilling free-form natural laws from experimental data</article-title>
          .
          <source>Science. 2009 Apr</source>
          <volume>3</volume>
          ;
          <issue>324</issue>
          (
          <issue>5923</issue>
          ):
          <fpage>81</fpage>
          -
          <lpage>5</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>NIH</given-names>
            <surname>Public Access Policy Details</surname>
          </string-name>
          , http://publicaccess.nih.gov/policy.htm
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Welcome</given-names>
            <surname>Trust</surname>
          </string-name>
          <article-title>Open access policy</article-title>
          , http://www.wellcome.ac.uk/Aboutus/Policy/Policy-and
          <string-name>
            <surname>-</surname>
          </string-name>
          position-statements/WTD002766.htm
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Andrew</surname>
          </string-name>
          J Vickers.
          <article-title>Making raw data more widely available</article-title>
          .
          <source>BMJ</source>
          .
          <year>2011</year>
          ;
          <volume>342</volume>
          :
          <fpage>d2323</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>11. Creative Commons, http://creativecommons.org/</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>12. CC0, http://creativecommons.org/publicdomain/zero/1.0/</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>13. ORCID, http://orcid.org/</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>14. Feudora Commons, http://fedora-commons.org/</mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Malte</surname>
            <given-names>Dreyer</given-names>
          </string-name>
          , Ulla Tschida, Natasa Bulatovic, Matthias Razum. eSciDoc
          <article-title>- a Scholarly Information and Communication Platform for the Max Planck Society</article-title>
          . German e-Science Conference, Baden-Baden.
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>16. ViBRANT, http://vbrant.eu</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Dave</surname>
            <given-names>Roberts</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Vince</given-names>
            <surname>Smith. ViBRANT - Virtual Biodiversity</surname>
          </string-name>
          Research and
          <article-title>Access Network for Taxonomy. Tools for identifying biodiversity: progress and problems</article-title>
          .
          <source>Proceedings of the International Congress, Paris. September 20-22</source>
          ,
          <year>2010</year>
          .
          <article-title>Edited by Pier Luigi Nimis and R gine Vignes Lebbe</article-title>
          . p.
          <fpage>54</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Vladimir</surname>
            <given-names>Blagoderov</given-names>
          </string-name>
          , Irina Brake, Teodor Georgiev, Lyubomir Penev, David Roberts,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Ryrcroft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ben</given-names>
            <surname>Scott</surname>
          </string-name>
          , Donat Agosti, Terry Catapano,
          <string-name>
            <surname>Vincent</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Streamlining taxonomic publication: a working example with Scratchpads and ZooKeys</article-title>
          . Zookeys.
          <year>2010</year>
          ;
          <volume>50</volume>
          :
          <fpage>1728</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Drupal</surname>
          </string-name>
          , http://drupal.org
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>20. MediaWiki, http://www.mediawiki.org</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>