<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Nanopublications to Incentivize the Semantic Exposure of Life Science Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mark Thompson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik A. Schultes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leiden University Medical Center</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The growing rate of data production in the life sciences creates an urgent need for semantic integration of information. Although the development of tools and infrastructure will make semantic data exposure easier with time, presently the e ort associated with creating linked data remains largely unrecognized by peer-review processes, publishers, and promotion committees. Here, we describe a novel data publishing framework called nanopublications that provides incentives for researchers to expose their data in semantic form. A nanopublication is the smallest unit of publishable information and is composed of an assertion (a semantic triple subject-predicate-object combination) and provenance metadata such as personal and institutional attribution (which also uses triples). As RDF named graphs, nanopublications are fully interoperable and machine readable, and need not be tethered to centralized databases, research articles or other schema for their retrieval and use. Hence, individual nanopublications can be cited and their impact tracked, creating powerful incentives for compliance with open standards and driving data interoperability.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Large, harmonized datasets, especially from heterogeneous sources, promise to
accelerate discovery in the life sciences, and o er new approaches to managing
intrinsically complex biomedical systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, the channels for data
consumption have not scaled with data production leading to the loss of valuable
data from scienti c discourse [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Given the magnitude and diversity of data
production in the life sciences, the identi cation of trends and the inference of
novel and relevant associations demands automated approaches to analysis and
reasoning. In turn, this requires the automatic and universal interoperability of
data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Semantic technologies have emerged that e ectively address these
issues, but the legacies of scholarly communication continue to preempt e orts of
data integration [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The narrative research article has been, for over a century,
the accepted unit of attribution and scienti c productivity. This made sense
when typical datasets were small enough to be included in the research article
itself (as tables or gures). However, as data production becomes increasingly
automated, large-scale datasets must necessarily be hosted independently of the
research article [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. In response, a dynamic ecosystem of technological
solutions to large-scale data deposition, archival, persistence, licensing, access and
attribution has emerged1. Yet, no consensus around data representation,
protocols for data linking or citation has crystallized among the research community
or publishers. Thus, the lack of data interoperability continues to persist as a
sociological, rather than as a technological problem.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>What is a nanopublication?</title>
      <p>
        Since 2009, the Dutch BioSemantics Group [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] has been developing a data format
standard that scales with the demands of Big Data. This standard, called
nanopublications [
        <xref ref-type="bibr" rid="ref5 ref8">5, 8, 9</xref>
        ], attaches to individual datum provenance metadata such that
data no longer need to be tethered to centralized databases, research articles
or other schema for their retrieval and use. Furthermore, exploiting o -the-shelf
semantic technology, nanopublications are fully interoperable and machine
readable. Hence, individual nanopublications can be cited and their impact tracked,
creating incentives for individuals and institutions to exchange appropriate data.
      </p>
      <p>Nanopublication packages individual datum as citable, stand-alone
publications using semantic representations. Nanopublication is a schema on top of
existing semantic technology using controlled vocabularies and ontologies. A
nanopublication has two parts: the assertion (datum) and provenance
(metadata). The assertion and provenance are RDF named graphs composed of
semantic triples (subject-predicate-object combinations) [10]. The assertion describes
a minimal unit of actionable scienti c information such as a controlled
observation (from the eld or the laboratory) or a simple hypothesis (that can later be
tested). The provenance describes how the assertion came to be, and includes
both supporting information (e.g., context, parameter settings, a description of
methods) and attribution information including ne-grained acknowledgment
of institutions supporting the work, funding sources and other information like
date and time stamps and certi cation [11{13]. A nanopublication represents
the smallest unit of actionable information and combines both the technical
solution for interoperability (semantic web representations) with the incentives
(attribution) as a single publishable unit.
3</p>
    </sec>
    <sec id="sec-3">
      <title>How to use nanopublications?</title>
      <p>Creating a nanopublication requires a one-time e ort to model the scienti c
assertion and provenance as RDF named graphs. After submission to an open,
decentralized nanopublication store (essentially a triple store), nanopublications
will be available as both human-readable and machine-readable information and
will be fully interoperable under semantic queries and to automated inferencing
engines. Nanopublications can be used to expose any data type whatsoever,
including quantitative and qualitative data, experimental data as well as
hypotheses, novel or legacy data and even negative results that usually go unpublished.
1 Some examples are available at www.datadryad.org, www.foaf-project.org,
www.thedatahub.org, www.datamarket.com, www.thedata.org and
www.gigasciencejournal.com
The nanopublication framework can be used to expose data streams from
curated databases as well as from instrumentation (sensors) and communication
sources (internet transactions, email, video, click streams, or other digital sources
available today and in the future). As a data publishing framework,
nanopublications are meant to augment (not replace) traditional narrative research articles,
although nanopublications can be used to expose individual assertions from
narrative text.</p>
      <p>By linking assertions and provenance using semantic representations, not only
do data become interoperable, but their value can be independently estimated.
Nanopublications provide a common currency for the exchange of data and thus
allow crowd sourced or market-driven assignment of value to individual datum
[4, 5, 14{17]. This is in contrast to traditional peer-review which has not scaled
with the demands of data production and increasingly shows signs of bias and
failure [18{22]. Based on this estimated value, nanopublications can be ltered
and prioritized for the purposes of search and inclusion in automated inferencing
algorithms. Large networks of custom nanopublication mash-ups from diverse
sources can be constructed and searched for novel (implied) associations that
would otherwise escape the human reasoning. Indeed, newly discovered
associations can themselves be represented and shared as nanopublications. In turn,
the value of individual datum can be translated into citation metrics, measures
of scienti c impact and other professional and economic indicators incentivizing
interoperability and sharing [14].
9. http://www.nanopub.org
10. Klyne, G., Carroll, J.J.: Resource Description Framework (RDF): Concepts and</p>
      <p>Abstract Syntax. Technical report
11. Taylor, C.F., Field, D., Sansone, S.A., Aerts, J., Apweiler, R., Ashburner, M., Ball,
C.A., Binz, P.A., Bogue, M., Booth, T., Brazma, A., Brinkman, R.R., Michael
Clark, A., Deutsch, E.W., Fiehn, O., Fostel, J., Ghazal, P., Gibson, F., Gray, T.,
Grimes, G., Hancock, J.M., Hardy, N.W., Hermjakob, H., Julian, R.K., Kane, M.,
Kettner, C., Kinsinger, C., Kolker, E., Kuiper, M., Novere, N.L., Leebens-Mack,
J., Lewis, S.E., Lord, P., Mallon, A.M., Marthandan, N., Masuya, H., McNally, R.,
Mehrle, A., Morrison, N., Orchard, S., Quackenbush, J., Reecy, J.M., Robertson,
D.G., Rocca-Serra, P., Rodriguez, H., Rosenfelder, H., Santoyo-Lopez, J.,
Scheuermann, R.H., Schober, D., Smith, B., Snape, J., Stoeckert, C.J., Tipton, K., Sterk,
P., Untergasser, A., Vandesompele, J., Wiemann, S.: Promoting coherent
minimum reporting guidelines for biological and biomedical investigations: the MIBBI
project. Nat Biotech 26(8) (August 2008) 889{896
12. http://isa-tools.org
13. http://www.w3.org/2004/02/skos
14. Bartolini, C., Vukovic, M.: Crowdsourcing human mutations. (2011)
15. Mons, B., Ashburner, M., Chichester, C., van Mulligen, E., Weeber, M., den
Dunnen, J., van Ommen, G.J., Musen, M., Cockerill, M., Hermjakob, H., Mons, A.,
Packer, A., Pacheco, R., Lewis, S., Berkeley, A., Melton, W., Barris, N., Wales, J.,
Meijssen, G., Moeller, E., Roes, P.J., Borner, K., Bairoch, A.: Calling on a million
minds for community annotation in WikiProteins. Genome biology 9(5) (January
2008) R89
16. Ho mann, R.: A wiki for the life sciences where authorship matters. Nat Genet
40(9) (September 2008) 1047{1051
17. Oprea, T.I., Bologa, C.G., Boyer, S., Curpan, R.F., Glen, R.C., Hopkins, A.L.,
Lipinski, C.A., Marshall, G.R., Martin, Y.C., Ostopovici-Halip, L., Rishton, G.,
Ursu, O., Vaz, R.J., Waller, C., Waldmann, H., Sklar, L.A.: A crowdsourcing
evaluation of the NIH chemical probes. Nature Chemical Biology 5(7) (2009)
441{447
18. Miller, A., Barwell, G.: Science and Technology Peer review in scienti c
publications Eighth Report of Session 201012. (July) (2011)
19. Mullard, A.: Reliability of 'new drug target' claims called into question. Nat Rev</p>
      <p>Drug Discov 10(9) (September 2011) 643{644
20. Prinz, F., Schlange, T., Asadullah, K.: Believe it or not: how much can we rely on
published data on potential drug targets? Nature reviews. Drug discovery 10(9)
(September 2011) 712
21. Booth, B.: Academic Bias &amp; Biotech Failures. http://lifescivc.com/2011/03/
academic-bias-biotech-failures/ (2011)
22. Fanelli, D.: How Many Scientists Fabricate and Falsify Research? A Systematic
Review and Meta-Analysis of Survey Data. PLoS ONE 4(5) (2009) e5738</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Wild</surname>
            ,
            <given-names>D.J.:</given-names>
          </string-name>
          <article-title>Mining large heterogeneous data sets in drug discovery</article-title>
          .
          <source>Expert Opinion on Drug Discovery</source>
          <volume>4</volume>
          (
          <issue>10</issue>
          ) (
          <year>2009</year>
          )
          <volume>995</volume>
          {
          <fpage>1004</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. http://www.nlm.nih.gov/bsd/medline_cit_counts_yr_pub.html</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Sansone</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocca-Serra</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Field</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maguire</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , C.,
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amaral-Zettler</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Begley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Booth</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bougueleret</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burns</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coleman</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Copeland</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Daruvar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Matos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dix</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edmunds</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evelo</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forster</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaudet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gilbert</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gri n</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacob</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinjans</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harland</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haug</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hermjakob</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sui</surname>
            ,
            <given-names>S.J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laederach</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGrath</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merrill</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reilly</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shamu</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shang</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinbeck</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trefethen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams-Jones</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolstencroft</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xenarios</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hide</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Toward interoperable bioscience data</article-title>
          .
          <source>Nature Genetics</source>
          <volume>44</volume>
          (
          <issue>2</issue>
          ) (
          <year>2012</year>
          )
          <volume>121</volume>
          {
          <fpage>126</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Marx</surname>
          </string-name>
          , V.:
          <article-title>My data are your data</article-title>
          .
          <source>Nat Biotech</source>
          <volume>30</volume>
          (
          <issue>6</issue>
          ) (
          <year>June 2012</year>
          )
          <volume>509</volume>
          {
          <fpage>511</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mons</surname>
            , B., van Haagen,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoen</surname>
          </string-name>
          , P.B.t., den Dunnen, J.T., van Ommen, G., van
          <string-name>
            <surname>Mulligen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hooft</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hammond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giardine</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velterop</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The value of data</article-title>
          .
          <source>Nat Genet</source>
          <volume>43</volume>
          (
          <issue>4</issue>
          ) (
          <year>April 2011</year>
          )
          <volume>281</volume>
          {
          <fpage>283</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. World Economic Forum:
          <article-title>Big Data , Big Impact : New Possibilities for International Development</article-title>
          .
          <source>Agenda</source>
          (
          <year>2012</year>
          ) 0{
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. http://www.biosemantics.org</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velterop</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The anatomy of a nanopublication</article-title>
          .
          <source>Information Services and Use</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ) (
          <year>2010</year>
          )
          <volume>51</volume>
          {
          <fpage>56</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>