<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>nanopub-java: A Java Library for Nanopublications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tobias Kuhn</string-name>
          <email>kuhntobias@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, VU University Amsterdam</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Humanities, Social and Political Sciences</institution>
          ,
          <addr-line>ETH Zurich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The concept of nanopublications was rst proposed about six years ago, but it lacked openly available implementations. The library presented here is the rst one that has become an o cial implementation of the nanopublication community. Its core features are stable, but it also contains uno cial and experimental extensions: for publishing to a decentralized server network, for de ning sets of nanopublications with indexes, for informal assertions, and for digitally signing nanopublications. Most of the features of the library can also be accessed via an online validator interface.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>This technical paper presents nanopub-java, which is a library for
nanopublications. Its source code can be found here:</p>
      <p>
        https://github:com/Nanopublication/nanopub-java
Nanopublications3 [
        <xref ref-type="bibr" rid="ref2 ref8">2,8</xref>
        ] are an approach to publish scienti c data and
metadata in RDF by subdividing them into small data snippets. They are a concrete
proposal to implement the visions of semantic publishing [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and linked science
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] by allowing for the publication and sharing of formally represented scienti c
resources and data that are semantically interlinked and provide provenance and
context information for their reliable integration and evaluation. Speci cally, a
nanopublication consists of three named graphs of RDF triples (plus a fourth
graph to keep them together): the assertion graph contains the actual content
of the nanopublication (e.g. a scienti c nding); the provenance graph contains
information about the provenance of the assertion (e.g. the scienti c method with
which the assertion was derived); and the publication information graph contains
meta-data about the nanopublication itself (e.g. its creator and a timestamp).
      </p>
      <p>The library presented here can be useful in a number of scenarios:
{ To represent and share small chunks of scienti c knowledge and metadata
in RDF in a provenance-aware manner (as nanopublications)
{ To make RDF content veri able and immutable (with trusty URIs)
3 http://nanopub:org
{ To de ne large or small datasets of RDF content where the data entries can
be individually addressed and recombined in new datasets (with
nanopublication indexes)
{ To quickly publish RDF snippets in a veri able and permanent manner
(relying on an existing server network)
{ To retrieve existing nanopublications from the network (5 millions and
counting)
{ To digitally sign RDF snippets (though this is still experimental)
Below the details of the library and its web interface are explained.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Implementation</title>
      <p>
        The library is built upon the Sesame library [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to validate, represent, and create
RDF structures. The features of the nanopublication library are centered around
a Java interface representing a nanopublication, called Nanopub, and the Java
class NanopubImpl provides a reference implementation of this interface. This
implementation checks the well-formedness of a nanopublication at the time of
its creation based on the latest version of the nanopublication guidelines,4 and
raises an exception in the case of a violation of these rules.
      </p>
      <p>
        Trusty URIs [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ] are the recommended way of how to make
nanopublications veri able and immutable, and to give them identi ers based on
cryptographic hash values. The nanopublication library uses for that purpose the
trustyuri-java library.5 In a nutshell, a trusty URI is a kind of URI reference
that contains a cryptographic hash value that is calculated on the digital
artifact it represents. This allows one to verify that a given content is really what
the URI was supposed to represent by its creator, and thereby to enforce the
immutability of digital artifacts such as nanopublications.
      </p>
      <p>The features of the library are made available through the Java API as well
as via a command line interface using the command np. The following features
are part of the core of the library, which means that they deal with stable and
agreed-upon structures as de ned by the community:
{ check / CheckNanopub reads a nanopublication or several of them and
checks whether any of the well-formedness criteria are violated. If a trusty
URI or a digital signature is found (see below), these are checked too.
{ mktrusty / MakeTrustyNanopub takes a nanopublication that does not yet
have a trusty URI and transforms it into one that is identi ed by a newly
created trusty URI.
{ fix / FixTrustyNanopub takes a nanopublication with a broken trusty URI
and xes it, i.e. assigns it a new trusty URI. This is useful when a
nanopublication has to be changed, which invalidates the hash. Running this command
creates a new nanopublication with a valid trusty URI. (Nanopublications</p>
      <sec id="sec-2-1">
        <title>4 http://nanopub:org/guidelines/working draft/ 5 https://github:com/trustyuri/trustyuri-java</title>
        <p>are immutable, so changing something necessarily leads to a new
nanopublication.)</p>
        <p>
          In addition to these core features, the library also contains a number of
uno cial extensions (which may or may not become o cial at some point). There
is code to validate informal assertions speci ed as AIDA sentences [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]; code for
creating index nanopublications to de ne small or large sets of nanopublications
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]; code for publishing and retrieving nanopublications from a decentralized
server network [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]; and experimental code for digitally signing nanopublications.
These features are accessible via the following commands and classes:
{ mkindex / MakeIndex takes a list of nanopublications and creates an index
that refers to them. A nanopublication index therefore represents a (possibly
large) set of nanopublications. Such indexes are themselves formatted as
nanopublications, and can therefore also be published to the server network
(see below).
{ publish / PublishNanopub uploads a given nanopublication that has a
trusty URIs to the server network. Such a nanopublication is then distributed
among the servers of the network (currently ve) and made available even if
some of the servers should be inaccessible at a certain point in time. In this
way, the nanopublication is made permanent and its publication cannot be
undone.
{ get / GetNanopub reliably retrieves a given nanopublication from the
decentralized server network. Nanopublications are veri ed according to their
trusty URI, and only veri ed nanopublications are returned by this
command. For nanopublication indexes, the whole set of nanopublications that
is de ned by the index can be downloaded.
{ status / NanopubStatus checks whether and how often a given
nanopublication (identi ed by its trusty URI) is found on the server network.
{ server / GetServerInfo returns some information about a given server in
the network, such as the number of nanopublications it contains.
{ mkkeys / MakeKeys creates a new key-pair to be used to sign
nanopublications.
{ sign / SignNanopub takes a nanopublication and signs it with a given
private key.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Examples</title>
      <p>Below, some of the most important commands are explained by examples based
on the np command line tool. The same functionality is also available via the
Java API. For the sake of these examples, let us assume that we have a le called
nanopubfile.trig that starts with the following RDF pre xes:
@prefix xsd: &lt;http://www.w3.org/2001/XMLSchema#&gt;.
@prefix dc: &lt;http://purl.org/dc/terms/&gt;.
@prefix pav: &lt;http://purl.org/pav/&gt;.
@prefix prov: &lt;http://www.w3.org/ns/prov#&gt;.
@prefix np: &lt;http://www.nanopub.org/nschema#&gt;.
@prefix ex: &lt;http://example.org/&gt;.
@prefix : &lt;http://example.org/np1#&gt;.
:assertion {</p>
      <p>ex:drugA ex:treats ex:diseaseB.
:provenance {</p>
      <p>:assertion prov:wasDerivedFrom ex:some_publication.
}
:pubinfo {
: pav:createdBy &lt;http://orcid.org/0000-0002-1267-0234&gt;.</p>
      <p>: dc:created "2015-08-18T15:36:22+01:00"^^xsd:dateTime.</p>
      <p>}
The provenance and publication info graphs provide meta-information about the
assertion and the entire nanopublication, respectively:
The lines above constitute a very simple but complete nanopublication. To make
this example a bit more interesting, let us assume that our le contains two more
nanopublications that have di erent assertions but are otherwise identical:
The de nition of the rst nanopublication in this le starts with the head graph
that de nes the structure of the nanopublication by linking to the other graphs:
:Head {
: a np:Nanopublication; np:hasAssertion :assertion;</p>
      <p>np:hasProvenance :provenance; np:hasPublicationInfo :pubinfo.</p>
      <p>The actual claim of the nanopublication is stored in the assertion graph:
}
}
@prefix : &lt;http://example.org/np2#&gt;.
...
:assertion {</p>
      <p>ex:Gene1 ex:isRelatedTo ex:diseaseB.
@prefix : &lt;http://example.org/np3#&gt;.
...
:assertion {</p>
      <p>ex:Gene2 ex:isRelatedTo ex:diseaseB.</p>
      <p>
        To check and validate these three nanopublications, we can now use the following
command:
$ np check nanopubfile.trig
Summary: 3 valid (not trusty);
$ np mktrusty nanopubfile.trig
These nanopublications can now be transformed into ones with trusty URIs using
the following command (resulting in a new le trusty.nanopubfile.trig):
Using the same command in verbose mode with the argument -v shows us the
newly generated trusty URIs for the three nanopublications:
$ np mktrusty -v nanopubfile.trig
Nanopub URI: http://example.org/np1#RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7M91ouhojsQ
Nanopub URI: http://example.org/np2#RA4xTdhe2gPctqvAwdgTU4eRiR1aTQlTYJcF3Sohe5Cus
Nanopub URI: http://example.org/np3#RAEjvXP0xTkeIa2mKmYT66i_PAJ-u-k0uRBd6_sMe9qG0
As they are tiny snippets of data, nanopublications are most useful when they
grouped and combined in small or large collections. We therefore need a simple
method to refer to collections or sets of nanopublications, which is achieved by
the experimental proposal of nanopublication indexes [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which are themselves
nanopublications. Such indexes can be used to group nanopublications that have
trusty URIs using the following command:
$ np mkindex trusty.nanopubfile.trig
      </p>
      <p>Index URI: http://np.inn.ac/RAFa_x4h0ng_NXtof35Ie9pQVsAY69Ab3ZQMir2NP8vGc
The nanopublications of the new index are saved in a le called index.trig
unless speci ed otherwise with the argument -o.</p>
      <p>Moving to the part that involves the server network, nanopublications that
have trusty URIs (which includes nanopublication indexes) can be published to
the network with the following command:
$ np publish trusty.nanopubfile.trig
3 nanopubs published at http://np.inn.ac/
The publication status of a given nanopublication can be checked like this:
$ np status -a http://example.org/np1#RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7M91ouhojsQ
URL: http://np.inn.ac/RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7M91ouhojsQ
URL: http://ristretto.med.yale.edu:8080/nanopub-server/RAHGB0WzgQijR88g_rIwtPCmzYgyO...
URL: http://nanopub-server.ops.labs.vu.nl/RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7M91ouhojsQ
URL: http://nanopubs.stanford.edu/nanopub-server/RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7...
URL: http://nanopubs.semanticscience.org/RAHGB0WzgQijR88g_rIwtPCmzYgyO4wRMT7M91ouhojsQ
Found on 5 nanopub servers.</p>
      <p>A given nanopublication that is published on the server network can be retrieved
via its URI:</p>
      <p>$ np get http://www.tkuhn.ch/bel2nanopub/RAhV9IpiUEjbentzGivp1Lbx0BVegp5sgE3BwS0S2RAYM
All the servers in the network are checked until the nanopublication is found and
successfully veri ed. This command is therefore reliable even if one or several
servers are down. Instead of the complete URI, it is also possible to just specify
the trusty URI artifact code:</p>
      <p>$ np get RAhV9IpiUEjbentzGivp1Lbx0BVegp5sgE3BwS0S2RAYM
To get the content of a nanopublication index (and not just the top-most index
nanopublication), argument -c can be used:</p>
      <p>$ np get -c -o content.trig RAtF0ivB9B8cb-u3K_zElgmRBxiDwfym1yVBRY6VAyWvE
Argument -o speci es again the name of the output le. The remaining
commands as introduced above are equally intuitive to use. Just entering the
command without any arguments will output a list of all argument options.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Web Interface</title>
      <p>Many of the features described above are made available through the
nanopublication validator Web interface,6 an instance of which can be accessed at
http://nanopub:inn:ac. Figure 1 shows a screenshot. With this interface, the
wellformedness of nanopublications can be checked as well as the adherance to a</p>
      <sec id="sec-4-1">
        <title>6 https://github:com/tkuhn/nanopub-validator</title>
        <p>number of patterns. They can furthermore be transformed into di erent RDF
serializations, and published to the server network. Loading of nanopublications
is possible via form input, upload, fetching from a URL, SPARQL endpoint
access, and retrieval from the server network. In general, this web interface and
the underlying library are supposed to support the development of best practices
for the nanopublication community by providing a solid basis for discussion, by
allowing for experimental features to be tested and discussed, and by facilitating
the implementation of prototypes.</p>
        <p>The current server network on which many of the uno cial features depend,
can be explored via a monitor interface at http://npmonitor:inn:ac.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The nanopub-java library provides a stable implementation of the
nanopublication concept, adhering to its speci ed guidelines. It is openly licensed under
the terms of the MIT license, is available on The Central Repository,7 and has
so far been used in about a dozen open-source codebases.8</p>
      <p>In general, we believe that this library can be a valuable resource for tools
that use RDF data in the context of provenance recording, reproducibility, data
publishing, data reuse, and reliable retrieval of Linked Data.
7 https://search:maven:org/#artifactdetailsjorg:nanopubjnanopubj1:7jjar
8 https://github:com/Nanopublication/nanopub-java#usage-tracking</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Broekstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kampman</surname>
          </string-name>
          , and
          <string-name>
            <surname>F. Van Harmelen. Sesame:</surname>
          </string-name>
          <article-title>A generic architecture for storing and querying rdf and rdf schema</article-title>
          .
          <source>In The Semantic Web | ISWC</source>
          <year>2002</year>
          , pages
          <fpage>54</fpage>
          {
          <fpage>68</fpage>
          . Springer,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>P.</given-names>
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gibson</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Velterop.</surname>
          </string-name>
          <article-title>The anatomy of a nano-publication</article-title>
          .
          <source>Information Services and Use</source>
          ,
          <volume>30</volume>
          (
          <issue>1</issue>
          ):
          <volume>51</volume>
          {
          <fpage>56</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kauppinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baglatzi</surname>
          </string-name>
          , and
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Ke ler. Linked science: interconnecting scienti c assets</article-title>
          .
          <source>Data Intensive Science</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Barbano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Nagy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Krauthammer</surname>
          </string-name>
          .
          <article-title>Broadening the scope of nanopublications</article-title>
          .
          <source>In The Semantic Web: Semantics and Big Data | ESWC</source>
          <year>2013</year>
          , pages
          <fpage>487</fpage>
          {
          <fpage>501</fpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chichester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krauthammer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          .
          <article-title>Publishing without publishers: a decentralized approach to dissemination, retrieval, and archiving of data</article-title>
          .
          <source>In The Semantic Web | ISWC 2015</source>
          . Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          . Trusty URIs:
          <article-title>Veri able, immutable, and permanent digital artifacts for linked data</article-title>
          .
          <source>In The Semantic Web: Trends and Challenges | ESWC</source>
          <year>2014</year>
          , pages
          <fpage>395</fpage>
          {
          <fpage>410</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>T.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          .
          <article-title>Making digital artifacts on the web veri able and reliable</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>27</volume>
          (
          <issue>9</issue>
          ),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>B.</given-names>
            <surname>Mons</surname>
          </string-name>
          , H. van Haagen,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chichester</surname>
          </string-name>
          , J. T. den Dunnen, G. van Ommen,
          <string-name>
            <surname>E. van Mulligen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hooft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Roos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hammond</surname>
          </string-name>
          , et al.
          <article-title>The value of data</article-title>
          .
          <source>Nature genetics</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ):
          <volume>281</volume>
          {
          <fpage>283</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Shotton</surname>
          </string-name>
          .
          <article-title>Semantic publishing: the coming revolution in scienti c journal publishing</article-title>
          .
          <source>Learned Publishing</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ):
          <volume>85</volume>
          {
          <fpage>94</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>