<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Placing landMarks in the Knowledge Space: crowd-sourcing landmark publications for benchmarking text-mined predictions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mark Thompson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Herman H.H.B.M. van Haagen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barend Mons</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik A. Schultes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leiden University Medical Center</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>It is generally perceived that text-mining systems have failed to deliver on the promise of predicting novel, meaningful relationships between biomedical concepts. Despite successes where novel relationships have been inferred and later con rmed by laboratory experiments, there are many more cases where text-mining did not predict the outcome of high-throughput experiments or population-based genetic studies. Here, we show that this apparent incongruity between text-mined predictions and experimental data results not from a failure of text-mining in principle, but rather, from the confounding of 4 distinct classes of data typically used in this research. Keeping this distinction in mind, and using a novel, crowd-sourced and crowd-curated test set of (among others) protein-protein interactions, we propose a more discriminating standard for the evaluation of text-mined predictions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Biological systems are composed of interactions between millions of components:
genes, regulatory elements, RNAs, proteins, metabolites, nutrients and drug
compounds which give rise to associated functions, healthy phenotypes and
disease states that dynamically emerge over the life cycle of the organism. Although
high-throughput methods routinely screen for associations between these
components, the very large scale of these datasets precludes their analysis except by
automated means [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In the last decade many text-mining systems have been
developed to assist biologists in nding new associations in large and
heterogeneous data. Typically, text-mining systems have two goals: (1) to annotate a
set of genes with literature-based information, or (2) to infer new associations
between concepts (e.g. a novel gene-disease relationships or a protein-protein
interactions) that have never before been explicitly stated in literature or recorded
in databases. For example, text-mining results were used to predict a relationship
between sh oil and Raynauds syndrome [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the physical interaction between
the proteins CAPN3 and PARVB [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and a novel gene involved in craniofacial
development from a 2-Mb chromosomal region, deleted in some patients with
      </p>
      <p>
        Authors contributed equally to this work
DiGeorge-like birth defects [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Tune Pers et al. used ve di erent sources to
annotate results from a genome-wide association study and found the causative
gene YWHAH for bipolar disorder [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Although these successes demonstrate the potential of exposing novel
associations from existing biomedical texts, there are also many examples where
text-mining was not able to predict experimental ndings from microarrays,
GWAS or other large-scale analyses. However, it is hard to evaluate the extent
of such failures of text-mining as these cases, viewed as negative results, are not
generally publishable. In any case, there is a growing consensus that text-mining
is unreliable, and has not delivered on its promise of automated knowledge
discovery [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
      </p>
      <p>Here, we show that the perceived incongruities between text-mined
predictions and laboratory studies often re ect confusion at a fundamental level about
what text-mining is doing and what text-mined inferences actually represent.
Essentially, text-mining exposes knowledge that is already there in the
knowledge store (but has yet to be recognized by researchers), while experimental
approaches can (and often do) establish novel associations that have no antecedents
whatsoever in existing knowledge stores. As such text-mined predictions are, in
general, not comparable to independent laboratory data and in such cases we
should expect little, if any signi cant overlap between the two. This does not
mean that text-mining systems can not be rigorously evaluated. To the contrary,
the performance of text-mining systems can be very accurately assessed but only
by directly testing the predictions in the laboratory.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Rede ning the knowledge space</title>
      <p>To help clarify these relations, we partition the knowledge space of potential
associations by evidence derived from text-mining analyses and laboratory
experiments (Figure 1). The evidence in both cases can be positive or negative,
creating four types of conceptual associations.</p>
      <p>Type I associations (top left) are cases where both the literature and
experiments have provided con rmatory evidence for the association, and therefore
represents well-established knowledge (Explicit Knowledge). Indeed, sometimes
multiple independent lines of evidence con rm a particular nding making it
more reliable. Typically literature is based on experimental evidence (e.g. a
publication describing the experiment) so that text-mined Type I associations are
often a re-discovery of what is already known, and in this way Type I
associations provide con rmation that the text-mining method is working as intended.
Although Type I associations enjoy consensus, they are not novel or surprising.
An example of a Type I association would be a high association score between
the gene huntingtin and Hungtintons disease.</p>
      <p>
        Type IV associations (Figure 1, right bottom) is the Negatome, those
associations that have no evidence supporting them, whether they have been explicitly
tested, or not. For example, a microarray experiment concludes that two genes
are not di erentially expressed, or that a SNP from GWAS is not signi cant. In
the text-mining case, a negative result may re ect a failure of the text-mining
system, or simply that there is insu cient information in the literature to
establish a signi cant association (a condition we call a Knowledge Vacuum). In any
case, like Type I associations, there is a consensus between text-mining and
experiments. Type IV associations are by far the largest class of associations and
are often treated as a null set of randomly chosen concept pairs in statistical
analysis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The remaining associations, Types II and III, are characterized by con icting
results between experiments and text-mining. Type II and III associations are
often confounded leading to confusion in the interpretation of text-mining results
and erroneous conclusions about text-mining performance.</p>
      <p>Type III associations (Figure 1, bottom left) is the case where text-mining
can be most e ectively used in knowledge discovery. Here, text-mining results
predict novel associations that have yet to be tested experimentally, or have been
tested but with negative results. In the former case, the predicted associations
are treated as hypotheses to be tested, which is the ultimate goal of text-mining.
In the latter case, as negative experimental results are always ambiguous, the
positive text-mining results can be used as leads looking for associations under
alliterative conditions. In either case, Type III associations can be viewed as
a prioritized list of Hypotheses guiding the next-step decisions of experimental
researchers.</p>
      <p>In the case of Type II associations (Figure 1, top right), ndings based on
experimental evidence are not supported by text-mining, yielding a
Contradiction. Many Type II associations come from high-throughput screens or GWAS
and are de novo discoveries such that no literature-based information is yet
available. Although it is always a possibility that a text-mining method may
simply be returning false negatives, failure to predict a positive experimental
result could also re ect a text-mining Knowledge Vacuum. In any case, Type
II associations necessitate further inquiry and possible trouble-shooting of the
text-mining system.</p>
      <p>Given the large number of biomedical concepts and their potential pair-wise
associations, the Knowledge Vacuum is likely to be a large fraction of the
Knowledge Space, that is, it is likely that the vast majority of associations have yet
to be represented either explicitly or implicitly in the literature. For example,
there are about 25,000 human genes yet only 12,000 of these entities have more
than 5 PubMed abstracts, making them visible to text-mining systems. Hence,
literature-based knowledge discovery is inherently limited to concepts that have
been well-published upon, and can not be used to predict associations between
concepts that have yet to be discussed in the literature. Although a large
number of associations can, and should be mined, they should not be compared
directly with experiments that test, de novo, a much wider class of associations.
Hence, Type II associations that involve high-throughput experimental screens
should not be viewed as a failure of text-mining, but rather text-mining and
high-throughput experiments should be seen as complementary approaches to
mapping the space of possible associations.
3</p>
    </sec>
    <sec id="sec-3">
      <title>An alternative evaluation method</title>
      <p>A more relevant evaluation of text-mining systems can be based directly on the
text-mined predictions themselves. We propose the use of retrospective analyses
that use benchmark sets of known associations that takes into consideration the
taxonomy of potential associations as shown in Fig. 1. In particular, the
benchmark datasets makes a distinction between associations that are (or can be)
inferred from the explicitome (Type I and III), and those that are not (Type II
and IV). We then perform a retrospective analysis using only the predicted
associations (Type III) until a certain date and evaluate the prediction by comparing
the result against the consensus knowledge after that date (Type I). In this way
the evaluation method will not discredit a text-mining result that fails to
predict relationships that are inherently unknowable due to a lack of information
available in literature.</p>
      <p>The key problem is to identify the set of benchmark associations that can
be inferred from the explicitome. Benchmark data sets have to de ne very
precisely the concepts that make up the association and the date of rst publication
of the association. We note some particular problems with trying to
automatically generate such a benchmark, for example, using automatically retrieved
rst co-occurrences. Although rst co-occurrences of terms can be determined
automatically, mapping those terms to concepts can still not be done with
complete accuracy. Moreover, a nding may have been reported rst in a publication
that was not in the co-occurrence data set or it may have been reported as a
hypothetical relationship, thus co-occurring before it is presented together with
any kind of evidence. For such reasons we propose building a curated benchmark
by means of crowd-sourcing.</p>
      <p>We have developed landMark a landmark publication crowd sourcing tool.
A landmark publication refers to the rst occurrence in literature of an
association between concepts for which experimental evidence is given. It has been
developed to allow easy and accurate registration of curated landmarks in a form
that we refer to as the "landmark claim": Article X is the rst to show a
relationship between concept A and concept B The target audience for this tool are
publication authors (who register their own landmark ndings) as well as, for
example, curators who may register landmark ndings on behalf of the authors.
From a su ciently large set of such landmark claims we will be able to derive
high quality curated benchmark test sets.</p>
      <p>These benchmark sets will be made publicly available as a valuable resource
for text-mining and knowledge discovery researchers worldwide. For the purpose
of simplicity we initially limit ourselves to protein-protein interactions,
genedisease relationships and drug-disease relationships. We think this represents an
important subset of landmark ndings while making for interesting targets in
the current state-of-the-art of knowledge discovery.</p>
      <p>
        The landMark interface presents the user with a very concise web-form that
helps the user to quickly and unambiguously enter a landmark claim.
Disambiguation is achieved in two steps: rst the user selects one of the three
relationship categories from the accordion widget, and types a term describing each of the
concepts. As the user types, an auto-complete feature queries the ConceptWiki
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] for concepts (within the selected category) that match the (partially) typed
term. In case multiple concepts match the term, the user can review their
ConceptWiki summaries to manually disambiguate them. The interface also has
elds asking the user to provide a DOI or PubMedID of the landmark paper,
its publication date and the rst author and his institutional a liation. The
optional curator eld identi es the curator making the claim on behalf of the
author. Two nal questions are asked to identify the type of discovery, from
which we can infer the Type of the landmark and thereby the suitability of this
landmark towards the evaluation of prediction mechanisms as discussed in the
previous section.
      </p>
      <p>A nal processing step is required to transform the entire collection of
landmark claims into the required benchmark set. For example, consider the
situation where an author challenges a previous claim by submitting a new claim
that refers to an earlier article. Due to the careful and unambiguous selection
of concepts, we can later identify whether two claims refer to the same concepts
and include only the information from the claim that refers to the earliest paper
in the nal benchmark test set.</p>
      <p>As with all curation e orts, the quality of the benchmark test set will depend
on the quality and the amount of contributions. We hope to incentivize authors
to make their landmark claims by o ering to turn their landmark claim in
nanopublications [9{12]. A nanopublication is a permanent, immutable, semantic-web
representation of the smallest unit of publishable information that consists of an
assertion and provenance. The landMark web tool will store each landmark claim
as a nanopublication assertion with the author, publication date and additional
information as nanopublication provenance. As the nanopublication becomes
part of the web of linked data, a landmark nanopublication o ers a simple way
for authors to gain attribution for key parts of their published research and for
curators to receive credit for the important (but often underappreciated) e ort
of data curation.</p>
      <p>Currently the landmark nanopublication web application is in an extensive
user testing phase at Leiden University Medical Center. We believe usability is
an important factor in the adoption of this tool. By reducing the e ort required
to submit a claim we make it easy for authors, curators and others to submit
claims and thus help the creation of a high-quality, curated benchmark test sets.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>Text-mining results can be partitioned by experimental evidence and text-mined
evidence. We clari ed that text-mining prediction always has literature as a
starting point and is therefore not particularly suitable for predicting
associations between concepts for which literature has no (or very little) information.
This is often the case for serendipitous ndings of high-throughput experiments,
such as for example, microarray experiments. We propose an alternative method
of evaluation based on a high-quality, curated benchmark data set of landmark
associations in literature. We demonstrated an implementation of a web tool
that will be made available to the community to crowd-source the creation of
such a benchmark set. We hope it will serve as a new and open standard for
text-mining and prediction research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Attwood</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kell</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McDermott</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marsh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thorne</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Calling international rescue: knowledge lost in literature and data landslide!</article-title>
          <source>Biochemical Journal</source>
          <volume>424</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <volume>317</volume>
          {
          <fpage>333</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Swanson</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Fish oil, Raynaud's syndrome, and undiscovered public knowledge</article-title>
          .
          <source>Perspectives in biology and medicine 30(1)</source>
          (
          <year>1986</year>
          )
          <volume>7</volume>
          {
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. van Haagen,
          <string-name>
            <surname>H.H.H.B.M.</surname>
          </string-name>
          , 't Hoen,
          <string-name>
            <given-names>P.A.C.</given-names>
            ,
            <surname>Botelho</surname>
          </string-name>
          <string-name>
            <given-names>Bovo</given-names>
            , A.,
            <surname>de Morree</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>van Mulligen</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          , den Dunnen, J.T., van Ommen,
          <string-name>
            <given-names>G.J.B.</given-names>
            ,
            <surname>van der Maarel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.M.</given-names>
            ,
            <surname>Kern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.M.</given-names>
            ,
            <surname>Mons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Schuemie</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.J.</surname>
          </string-name>
          :
          <article-title>Novel protein-protein interactions inferred from literature context</article-title>
          .
          <source>PLoS ONE</source>
          <volume>4</volume>
          (
          <issue>11</issue>
          ) (
          <year>11 2009</year>
          ) e7894
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Aerts</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lambrechts</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maity</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Van Loo,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Coessens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>De Smet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Tranchevent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.C.</given-names>
            ,
            <surname>De Moor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Marynen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Carmeliet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Moreau</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Gene prioritization through genomic data fusion</article-title>
          .
          <source>Nat Biotechnol</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ) (May
          <year>2006</year>
          )
          <volume>537</volume>
          {
          <fpage>544</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pers</surname>
            ,
            <given-names>T.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>N.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lage</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koefoed</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dworzynski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flint</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mellerup</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dam</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andreassen</surname>
            ,
            <given-names>O.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Djurovic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melle</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , B rglum, A.D.,
          <string-name>
            <surname>Werge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Purcell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kouskoumvekaki</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Workman</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mors</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Meta-analysis of heterogeneous data sources for genome-scale identi cation of risk genes in complex phenotypes</article-title>
          .
          <source>Genetic Epidemiology</source>
          <volume>35</volume>
          (
          <issue>5</issue>
          ) (
          <year>2011</year>
          )
          <volume>318</volume>
          {
          <fpage>332</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Veuthey</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pillet</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yip</surname>
            ,
            <given-names>Y.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruch</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Text mining for swiss-prot curation: A story of success and failure</article-title>
          . In: Nature Precedings. (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>H.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>Y.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tzong-Han</surname>
            <given-names>Tsai</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.L.</surname>
          </string-name>
          :
          <article-title>New challenges for biological text-mining in the next decade</article-title>
          .
          <source>Journal of Computer Science and Technology</source>
          <volume>25</volume>
          (
          <year>2010</year>
          )
          <volume>169</volume>
          {
          <fpage>179</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. http://www.conceptwiki.org</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velterop</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The anatomy of a nanopublication</article-title>
          .
          <source>Information Services and Use</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ) (01
          <year>2010</year>
          )
          <volume>51</volume>
          {
          <fpage>56</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mons</surname>
            , B., van Haagen,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chichester</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoen</surname>
          </string-name>
          , P.B.t., den Dunnen, J.T., van Ommen, G., van
          <string-name>
            <surname>Mulligen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hooft</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hammond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giardine</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velterop</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The value of data</article-title>
          .
          <source>Nat Genet</source>
          <volume>43</volume>
          (
          <issue>4</issue>
          ) (
          <year>Apr 2011</year>
          )
          <volume>281</volume>
          {
          <fpage>283</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mons</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultes</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Data publishing using nanopublications</article-title>
          .
          <source>In: Tiny Transaction on Computer Science (TinyToCS)</source>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>12. http://www.nanopub.org</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>