<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Nanopublication Framework for Biological Networks using Cytoscape.js</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>James P. McCusker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rui Yan</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kusum Solanki</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Erickson</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cynthia Chang</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michel Dumontier</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jonathan S. Dordick</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deborah L. McGuinness</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>5AM Solutions, Inc</institution>
          ,
          <addr-line>Rockville, MD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Chemical &amp; Biological Engineering, Rensselaer Polytechnic Institute</institution>
          ,
          <addr-line>Troy, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Computer Science</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Rensselaer Polytechnic Institute</institution>
          ,
          <addr-line>Troy, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Stanford University</institution>
          ,
          <addr-line>Stanford, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>90</fpage>
      <lpage>92</lpage>
      <abstract>
        <p>-We leverage semantic technologies and Cytoscape.js to create a provenance-aware, probabilistic analysis platform for systems biology and evaluate its usefulness in discovering links between drugs and diseases. In our efforts to create a systematic approach to discovering new uses for existing drugs, we have developed Repurposing Drugs with Semantics (ReDrugS). ReDrugS is a data curation and publication framework that accepts data from nearly any database containing biological or chemical entity interactions and produces visualizations using Cytoscape.js. A semantic web service API is provided that enables search, traversal, and provides composite probabilities for the resulting graph of biological entities using the SADI web service framework and Nanopublications. We show how associations between a postive control, topiramate, allows us to independently reconstruct a positive control of epilepsy and migraine, and potential consequences on bone health.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>Drug repurposing can often lead to effective new treatments
for diseases. The ReDrugs system we are developing can
assist in this procedure through the integration of multiple
systems biology, pharmacology, disease association, and gene
expression databases into a coherent repository of
individuallysupported assertions that can each be assigned their own
probabilistic value. We have developed an initial database that
includes drug/protein, protein/protein, and protein/biological
process associations that is providing us a view into how drugs
have the effects that they do.</p>
    </sec>
    <sec id="sec-2">
      <title>II. METHODS</title>
      <p>
        We deployed an instance of the RPI semantic web
toolsuite, Prizms, at http://redrugs.tw.rpi.edu and the
Comprehensive Knowledge Archive Network (CKAN) to
http://data.melagrid.org to catalog the available datasets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Cataloging is an ongoing process, but initial datasets were
added to the catalog, initializing the Prizms conversion
process. We were then able to use the Prizms infrastructure
to generate RDF for publication to our SPARQL endpoint.
We used the BigData RDF store with named graph and text
indexing support enabled.
      </p>
      <p>A. Inferring Probabilities</p>
      <p>
        Two molecular biology and biochemistry experts, Michel
Dumontier and Pascale Gaudet, assigned a score from low to
high confidence of 1-3, evidence and/or technique associated
with the interaction. The confidence measure was based on the
comparative analyses of techniques [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and experience
of the experts in reviewing data of this kind. The confidence
assignment is based on a number of factors including degree
of indirection in the assay, sensitivity and specificity of the
approach, and reproducibility of results under different
conditions. The confidence scores for both experts were encoded as
classes of evidence, where each experimental method class was
assigned two superclasses, one for each expert. This ontology
was created from a spreadsheet and expanded to full inferences
using Pellet [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. At the same time, SPARQL-based reasoning
is used to classify nanopublication assertions by their available
evidence, and thereby assign a class of confidence codes to it.
      </p>
      <sec id="sec-2-1">
        <title>B. SADI Web Service Interface</title>
        <p>We developed four Semantic Automated Discovery and
Integration (SADI) web services in Python1 to support easy
access to the nanopublications. We use SADI to provide a
discoverable, consistent API that can be re-used in other
applications or directly consumed by analytical tools.</p>
        <p>The services perform these computational tasks that would
otherwise be difficult to perform with SPARQL queries. The
services return only one interaction for each triple (source,
interaction type, target) but multiple, probabilities per
interaction, and more than one interaction per interaction type.
This is because the interaction may have been recorded in
multiple databases, based on different experimental methods.
To provide a single probability score for each triple, the
interactions are combined. This is done to indicate that multiple
experiments that produce the same results reinforce each other,
and should therefore give a higher overall probability than
would be indicated by taking their mean.</p>
        <p>P (x1...n) = CDF
n
X CDF 1 (P (xi))</p>
        <p>!
i=1
1For further information on developing web services in Python
using SADI, see this tutorial: https://code.google.com/p/sadi/wiki/
BuildingServicesInPython</p>
      </sec>
      <sec id="sec-2-2">
        <title>C. User Interface</title>
        <p>Users can search for biological entities and processes, which
can then be autocompleted to specific entities that are in
the ReDrugS graph. Users can then add those entities and
processes to the displayed graph and retrieve upstream and
downstream connections and link out to more details for every
entity. Cytoscape.js is used as the main rendering and network
visualization tool, and provides node and edge rendering,
layout, and network analysis capabilities.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>III. EVALUATION</title>
      <p>In order to evaluate this knowledge base, we developed
a demonstration web interface2 based on the Cytoscape.js3.
It lets users enter biological entity names, and as the user
types, the text is resolved to a list of entities to be selected.
After that, the entity is submitted to all three SADI services
via a basic JavaScript SADI client.4 The resulting interactions
and nodes are added to the Cytoscape.js graph, which can be
laid out according to a number of algorithms. Users are also
able to select nodes and populate upstream or downstream
connections. An example of this is shown in Figure 1. This
figure was obtained by putting “Topiramate” as a query in
the search box, which returned all of the biological entities
that topiramate is directly associated with. We then expanded
the network downstream to see what biological entities are
affected by topiramate’s targets.</p>
      <p>We are able to successfully navigate a protein-drug-disease
interaction graph that is a consensus of 16 diverse sources, to
infer prior probabilities for more than three million individual
assertions using their provenance and experts’ confidence in
different experimental methods and to find drug/disease
associations that are not directly expressed by any one database.</p>
      <p>
        We plan to add further data sources, especially those that
provide direct experimental results that predict protien (or
gene)/disease associations like the Gene Expression Atlas (in
progress) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Further, we are very interested in integrating
the newest version of the Connectivity Map dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], as
it provides gene expression signature similarities for a large
number of chemical and genetic perturbations. Finally, as we
develop new hypotheses about potential new drug effects,
we plan to test them using a new three-dimensional cellular
microarray to perform high throughput drug screening [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] with
reference samples.
      </p>
    </sec>
    <sec id="sec-4">
      <title>V. CONCLUSION</title>
      <p>We have developed a framework for collecting, searching,
analyzing, and visualizing important components of biological
systems. We were able to build this by converting existing
databases into a common nanopublication structure that uses
the provenance of the database records to determine the quality
of any given piece of information through the methods used
to provide it. We use the Semantic Automated Discovery and
Integration framework to provide simple access to data, and
can visualize results using an existing interaction graph tool.
The resulting application makes it easy to search for biological
entities and see how they interact. We have already found
some hypotheses of proteins through which drugs influence
disease conditions. We plan to expand the loaded set of data
with protein/disease associations as well as gene expression
profiles, and will be using ReDrugS to produce prospective
testable hypotheses.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>A special thanks to Pascale Gaudet, who, with Michel
Dumontier, evaluated the experimental methods and evidence
codes listed in the Protein/Protein Interaction Ontology and
Gene Ontology.</p>
      <p>Tall version (Preferred)
Use this version in the majority of cases.
5AMwwSaoaslsuGtieoennnseeLrroaagtteeoddUBBsyayge Guidelines</p>
      <p>G</p>
      <p>XX
aaMdI_ir0e4c0t7interaction
hhasa-st-atragregte:t: SLSCL4CA48A8
hhasa-sp-apratrictiicpiapnatn:t:CAC2A2</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCusker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lebo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krauthammer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. L.</given-names>
            <surname>McGuinness</surname>
          </string-name>
          , “
          <article-title>Next Generation Cancer Data Discovery, Access, and Integration Using Prizms and Nanopublications,” in Data Integration in the Life Sciences</article-title>
          . Springer,
          <year>2013</year>
          , pp.
          <fpage>105</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Obenauer and M. B. Yaffe</surname>
          </string-name>
          , “
          <article-title>Computational prediction of proteinprotein interactions,” in Protein-Protein Interactions</article-title>
          . Springer,
          <year>2004</year>
          , pp.
          <fpage>445</fpage>
          -
          <lpage>467</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sprinzak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sattath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Margalit</surname>
          </string-name>
          , “
          <article-title>How reliable are experimental protein-protein interaction data?” Journal of molecular biology</article-title>
          , vol.
          <volume>327</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>919</fpage>
          -
          <lpage>923</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sirin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Parsia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalyanpur</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Katz</surname>
          </string-name>
          , “
          <article-title>Pellet: A practical owl-dl reasoner,” Web Semantics: science</article-title>
          ,
          <source>services and agents on the World Wide Web</source>
          , vol.
          <volume>5</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>51</fpage>
          -
          <lpage>53</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Petryszak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Burdett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fiorelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Fonseca</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. GonzalezPorta</surname>
            , E. Hastings,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Huber</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Jupp</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Keays</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Kryvych</surname>
          </string-name>
          , and et al.,
          <article-title>“Expression Atlas update-a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments</article-title>
          ,
          <source>” Nucleic Acids Research</source>
          , vol.
          <volume>42</volume>
          , no.
          <source>D1</source>
          , p.
          <fpage>D926</fpage>
          -
          <lpage>D932</lpage>
          ,
          <year>Jan 2014</year>
          . [Online]. Available: http://dx.doi.org/10.1093/nar/gkt1270
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lamb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Crawford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Peck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Modell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. C.</given-names>
            <surname>Blat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Wrobel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lerner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-P.</given-names>
            <surname>Brunet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Subramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Ross</surname>
          </string-name>
          et al., “
          <article-title>The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease,” science</article-title>
          , vol.
          <volume>313</volume>
          , no.
          <issue>5795</issue>
          , pp.
          <fpage>1929</fpage>
          -
          <lpage>1935</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.-Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Sukumaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Hogg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Clark</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Dordick</surname>
          </string-name>
          , “
          <article-title>Three-dimensional cellular microarray for high-throughput toxicology assays</article-title>
          ,
          <source>” Proceedings of the National Academy of Sciences</source>
          , vol.
          <volume>105</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>63</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>