<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NAPRALERT, from an historical information silo to a linked resource able to address the new challenges in Natural Products Chemistry and Pharmacognosy.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jonathan Bisson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James McAlpine</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James Graham</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido F. Pauli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Natural Product Technologies, University of Illinois at Chicago</institution>
          ,
          <country country="US">United States</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>- NAPRALERT (https://www.napralert.org) is a database on natural products, including data on ethnobotany, chemistry, pharmacology, toxicology, and clinical trials from literature dating back to the 19th century. Established in 1975 by Norman R. Farnsworth, it became a web accessible resource in 2005 but soon became stagnant while literature grew exponentially. After a complete rewrite of the platform, the focus is now on connecting this resource to the rest of the existing databases and expanding its usability. The creation of a Pharmacognosy/Natural Product ontology will foster better understanding of this domain, its linking potential with other resources and the ability to automatize literature annotation and entry efficiently.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        The late Norman R. Farnsworth established
NAPRALERT[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as a tool to survey Natural Product research
in 1975, when relational databases were initially being
adopted. Before formal methodologies for developing
ontologies existed, Professor Farnsworth developed his own
simple but somewhat exhaustive hierarchical classification
system and used it to annotate information gathered from the
literature by him and his team for the following 30+ years.
      </p>
      <p>NAPRALERT became web-accessible in 2005 with more
than 200,000 citations covered as of 2015, but the informal
classification system itself remained hidden behind the
interface and was only accessible through customized manual
queries. NAPRALERT became quiescent due to budget
constraints around 2004, at a time of exponential expansion of
the literature and development of new resources, new tools and
new knowledge.</p>
      <p>In 2015, a complete rewrite of the code and database
schema (Fig 1.) was undertaken. A new web version (October
1, 2015) provides limited free searches to academics, industry,
and governmental agencies. Currently this system is being
enhanced with improved search and data entry functionalities.
However, the fundamental limitations of the current approach
are fully realized as it is now time to connect with existing
repositories of bibliographical data, chemical structures, and
biological activities including data outside the Natural Product
literature previously covered by NAPRALERT.</p>
    </sec>
    <sec id="sec-2">
      <title>III. LINKING RESOURCES TO THE DATABASE</title>
      <p>Pharmacognosy is a domain at the intersection of Biology,
Biochemistry, Botany, Ethnobotany, Pharmacy and Chemistry.
The data sources required relative to each discipline are diverse
and cover at the same time a wider and a narrower range than
what is relevant to this field, and thereby requiring very careful
mapping. An example of the apparent dissonance with this
domain and existing resources is ChEBI, the ontology of which
covers many but not all kinds of Natural Products and at the
same time covers non-natural “molecular entities”. However,
creating such entities de-novo in this domain ontology would
only reduce both its interoperability and maturity. Moreover,
whenever these resources provide a formal and interoperable
ontology some inference can be made to enrich the queries of
users and potentially the coverage of the domain.</p>
      <p>IV. MACHINE LEARNING BASED LITERATURE ANNOTATION</p>
      <p>
        During the first two decades of NAPRALERT, it was still
feasible for a relatively small team to work on the annotation
and entry of literature data. Nowadays, such an approach is
unrealistic both in terms of managing the coherence and
validity of the data entry and, particularly, due to the
tremendous human resources required. One of the approaches
considered for the future of NAPRALERT is similar to what is
currently achieved with projects such as GeoDeepDive
(https://geodeepdive.org) and PaleoDeepDive [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These
projects demonstrated the efficiency of mixing Natural
Language Processing, Machine Learning, and their already
annotated corpus of publications, making it possible to
annotate literature more rapidly but still accurately.
      </p>
      <p>
        One of the expected use cases of this ontology is to
determine whether or not the compounds identified in natural
product extracts are truly responsible for, or involved in, the
observed bioactivity. This requires the ability to compare and
study bioactivity data broadly, from many different data
sources and with respect to chemical composition, and to rate
the confidence in both. Moreover, as compounds with
promiscuous and unspecific activities are surprisingly often
mistaken for active principles, it is important to identify them
and take their characteristics into account when searching for
bioactives [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For this important use case, the new ontology
will facilitate ways to annotate and link data from other
resources.
      </p>
    </sec>
    <sec id="sec-3">
      <title>ACKNOWLEDGMENT</title>
      <p>The authors acknowledge support by grant U41 AT008706
from NCCIH and ODS/NIH. The authors also appreciate the
creative advice of Sheila Miguez and Aaron Lav (Pumping
Station: ONE, Chicago, IL).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Graham</surname>
          </string-name>
          and
          <string-name>
            <given-names>N. R.</given-names>
            <surname>Farnsworth</surname>
          </string-name>
          , “
          <article-title>The NAPRALERT database as an aid for discovery of novel bioactive compounds.” in Comprehensive Natural Products</article-title>
          , vol. II, H.-W. Liu and L. Mander, Eds. Elsevier: Amsterdam,
          <year>2010</year>
          , pp
          <fpage>81</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Livny</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Ré</surname>
          </string-name>
          ,
          <string-name>
            <surname>“</surname>
          </string-name>
          <article-title>A machine reading system for assembling synthetic paleontological databases</article-title>
          .”
          <source>PLoS ONE</source>
          , vol
          <volume>9</volume>
          (
          <issue>12</issue>
          ), pp:
          <fpage>e113523</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bisson</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. B. McAlpine</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          <string-name>
            <surname>Friesen</surname>
            , SN Chen,
            <given-names>J. G.</given-names>
          </string-name>
          <string-name>
            <surname>Graham</surname>
            ,
            <given-names>G. F.</given-names>
          </string-name>
          <string-name>
            <surname>Pauli</surname>
          </string-name>
          , “
          <article-title>Can invalid bioactives undermine Natural Product-based drug discovery</article-title>
          ”
          <source>Journal of Medicinal Chemistry</source>
          , vol
          <volume>59</volume>
          (
          <issue>5</issue>
          ), pp
          <fpage>1671</fpage>
          -
          <lpage>1690</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>