<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>KG-Microbe: A Reference Knowledge-Graph and Platform for Harmonized Microbial Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcin P. Joachimiak</string-name>
          <email>MJoachimiak@lbl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harshad Hegde</string-name>
          <email>hhegde@lbl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William D. Duncan</string-name>
          <email>wdduncan@lbl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Justin T. Reese</string-name>
          <email>JustinReese@lbl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Cappelletti</string-name>
          <email>luca.cappelletti1@unimi.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anne E. Thessen</string-name>
          <email>thessena@oregonstate.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher J. Mungall</string-name>
          <email>CJMungall@lbl.gov</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Environmental Genomics and Systems Biology Division, Lawrence Berkeley National Laboratory</institution>
          ,
          <addr-line>Berkeley, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Oregon State University</institution>
          ,
          <addr-line>Beaverton, Oregon</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Milan</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Microorganisms (microbes) are incredibly diverse, spanning all major divisions of life, and represent the greatest fraction of known species. A vast amount of knowledge about microbes is available in the literature, across experimental datasets, and in established data resources. While the genomic and biochemical pathway data about microbes is wellstructured and annotated using standard ontologies, broader information about microbes and their ecological traits is not. We created the KG-Microbe (github.com/Knowledge-GraphHub/kg-microbe) resource in order to extract and integrate diverse knowledge about microbes from a variety of structured and unstructured sources. Initially, we are harmonizing and linking prokaryotic data for phenotypic traits, taxonomy, functions, chemicals, and environment descriptors, to construct a knowledge graph with over 266,000 entities linked by 432,000 relations. The effort is supported by a knowledge graph construction platform (KG-Hub) for rapid development of knowledge graphs using available data, knowledge modeling principles, and software tools. KG-Microbe is a microbe-centric Knowledge Graph (KG) to support tasks such as querying and graph link prediction in many use cases including microbiology, biomedicine, and the environment. KG-Microbe fulfills a need for standardized and linked microbial data, allowing the broader community to contribute, query, and enrich analyses and algorithms.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Knowledge graph</kwd>
        <kwd>microbiology</kwd>
        <kwd>ontology</kwd>
        <kwd>graph learning</kwd>
        <kwd>data standardization</kwd>
        <kwd>data science</kwd>
        <kwd>semantic technology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Not only are microbes the most abundant and diverse life forms, but they are also found in the
greatest range of environments and possess the largest metabolic and functional potential which is just
beginning to be harnessed for biomedicine and biomanufacturing. A vast amount of knowledge about
microbes is available in the literature, across experimental datasets, and in established data resources.
While the genomic and biochemical pathway data about microbes is well-structured and annotated
using standard ontologies, broader information about microbes and their ecological traits is not.</p>
      <p>
        We draw inspiration from the biomedical domain, which has a rich set of ontologies, controlled
vocabularies, and data schemas, which have been deployed in multiple biomedical knowledge
resources. Examples include data schemas such as OMOP [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], MeSH [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and the Biolink Model [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
collections of ontologies such as the OBO Foundry [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and the NCBO Bioportal [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and a growing
ecosystem of graph and Natural Language Processing (NLP) tools. These resources have helped drive
the standardization and interoperability of information across research domains. This set of concepts
and tools provides a path to useful semantic harmonization of knowledge in other domains.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Main</title>
      <p>
        The KG-Microbe knowledge graph (KG) resource was created to support extraction and
integration of diverse knowledge about microbes. Initially, our focus has been on harmonizing and
linking prokaryotic data for phenotypic traits, taxonomy, functions, chemicals, and environment
descriptors. Based on this information we constructed a knowledge graph, which in its current release
(9/2/21) contains over 266,000 entities linked by 432,000 relations, classified into 9 and 31 Biolink
Model entity and relation categories, respectively. The KG-Microbe knowledge graph effort is
supported by a knowledge graph construction platform, KG-Hub [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], designed for rapid development
and deployment of knowledge graphs using available data, common knowledge modeling principles,
and software tools. One key concept for KG-Hub is to connect unstructured data into broader
knowledge using links to structured data such as ontologies.
      </p>
      <p>KG-Microbe (github.com/Knowledge-Graph-Hub/kg-microbe) is a microbe-centric Knowledge
Graph (KG) to support tasks such as querying and graph link prediction in a variety of use cases
including microbiology, biomedicine, and the environment. We use Named Entity Recognition (NER)
and Natural Language Processing (NLP) tools to identify, annotate, and normalize terms found in raw
data. The harmonized data contained in the KG-Microbe knowledge graph provides rich and
standardized labeling for building, training, and evaluating machine learning models. The resulting
KG-Microbe graph is able to answer questions like which microbes are enriched in soil environments.
It can also be used to train models for various microbial trait predictions and it can report enriched
features for a given set of taxa or taxa features. We demonstrate example applications of KG-Microbe
with predictive models for microbial shape and metabolism using embeddings (Figure 1) from graph
learning. Many other types of link predictions are possible based on the available KG-Microbe entity
and relation categories, allowing predictions for data, which is difficult to obtain without resource
intensive field and laboratory experiments such as metabolic characterization or cell imaging.
KGMicrobe fulfills a need for standardized and linked microbial data, allowing the broader community to
contribute as well as enrich analyses and algorithms.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Acknowledgements</title>
      <p>This work was supported by a grant from the Laboratory Directed Research and Development
(LDRD) Program of Lawrence Berkeley National Laboratory under U.S. Department of Energy
Contract No. DE-AC02-05CH11231.</p>
    </sec>
    <sec id="sec-4">
      <title>4. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.A.</given-names>
            <surname>Voss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Makadia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Matcho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ma</surname>
          </string-name>
          , C. Knoll,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuemie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.J.</given-names>
            <surname>DeFalco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Londhe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.B.</given-names>
            <surname>Ryan</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Feasibility and Utility of Applications of the Common Data Model to Multiple, Disparate Observational Health Databases</article-title>
          .
          <source>Journal of the American Medical Informatics Association: JAMIA</source>
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <fpage>553</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.B.</given-names>
            <surname>Searls</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Can Literature Analysis Identify Innovation Drivers in Drug Discovery? Nature Reviews</article-title>
          .
          <source>Drug Discovery</source>
          <volume>8</volume>
          (
          <issue>11</issue>
          ):
          <fpage>865</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.J.</given-names>
            <surname>Mungall</surname>
          </string-name>
          , et al.
          <year>2021</year>
          .
          <article-title>Biolink Model</article-title>
          . URL: https://github.com/biolink/biolink-model.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashburner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rosse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Bug</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ceusters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.J.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          , et al.
          <year>2007</year>
          .
          <article-title>The OBO Foundry: Coordinated Evolution of Ontologies to Support Biomedical Data Integration</article-title>
          .
          <source>Nature Biotechnology</source>
          <volume>25</volume>
          (
          <issue>11</issue>
          ):
          <fpage>1251</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.L.</given-names>
            <surname>Whetzel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.F.</given-names>
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.H.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.R.</given-names>
            <surname>Alexander</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nyulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tudorache</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Musen</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>BioPortal: Enhanced Functionality via New Web Services from the National Center for Biomedical Ontology to Access and Use Ontologies in Software Applications</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>39</volume>
          (
          <issue>Web Server issue</issue>
          ):
          <fpage>W541</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.J.</given-names>
            <surname>Mungall</surname>
          </string-name>
          , et al.
          <year>2021</year>
          .
          <article-title>KG-Hub: a knowledge graph hub</article-title>
          . URL: https://knowledge-graphhub.github.io/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>L. van der Maaten</surname>
          </string-name>
          , L. and
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Visualizing Data Using T-SNE</article-title>
          .
          <source>Journal of Machine Learning Research: JMLR</source>
          <volume>9</volume>
          (Nov):
          <fpage>2579</fpage>
          -
          <lpage>2605</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappelletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fontana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Casiraghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ravanmehr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Callahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Joachimiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Mungall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Reese</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Valentini</given-names>
            <surname>Giorgio</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>GraPE: fast and scalable Graph Processing and Embedding</article-title>
          . arXiv:
          <volume>2110</volume>
          .
          <fpage>06196</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>