<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mondo Disease Ontology: harmonizing disease concepts across the world</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicole Vasilevsky</string-name>
          <email>vasilevs@ohsu.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shahim Essaid</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nico Matentzoglu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nomi L. Harris</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Melissa Haendel</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Robinson</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher J. Mungall</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>European Bioinformatics Institute</institution>
          ,
          <addr-line>Hinxton</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lawrence Berkeley National Laboratory</institution>
          ,
          <addr-line>Berkeley, CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Oregon Health &amp; Science University</institution>
          ,
          <addr-line>Portland, OR</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Oregon State University</institution>
          ,
          <addr-line>Corvallis, OR</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>The Jackson Laboratory</institution>
          ,
          <addr-line>Farmington, CT</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <kwd-group>
        <kwd />
        <kwd>Disease</kwd>
        <kwd>Rare disease</kwd>
        <kwd>Ontology</kwd>
        <kwd>Clinical Terminology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Many terminologies and ontologies currently exist for classifying and describing
diseases, each of which has been designed for a particular purpose, and as such has
different strengths. Examples include the National Cancer Institute Thesaurus (NCIt),
the Online Mendelian Inheritance in Man (OMIM), SNOMED CT, ICD, ICD-O,
OncoTree, MedGen, Disease Ontology, and numerous others. Our census of disease
ontologies identified 37 such resources. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] However, these standards partially overlap
and often conflict, making it difficult to align knowledge sources. This need to integrate
information has resulted in a proliferation of mappings and cross-references between
disease entries in different resources; these mappings lack completeness, accuracy, and
precision, and are often inconsistent between resources.
      </p>
      <p>In order to computationally leverage the wide array of available knowledge sources
for diagnostics and to reveal underlying mechanisms of diseases, we need to understand
which terms are truly equivalent across different resources. This will allow integration
of associated information, such as treatments, genetics, phenotypes, etc. We therefore
created the Mondo Disease Ontology to provide a logic-based structure for unifying
multiple disease resources. The goal of Mondo is to integrate the classifications and
relationships of commonly used disease ontologies into a single semantically coherent
resource to enable the aggregation and analysis of disparate clinical data repositories and
facilitate the discovery of relationships between disease concepts across ontologies</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <p>
        Mondo was created by a combination of algorithmic equivalency determination
using the k-BOOM algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and expert curation. Mondo spans both rare and
‘common’ disease, comprising monogenic and polygenic (complex, common) diseases,
infectious diseases, trauma, and cancer. It provides equivalence mappings to other
disease resources, but in contrast to other mapping sets, Mondo precisely annotates each
mapping using strict semantics, so that we know when two diseases are precisely
equivalent or merely closely related - allowing computational integration of associated
data.
      </p>
      <p>Mondo provides a hierarchical structure which can be used to annotate data at
different levels of precision. Terms are classified under multiple parents, as needed, and
external ontologies are imported for use in equivalence and subclass of axioms.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Discussion</title>
      <p>Mondo currently contains over 20,000 disease classes. It is iteratively developed
with the community and is under continuous revision, with future plans to further
revise the top-level classes. Mondo is an OBO Foundry ontology and can be
viewed on the OBO website (obofoundry.org/ontology/mondo.html), or with other
ontology browsers such as the Ontology Lookup Service (OLS)
(ebi.ac.uk/ols/ontologies/mondo). The ontology files and current releases are
available on GitHub (https://github.com/monarch-initiative/mondo), with releases
made on a monthly basis. We invite the community to contribute to Mondo; visit
our website, GitHub repository and/or join our mailing list
(mondousers@googlegroups.com).</p>
      <p>
        Mondo is being utilized in diverse applications and resources such as the Monarch
Initiative [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], ClinGen [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and Gabriella Miller Kids First Data Resource [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
which is a curated database of clinical and genetic sequence data from pediatric
patients with structural abnormalities or childhood cancers. The Experimental
Factor Ontology (EFO) employs Mondo’s classifications and axioms to facilitate
annotation of disease information in applications that use EFO. By integrating
information from multiple disease resources, amassed through years of work by
researchers, clinicians, ontologists, and other scientists from around the world,
Mondo aims to make this knowledge readily available to the scientific community
and grow its value through logical connections across resources.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Haendel</surname>
            <given-names>MA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McMurry</surname>
            <given-names>JA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Relevo</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            <given-names>PN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            <given-names>CG</given-names>
          </string-name>
          .
          <article-title>A Census of Disease Ontologies</article-title>
          .
          <source>Annu Rev Biomed Data Sci</source>
          .
          <year>2018</year>
          ;
          <volume>1</volume>
          :
          <fpage>305</fpage>
          -
          <lpage>331</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Mungall</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koehler</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haendel</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>k-BOOM: A Bayesian approach to ontology structure inference, with applications in disease ontology construction</article-title>
          .
          <source>bioRxiv</source>
          .
          <year>2019</year>
          . p.
          <volume>048843</volume>
          . doi:
          <volume>10</volume>
          .1101/048843
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Shefchek</surname>
            <given-names>KA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            <given-names>NL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gargano</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matentzoglu</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Unni</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brush</surname>
            <given-names>M</given-names>
          </string-name>
          , et al.
          <article-title>The Monarch Initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species</article-title>
          .
          <source>Nucleic Acids Res</source>
          .
          <year>2020</year>
          ;
          <volume>48</volume>
          :
          <fpage>D704</fpage>
          -
          <lpage>D715</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Clinical</given-names>
            <surname>Genome Resource</surname>
          </string-name>
          . Welcome to ClinGen. Available: https://clinicalgenome.org/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Working</given-names>
            <surname>Together</surname>
          </string-name>
          to Put Kids First. Available: https://kidsfirstdrc.org/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>