<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building a molecular glyco-phenotype ontology to decipher undiagnosed diseases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jean-Philippe Gourdine*, David M. Koeller*,</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas O. Metz*</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Matthew H. Brush, Melissa A. Haendel*, Oregon Health and Science University</institution>
          ,
          <addr-line>Portland, Oregon</addr-line>
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Pacific Northwest National Laboratory</institution>
          ,
          <addr-line>Richland, WA</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Hundreds of rare diseases are due to mutation on genes related to glycans synthesis, degradation or recognition. These glycan-related defects are well described in the literature but largely absent in ontologies and databases of chemical entities and phenotypes, limiting the application of computational methods and ontology-driven tools for characterization and discovery of glycan related diseases. We are curating articles and textbooks in glycobiology related to genetic diseases to inform the content and the structure of an ontology of Molecular GlycoPhenotypes (MGPO). MGPO will be applied toward use cases including disease diagnosis and disease gene candidate prioritization, using semantic similarity and pattern matching at the glycan level with glycomics data from patient of the Undiagnosed Diseases Network.</p>
      </abstract>
      <kwd-group>
        <kwd>rare diseases</kwd>
        <kwd>glycans</kwd>
        <kwd>phenotypes</kwd>
        <kwd>ontology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Rare diseases and disorders affect 25 million people in the
United States, 30 million people Europe and 350 million
worldwide [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. 7,000 rare diseases and disorders have been
identified and 80% of them are related to genetic defects [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
More than 100 of these genetic defects are related to glycan
biology, affecting glycan synthesis, degradation or recognition
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Glycans, also referred to as sugars, carbohydrates,
monosaccharides or polysaccharides, can be found free or
attached to macromolecules to form glycosphingolipids when
attached to lipids, N or O-linked glycoproteins when attached
to proteins or glycophosphatidylinositol (GPI)-anchored
glycoproteins when glycoproteins are attached to GPI.
      </p>
      <p>
        Two broad categories define their biological roles:
structural/modulatory function and intrinsic/extrinsic
molecular recognition [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Glycan-related diseases can be
observed by their glycan markers or glyco-phenotypes,
defined as an abnormality related to glycan structure, level,
activity, and processing. Many of these glycophenotypes can
be detected in patients’ bodily fluids or cells (e.g. urine [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
      </p>
      <p>
        The Human Phenotype Ontology (HPO) is a structured
clinical vocabulary that has been used for deep clinical
phenotyping and allows the integration of data across sources
and organisms. The ontology is the de facto standard for
clinical phenotyping for rare disease diagnostics using
phenotypic profile comparison [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Glycan-related defects and
the molecules they affect are well described in the literature,
but largely absent in ontologies such as the HPO, and from
B. Development of initial MGPO prototype.
      </p>
      <p>With a rich set of curated data in hand, the next challenge is
to synthesize this information to identify the key dimensions of
glycan phenotypes on which the ontology would be classified.
We begin by listing each distinct glyco-phenotype curated
from the sources above, and looked for patterns that would
reveal important features for their categorization. Four essential
dimensions reveal themselves through this exercise:
1. Affected glycan characteristic. This dimension
distinguishes phenotypes based on the general type of
glycan feature or process affected: (e.g. glycan structure,
occupancy, composition, levels, activity).
2. Affected glycan type. This dimension distinguishes
phenotypes based on the class of glycan exhibiting
abnormal characteristics.
3. Affected glycosylation target. This dimension
distinguishes phenotypes based on the aglycone target
affected by a defect.
4. Affected locus of abnormality. This dimension
distinguishes phenotypes based on the subcellular
location of the defect and/or the tissue or fluid where it
was measured.</p>
      <p>The next step was applying these dimensions as
classification axes to build a hierarchical model. Careful
consideration was given to the order in which the
dimensions should be applied to confer the optimal
structure to the ontology. An initial straw man was
generated for testing and eliciting expert feedback for
iterative improvement, and is available at
http://bit.ly/23kdoyj
C. Evaluation and feedback provided to related databases
and ontologies such as ChEBI and GO</p>
      <p>We next evaluated glycan-related coverage of orthogonal
efforts including ChEBI, GO, HPO, and various glycan
databases. Here we noted an underrepresentation of concepts
needed to describe glyco-phenotypes like those we had
curated form the literature. In GO, grouping categories are
missing, for instance, glycan binding proteins (R-type lectin,
L-type lectin, C-type lectin, P-type lectin, I-type lectin,
galectin). Also, glycan-related synonyms are not present (e.g.
“core 1” = T antigen).</p>
      <p>
        In ChEBI, labels are not always aligned with common
domain use - for instance ChEBI’s use of ‘glycans' as
synonym for 'carbohydrate'. There are key gaps in
representation, for instance, no representation of N- and
Oglycans as groups, absence of many molecular species
reported as diseases markers, absence of synonyms commonly
used in glycobiology, etc. In HPO/MP, glyco-phenotypes are
also underrepresented. For instance, a search for terms
“oligosacchariduria” shows only sialidated and disaccharides
excretion while more than fifty glycophenotypes were
detected in the urine from ten different diseases [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In
metabolomics databases, glycans with masses under 1,500 da
are severely underrepresented with mostly only mono or
disaccharides and their derivatives. For instance, a search of the
term “glycan” in the ‘Metabolomics Workbench’ database [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
leads to only 6 molecules, while more than 30 N-glycan
precursors can be identified in the Endoplasmic Reticulum [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        This contrasts with the abundance of chemical
characteristics of glycans in glycobiologist databases such as
UniCarbKB, GlycomeDB and Japan Consortium For
Glycobiology and Glycotechnology Database (JCGGDB) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Nevertheless, links between glycophenotypes and diseases are
still lacking in these glycan databases.
      </p>
    </sec>
    <sec id="sec-2">
      <title>III. APPLICATIONS AND FUTURE DIRECTIONS</title>
    </sec>
    <sec id="sec-3">
      <title>The MGPO will be integrated with existing</title>
      <p>
        phenotype ontologies such as the Human Phenotype Ontology
(HPO), and Mammalian Phenotype Ontology (MP), to support
research and clinical application of ontology-driven methods.
MGPO will serve as a bridge between existing glycan related
databases and metabolomics, databases, between glycan
databases and diseases databases. As part of the Monarch
Initiative [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the MGPO will be applied in semantic similarity
approaches for the purposes of (1) disease diagnosis, (2)
identification of patient cohorts for clinical studies, (3)
discovery of model organisms to researching rare disease, and
(4) identifying candidate genes underlying undiagnosed
diseases. The MGPO will also be applied to annotate
glycomics data from our partners in the Undiagnosed Disease
Network, and support its use in semantic similarity and pattern
matching analyses that can facilitate disease diagnosis and
characterization.
      </p>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGMENT</title>
      <p>We thank the Ontology Development Group (ODG) at the
OHSU’s library for their feedback.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] https://globalgenes.org/rare-diseases
          <string-name>
            <surname>-</surname>
          </string-name>
          facts-statistics/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] https://www.genome.gov/27531963/faq-about
          <string-name>
            <surname>-</surname>
          </string-name>
          rare-diseases/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Freeze</surname>
            <given-names>HH.</given-names>
          </string-name>
          “
          <article-title>Genetic defects in the human glycome</article-title>
          .
          <source>” Nat Rev Genet</source>
          .
          <source>2006 Jul;7</source>
          (
          <issue>7</issue>
          ):
          <fpage>537</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Varki</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cummings</surname>
            <given-names>RD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esko</surname>
            <given-names>JD</given-names>
          </string-name>
          , et al., editors. Cold Spring Harbor (NY): Cold Spring Harbor Laboratory Press;
          <year>2009</year>
          . “Chapter 6,
          <string-name>
            <given-names>Biological</given-names>
            <surname>Roles</surname>
          </string-name>
          of Glycans,
          <source>Essentials of Glycobiology.” 2nd edition</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Xia</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asif</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arthur</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pervaiz</surname>
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>R.</given-names>
          </string-name>
          , et al. “
          <article-title>Oligosaccharide analysis in urine by maldi-tof mass spectrometry for the diagnosis of lysosomal storage diseases</article-title>
          .
          <source>” Clin Chem</source>
          .
          <year>2013</year>
          Sep;
          <volume>59</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1357</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Haendel</surname>
            <given-names>MA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasilevsky</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brush</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hochheiser</surname>
            <given-names>HS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobsen</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oellrich</surname>
            <given-names>A</given-names>
          </string-name>
          , et al. “
          <article-title>Disease insights through cross-species phenotype comparisons</article-title>
          .
          <source>” Mamm Genome</source>
          . 2015 Oct;
          <volume>26</volume>
          (
          <issue>9</issue>
          -10):
          <fpage>548</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Sud</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fahy</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cotter</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azam</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vadivelu</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burant</surname>
            <given-names>C</given-names>
          </string-name>
          , et al. “
          <article-title>Metabolomics Workbench: An international repository for metabolomics data and metadata, metabolite standards, protocols, tutorials and training, and analysis tools</article-title>
          .
          <source>” Nucleic Acids Res</source>
          .
          <source>2016 Jan</source>
          <volume>4</volume>
          ;
          <fpage>44</fpage>
          (
          <issue>D1</issue>
          ):
          <fpage>D463</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ng</surname>
            <given-names>B.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolfe</surname>
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ichikawa</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markello</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tifft C.J.</surname>
          </string-name>
          , et al. “
          <article-title>Biallelic mutations in CAD, impair de novo pyrimidine biosynthesis and decrease glycosylation precursors</article-title>
          .”
          <source>Hum Mol Genet</source>
          .
          <source>2015 Jun</source>
          <volume>1</volume>
          ;
          <issue>24</issue>
          (
          <issue>11</issue>
          ):
          <fpage>3050</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Aoki-Kinoshita</surname>
            <given-names>KF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolleman</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campbell</surname>
            <given-names>MP</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawano</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            <given-names>JD</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lütteke</surname>
            <given-names>T</given-names>
          </string-name>
          , et al. “
          <article-title>Introducing glycomics data into the Semantic Web</article-title>
          ”
          <string-name>
            <given-names>J. Biomed</given-names>
            <surname>Semantics</surname>
          </string-name>
          .
          <source>2013 Nov</source>
          <volume>26</volume>
          ;
          <issue>4</issue>
          (
          <issue>1</issue>
          ):
          <fpage>39</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>