<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>EGO: a biomedical ontology for integrative epigenome representation and analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yongqun He</string-name>
          <email>yongqunh@med.umich.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jie Zheng</string-name>
          <email>jiezheng@upenn.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhaohui Qin</string-name>
          <email>zhaohui.qin@emory.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Emory University</institution>
          ,
          <addr-line>Atlanta, GA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Philadelphia</institution>
          ,
          <addr-line>PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Epigenomics is crucial to understand biological mechanisms beyond genome DNA. To better represent epigenomic knowledge and support data integration, we developed a prototype Epigenome Ontology (EGO). EGO top level hierarchy and design pattern are provided with a use case illustration. EGO is proposed to be used for statistically analyzing enriched epigenomic features based on given sequence data input using statistical methods.</p>
      </abstract>
      <kwd-group>
        <kwd>Epigenome</kwd>
        <kwd>ontology</kwd>
        <kwd>EGO</kwd>
        <kwd>enrichment analysis</kwd>
        <kwd>ENCODE</kwd>
        <kwd>ChIP-seq</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>The majority of eukaryotic genomes such as those of the
human and mouse is noncoding. In eukaryote genomes, basic
biological functions such as gene expression are affected by
many regulatory elements located outside the coding region of
the genome. Unlike the genome, which is largely static across
tissues and cells within an individual, the epigenome is cell
type specific and can be dynamically altered by environmental
conditions. Better epigenomic understanding is critical for
uncovering biological mechanisms and disease etiology.</p>
    </sec>
    <sec id="sec-2">
      <title>Our understanding of the epigenome has dramatically</title>
      <p>improved in the past decade thanks to the efforts of several
large international consortia, e.g., the Encyclopedia of DNA
Elements (ENCODE; https://www.encodeproject.org/). As
more and more cells and cell lines are being profiled, combined
with more factors being studied, it becomes difficult to track all
the experiments and resulting knowledge. In particular, many
of the experiments are related due to the fact that either the cell
types, or the experimental factors are related or both. Thus it is
inadvisable to treat these data as independent.</p>
      <p>For better interpretation, the subtle and complicated
relationships among these experiments should be more fully
considered to achieve a better interpretation. Specifically, a
well-defined ontology system that handles complex
relationships within a rigorous framework and offers
annotation at various levels of granularity would be
particularly effective. Towards such goals, we developed a new
ontology named Epigenome Ontology (EGO).</p>
    </sec>
    <sec id="sec-3">
      <title>II. METHODS</title>
      <p>A. EGO ontology development</p>
    </sec>
    <sec id="sec-4">
      <title>EGO development follows the OBO Foundry principles</title>
      <p>(http://obofoundry.org/). Existing terms from other ontologies
were imported to EGO using OntoFox
(http://ontofox.hegroup.org). New terms were added and edited
using the Protégé OWL editor.</p>
    </sec>
    <sec id="sec-5">
      <title>An EGO ontology design pattern was generated with a use case in ChIP-seq, a method that combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify DNA-protein binding sites [1].</title>
    </sec>
    <sec id="sec-6">
      <title>EGO is open and freely available https://github.com/EGO-ontology/EGO. at</title>
    </sec>
    <sec id="sec-7">
      <title>GitHub:</title>
      <sec id="sec-7-1">
        <title>B. EGO applications</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Beyond the ChIP-seq use case illustration, different EGO applications were identified in this study. III. RESULTS</title>
      <p>continuant</p>
      <p>(BFO)
entity
(BFO)
occurrent
(BFO)
realizable
entity
(BFO)
immaterial
entity
(BFO)
material
entity
(BFO)
process
(BFO)
disposition
(BFO)
site
(BFO)
cell
(CL)
genomic
location
(EGO)
cell in
vitro
disease
(DOID)
molecular
function (GO)
human
genomic
location
(EGO)
cultured
cell
protein
binding
(GO)</p>
      <p>human
chromosome
I6 location
(EGO)</p>
      <sec id="sec-8-1">
        <title>FOXM1 binding</title>
        <p>in U2OS cell
(EGO)
position
of human
Chr I6
(EGO)
cell line
cell (CLO)</p>
        <p>U2OS
cell (CLO)
epigenome
(EGO)</p>
        <p>eukayotic
epigenome (EGO)
human epigenome
(EGO)
organism ‘has part’ (protein ‘has DNA template’ gene)
biological
process
(GO)
planned
process (OBI)
cellular
process
(GO)
assay
(OBI)
chromatin
organization</p>
        <p>(GO)
sequencing
assay (OBI)</p>
        <p>histone
modification</p>
        <p>(GO)</p>
      </sec>
      <sec id="sec-8-2">
        <title>ChIP-seq assay (OBI)</title>
        <p>Current prototype EGO includes &gt;600 terms. EGO is
aligned with the BFO (http://ifomis.uni-saarland.de/bfo/) (Fig.
1). EGO also imports many terms and relations from existing
ontologies (Fig. 1). EGO-specific terms focus on terms in the
domain of epigenome. EGO defines ‘epigenome’ as a material
entity that is made up of chemical compounds and proteins that
can attach to DNA and direct various actions (e.g., turning
genes on or off) (https://www.genome.gov/27532724). Since
EGO targets genetic intervals at base-pair resolution, we will
represent different positions of human and
chromosomes and the events that occur at these positions.
mouse
B. EGO ontology design pattern with example</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>EGO is aimed to catalog and organize a large amount of</title>
      <p>
        high quality, genome-wide profiling data and label every base
with observed “events” denoted as “EGO terms” such as
ChIPseq read coverage for a transcription factor (TF) in a specific
cell line. Fig. 2 illustrates an example design of how EGO
represents ChiP-seq read results for human Forkhead TF
FOXM1 binding in the cell line cell U2OS, which occurs at
different chromosomal positions of the cell line genome DNA
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For example, FOXM1 binds at a promoter region of gene
PLK1 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] (Fig. 2). FOXM1 is represented with the Protein
Ontology (PR), and its corresponding gene is represented by
the Ontology of Genes and Genomes (OGG). EGO marks the
chromosome positions where FOXM1 binding is indicated by
the landing of a ChIP-seq read in the cell line. EGO includes
shortcut relations such as ‘function in cell’ and ‘binding at
position’ to lay out semantic axioms and make queries (Fig. 2).
material entity
(BFO)
is a
protein (PR)
is a
FOXM1
(PR) is a
transcription
      </p>
      <p>factor
has_gene_template
gene FOXM1
(OGG)
function
of</p>
      <p>pbr(ioOncdBeinIs)gs ouhtapsut prcoot(emGinOp-lD)exNA
inhpaust is_realized_in
transcription molecular
factor binding is a function (GO)
(GO) functioisna(BFO)
is a
FOXM1
binding
(EGO)
is a
gene
(OGG)</p>
      <p>ChiP-seq
assay (OBI)</p>
      <p>has
is about specified output</p>
      <p>ChiP-seq is a data item
reads (OBI)
reads on position
isfuanctFioOinnXUi(nME2c1GOebOSllin)(csdehilnlogrtcub(sptin)hodosirintticgounat)t PisCohaosrfs2Ihi63tue5(mE9(Ba7GFn6OO3)7) paorft gseprietnreogeismi(oPBanoLFtoeKOfr1),</p>
      <p>Ul2i(nCOeLScOec)lelll is a ceclell(lClinLeO) is a cell (CL)</p>
      <sec id="sec-9-1">
        <title>C. EGO applications</title>
        <p>
          With the EGO support, we will be able to leverage dense,
genome-wide annotations of various types to facilitate
enrichment analysis of “EGO terms” for any give genomic
region(s), similar to the GO term enrichment analyses or gene
set enrichment analysis (GSEA). As an illustration, suppose a
researcher conducted ATAC-seq experiments [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] on primary
tumor cells of prostate cancer as well as normal surrounding
tissues. She is interested in finding out in the thousands of
genomic loci that harbor a peak in the tumor cells but not in the
normal cells, what transcription factor (TF) or histone marks
showed elevated in vivo binding in relevant cell types (such as
CLO_0000748: human prostate gland-derived cell lines
http://purl.obolibrary.org/obo/CLO_0000748). EGO
enrichment analysis will return a ranked list of all EGO terms
satisfied with the cell type constraint and highlight the EGO
terms that show significant enrichment. In another example,
suppose a researcher has identified two lists of genes that show
differential expression between psoriasis skin tissue and
adjacent normal tissue (one up-regulated and one
downregulated). She can then collect the promoter region (10 kb
upstream to 10 kb downstream of the transcription start site)
for the two sets of genes, and ask the question which TF or TFs
are bind preferentially to each set of these regions. Such
information will inform whether different TFs are likely to be
responsible for turning on and off these genes during the
process of disease pathogenesis.
        </p>
        <p>Such an enrichment analysis enables researchers to learn
the key properties of a set of genomic regions, collectively,
avoid the distraction of noise and diversity that are found
common place in set of regions identified through
highthroughput experiments. A key advantage of using an ontology
system is that annotation is provided at different level of
granularity, beyond individual dataset level. For example,
instead of querying individual TFs in which many are highly
related, EGO analysis enables query against TF families which
provide concise and clear interpretation. Similarly, often times
it is highly desirable to compare the binding patters of a factor
across a handful of cell type classes over many closely-related
individual cell types.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>IV. DISCUSSION</title>
    </sec>
    <sec id="sec-11">
      <title>EGO aims to represent the terms and complex relations in</title>
      <p>the domain of epigenome. With EGO, we are hoping to shed
light on the less understood, “dark matter” regions of the
genomes while producing insights, enhancing interpretation,
and generating new hypotheses for genomic regions of interest.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Quaas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stutchbury</surname>
          </string-name>
          , et al.,
          <article-title>"The forkhead transcription factor FOXM1 controls cell cycledependent gene expression through an atypical chromatin binding mechanism,"</article-title>
          <source>Mol Cell Biol</source>
          , vol.
          <volume>33</volume>
          , pp.
          <fpage>227</fpage>
          -
          <lpage>36</lpage>
          ,
          <year>Jan 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Buenrostro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. G.</given-names>
            <surname>Giresi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Zaba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Greenleaf</surname>
          </string-name>
          ,
          <article-title>"Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position,"</article-title>
          <source>Nat Methods</source>
          , vol.
          <volume>10</volume>
          , pp.
          <fpage>1213</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>Dec 2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>