<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards an integrated knowledge system for capturing gene expression events</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aravind Venkatesan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Mironov</string-name>
          <email>mironov@nt.ntnu.no</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Kuiper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biology, NTNU</institution>
          ,
          <addr-line>7491 Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Transcriptional regulation of gene expression is an important mechanism in many biological processes. Aberrations in this mechanism have been implicated in cancer and other diseases. Effective investigation of gene expression mechanisms requires a system-wide integration and assessment of all available knowledge of the underlying molecular networks. This calls for a method that effectively manages and integrates the available data. We have built a semantic web based knowledge system that constitutes a significant step in this direction: the Gene Expression Knowledge Base (GeXKB). The GeXKB encompasses three application ontologies: the Gene Expression Ontology (GeXO), the Regulation of Gene Expression Ontology (ReXO), and the Regulation of Transcription Ontology (ReTO). These three ontologies, respectively, integrate gene expression information that is increasingly more specific, yet decreasing in coverage, from a variety of sources. The system is capable of answering complex biological questions with respect to gene expression and in this way facilitates the formulation or assessment of new hypothesis. Here we discuss the architecture of these ontologies and the data integration process and provide examples demonstrating the utility thereof. The knowledge base is freely available for download and can be queried through a SPARQL endpoint (http://www.semanticsystems-biology.org/apo/).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Research in the Life Sciences is supported by a plethora of
databases (see overview at www.pathguide.org). Moreover,
the continuing advancements in functional genomics
technologies make it possible to create an overwhelming
amount of data in a single experiment. The many
hypotheses that can be derived from such experiments must
be assessed against a multitude of information and
knowledge bases, often represented in a variety of formats.
Scientists therefore become increasingly dependent on
sophisticated computer technologies to integrate and
manage all the available information. Furthermore, the
drastic increase in the available information and a lack of
adhering to accepted formal representations across all
disparate knowledge bases allows only a fraction of the
knowledge to be easily considered in the analysis of new
data, or causes a user to query many databases individually,
sometimes even without the support of ontology terms that
would warrant a common semantics of queries in different
databases. As discussed by
        <xref ref-type="bibr" rid="ref2">Antezana et al. (2009)</xref>
        ,
application ontologies can facilitate the query process itself
as the ontology ensures a uniform semantics across all data.
1.1
      </p>
      <p>Need for an integrated resource that captures
gene expression knowledge</p>
      <sec id="sec-1-1">
        <title>Transcriptional gene expression and its regulation depend</title>
        <p>on a large variety of cellular processes that control the
timing and level of transcription of an individual gene, often in
a cell- or condition specific manner. Regulation of the
expression of protein coding genes is extensively studied.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Gene expression falls into two main phases, i.e. transcrip</title>
        <p>tion and translation. During the process of transcription,
proteins called transcription factors bind to specific DNA
sequence motifs (binding sites) of a gene, playing a key role
in initiating or inhibiting the formation of an active RNA
Polymerase II transcription complex. Active transcription
produces pre-mRNAs which are subsequently processed
(removal of introns, and polyadenylation of the transcript)
upon which mature mRNAs are transported from the
nucleus to the cytoplasm where the mRNA is translated into a
protein. Regulatory processes of gene expression occur at
different levels, enabling the cell to adapt to different
conditions by controlling its structure and function. Furthermore,
the process of gene expression may also be influenced at the
epigenetic level, where nucleotide or protein modifications
can cause heritable changes in expression of otherwise
identical gene sequences. Abnormalities in the regulation of
gene expression can cause diseases such as the occurrence
of malignant cell proliferation.</p>
      </sec>
      <sec id="sec-1-3">
        <title>The knowledge required to decipher the various processes involved in gene expression continues to grow. However,</title>
        <p>
          for a systems-wide understanding of gene regulation, there
is a need for efficiently capturing knowledge of this domain
in its entirety and to further facilitate efficient querying of
this data. For instance, the complex one-to-many
relationships of a transcription factor like Myc includes thousands
of target genes, representing a wide variety of functions and
processes. An ontology-driven approach would best solve
the issue of knowledge querying, representation and
management. Previously, attempts have been made to model the
gene regulation process; resulting in the Gene Regulation
Ontology (GRO)
          <xref ref-type="bibr" rid="ref5">(Beisswanger et al., 2008)</xref>
          . GRO provides
a conceptual model to represent common knowledge about
the gene regulation domain. However, it was primarily built
as a scaffold for knowledge intensive natural language
processing (NLP) tasks and lacks the granularity in concepts
much needed for advanced querying and hypothesis
generation.
        </p>
      </sec>
      <sec id="sec-1-4">
        <title>We have built a system that integrates existing ontologies</title>
        <p>relevant for the domain of gene expression to support the
discovery of new scientific knowledge. We have named this
knowledge system: the Gene Expression Knowledge Base
(GeXKB). This system is conceived as part of the Semantic</p>
      </sec>
      <sec id="sec-1-5">
        <title>Systems Biology (SSB) (http://www.semantic-systems</title>
        <p>biology.org) initiative and comprises at the current stage
three application ontologies that capture the knowledge
about gene expression, namely the Gene Expression</p>
      </sec>
      <sec id="sec-1-6">
        <title>Ontology (GeXO), Regulation of Gene Expression</title>
      </sec>
      <sec id="sec-1-7">
        <title>Ontology (ReXO) and the Regulation of Transcription</title>
      </sec>
      <sec id="sec-1-8">
        <title>Ontology (ReTO).</title>
        <p>2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>GEXKB OBJECTIVES</title>
    </sec>
    <sec id="sec-3">
      <title>PRINCIPLES AND</title>
    </sec>
    <sec id="sec-4">
      <title>DESIGN</title>
      <p>GeXKB is designed to provide the molecular biologist with
a knowledge system that captures knowledge on a variety of
aspects of the gene expression process. To this end it should
be able to provide answers to questions like:
• ‘Which are the proteins that act as chromatin
remodeling proteins and as modulators of
transcription factor activity?’
• ‘Which are the proteins that participate in two
successive regulatory pathways?’.
• ‘Which are the transcription factors (Human)
that are located in the cytoplasm?’.</p>
      <p>The following design principles were followed in the
process of GeXKB development:
• 'is a' completeness
• 'all-some' semantics
• only classes used for modelling of the domain of
discourse (see Table 1)
• maximal flexibility both for users and for future
extensions
3</p>
    </sec>
    <sec id="sec-5">
      <title>GEXKB ARCHITECTURE</title>
    </sec>
    <sec id="sec-6">
      <title>CONSTRUCTION</title>
      <sec id="sec-6-1">
        <title>The core of the three ontologies is built of terms from a</title>
        <p>
          number of well established biomedical ontologies, first of
all GO
          <xref ref-type="bibr" rid="ref3">(Ashburner et al., 2000)</xref>
          and Molecular Interactions
ontology
          <xref ref-type="bibr" rid="ref11 ref12">(Kerrien et al., 2007)</xref>
          , The core is used to integrate
data from GOA
          <xref ref-type="bibr" rid="ref4">(Barrell et al., 2009)</xref>
          , IntAct database
          <xref ref-type="bibr" rid="ref11 ref12">(Kerrien et al., 2007)</xref>
          , KEGG
          <xref ref-type="bibr" rid="ref10">(Kanehisa and Goto, 2000)</xref>
          ,
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>UniProtKB (Magrane and Uniprot consortium, 2011) and</title>
      </sec>
      <sec id="sec-6-3">
        <title>NCBI Gene (Wheeler et al., 2005). In the subsequent sections we describe the architecture and the main features of the ontologies.</title>
        <p>3.1</p>
        <p>Data integration pipeline</p>
      </sec>
      <sec id="sec-6-4">
        <title>The ontologies were built using an automated pipeline implemented with the use of the library ONTO-PERL (Antezana et al., 2008).</title>
        <p>AND</p>
        <sec id="sec-6-4-1">
          <title>3.1.1 Building seed ontologies:</title>
        </sec>
      </sec>
      <sec id="sec-6-5">
        <title>GeXO, ReXO and ReTO share a common Upper Level</title>
      </sec>
      <sec id="sec-6-6">
        <title>Ontology (ULO), which provides a general scaffold for data</title>
        <p>integration. It was developed on the basis of the Science</p>
      </sec>
      <sec id="sec-6-7">
        <title>Integrated Ontology (SIO) (http://code.google.com/p/seman</title>
        <p>ticscience/wiki/SIO) with the addition of few terms from
other ontologies. The origin of the terms is preserved in
external references. The ULO is generated on the fly by the
pipeline and does not exist as an individual artifact. The
upper level term IDs are of the form ‘SSB:nnnnnnn’.</p>
      </sec>
      <sec id="sec-6-8">
        <title>The ULO is then merged with GO (domain specific</title>
        <p>fragments of Biological Process, complete Cellular</p>
      </sec>
      <sec id="sec-6-9">
        <title>Component, complete Molecular Function), MI ('interaction</title>
        <p>type' branch), and the Biorel ontology (Blondé et al. 2011).</p>
      </sec>
      <sec id="sec-6-10">
        <title>This yields three ontologies referred to as seed ontologies.</title>
      </sec>
      <sec id="sec-6-11">
        <title>To be more specific, in order to build the seed ontology for</title>
        <p>GeXO, the term ‘gene expression’ (GO:0010467) and all its
descendants are imported. For ReXO and ReTO the
corresponding GO terms are: 'regulation of gene expression'
(GO:0010468) and 'regulation of transcription, DNA
dependent' (GO:0006355). We refer to these three terms as
sub-roots. Each of them is connected to the ULO as a
subclass of 'biological process'. To ensure 'is a'
completeness, each of the ontologies is complemented with
an auxiliary term - (‘gene expression process’
(GeXO:0000001), 'process of regulation of gene expression’
(ReXO:0000001), ‘process of regulation of DNA-dependent
transcription’ (ReTO:0000001)), which becomes the parent
of all the terms that did not have an 'is a' path to the
subroot. Apart from this, the three seed ontologies are
structurally identical (Figure 1).
3.1.2 Building species specific intermediate ontologies:</p>
      </sec>
      <sec id="sec-6-12">
        <title>The GeXKB ontologies support three model organisms:</title>
        <p>Homo sapiens, Mus musculus and Rattus norvegicus.
The corresponding three species-specific intermediate
ontologies were developed in the following steps:
(1) For each species GOA annotations are used to
extract all the associations involving domain
specific Biological Process terms incorporated in
the previous phase. The corresponding proteins are
added as child terms to the upper level term
‘protein’ (SSB:0001211) and referred to as 'core
proteins' hereafter.
(2) From the IntAct database all the interactions
involving at least one of the core proteins are
retrieved and incorporated into the knowledge base
along with their pertinent information. This results
in a further extension of the set of proteins in the
KB.</p>
        <sec id="sec-6-12-1">
          <title>3.1.3 Building the complete ontologies:</title>
          <p>This is the final phase in the generation of the ontologies
which proceed as follows:
(1) The species specific ontologies (from the previous
step) are merged together.
(2) From the KEGG database all the pathways
involving at least one of the core proteins are
extracted and incorporated in the KB along with the
pertinent information. The pathway terms become
children of the term 'SSB:0011221' ( 'pathway',
'BioPAX:Pathway'). The corresponding KEGG
orthology groups are incorporated as children of the
term 'protein cluster' (SSB:0001122). This step
results in a second extension of the set of proteins.
(3) Putative orthology relationships were computed
with the use of the high-performance library</p>
        </sec>
      </sec>
      <sec id="sec-6-13">
        <title>TurboOrtho (Ekseth et al., 2010), a multi-threaded</title>
        <p>
          C++ implementation of the OrthoMCL algorithm
          <xref ref-type="bibr" rid="ref13">(Li et al., 2003)</xref>
          . The relations including core
proteins are added to the KB, leading to the final
extension of the set of proteins.
(4) The set of proteins in the GeXKB was finally
augmented with:
•
•
•
        </p>
      </sec>
      <sec id="sec-6-14">
        <title>GOA annotations for Cellular Components and</title>
      </sec>
      <sec id="sec-6-15">
        <title>Molecular Functions,</title>
      </sec>
      <sec id="sec-6-16">
        <title>Additional information modifications) from UniProtKB, (e.g. protein</title>
      </sec>
      <sec id="sec-6-17">
        <title>The corresponding genes along with the pertinent information from NCBI. The final result is the three ontologies in the OBO (Smith et al., 2007) format.</title>
        <p>3.1.4 Enhancing the utility of the ontologies:
(1) Transitive closures were constructed with the use of
the library ONTO-PERL for the following relation
types: 'is a', 'part of', ‘regulates’.
(2) The ontologies were exported in a number of formats:</p>
        <p>RDF, OWL, XML, and DOT.
(3) The RDF exports were used to populate a triple store,
refer Table 2 (Virtuoso Open Link).</p>
        <p>Ontology
GeXO</p>
      </sec>
      <sec id="sec-6-18">
        <title>The Semantic Web (Berners-Lee and Hendler, 2001) is an</title>
        <p>
          extension of the WWW which aims at building a web of
data accessible both by computers and human beings. This
new technology is increasingly gaining momentum, in
particular in the domain of Life Sciences
          <xref ref-type="bibr" rid="ref2">(Antezana et al.,
2009)</xref>
          .
        </p>
      </sec>
      <sec id="sec-6-19">
        <title>In order to make use of these new technologies, the RDF</title>
        <p>
          versions of the ontologies have been loaded into Open Link
Virtuoso (http://virtuoso.openlinksw.com) and can be
accessed via a SPARQL query page
(http://www.semanticsystems-biology.org/apo/queryingcco/sparql). In contrast to
other Semantic Web formalisms, such as OWL, RDF
enables handling of large amounts of knowledge due to its
simple and flexible syntax, making querying tractable.
However, on the downside the low expressivity of RDF/RDFS
imposes limitations on the inferencing over the knowledge
base. To overcome this limitation,
          <xref ref-type="bibr" rid="ref7">Blondé et al. (2011)</xref>
          have
developed a novel approach for semi-automated reasoning
on RDF stores with the use of the SPARUL update language
(http://www.w3.org/TR/sparql11-update/). This allows for
pre-computing the inferences supported by the store, thus
making implicit knowledge explicit and available for
querying. In order to provide maximum flexibility for querying,
two graphs are available for each of the ontologies - with or
without closures (e.g. GeXO-tc and GeXO, 'tc' standing for
'total closure').
        </p>
      </sec>
      <sec id="sec-6-20">
        <title>The most convincing evidence of the success of the Seman</title>
        <p>
          tic Web is the quick expansion of the Linked Data cloud
          <xref ref-type="bibr" rid="ref14 ref9">(Heath and Bizer, 2011)</xref>
          . In the course of the design of
        </p>
      </sec>
      <sec id="sec-6-21">
        <title>GeXKB a number of decisions were made to facilitate the</title>
        <p>migration of GeXKB eventually to the Linked Data cloud.</p>
      </sec>
      <sec id="sec-6-22">
        <title>For instance, we have re-used original IDs as much as pos</title>
        <p>sible. If the original IDs include a name-space (e.g. GO, MI)
they were adopted without any modifications, otherwise the
IDs were prepended with a name-space (for example UPKB
for UniProtKB or NCBIgn for NCBI Gene), separated by a
colon from the original ID (the colons are replaced with
underscores in the RDF renderings). The re-use of the IDs
benefits as well the users due to faster query execution and
the familiarity of the IDs. Furthermore, in compliance with
the Linked Data recommendations we minted the URIs in
our own common name-space:
http://www.semanticsystems-biology.org/ and have consistently used rdfs:label
properties to aid human readability of the results.</p>
        <p>RDF
graphs
No. of
triples</p>
        <p>GeXO
~3.3
million</p>
        <p>GeXOtc
~23
million</p>
        <p>ReXO
~3
million</p>
        <p>ReXOtc
~19.9
million</p>
        <p>ReTO
~2.8
million</p>
        <p>ReTOtc
~19.1
million
In this section we demonstrate the utility of GeXKB with
the help of a few example SPARQL queries. These queries
are available as a part of a list of sample queries provided on
the query page
(http://www.semantic-systemsbiology.org/apo/queryingcco/sparql). To query GeXKB, the
base URI and the prefixes are set and the SELECT block
specifies the variables to be part of the solution. The RDF
triple pattern queried is defined in the WHERE block. The
queries are as follows:</p>
      </sec>
      <sec id="sec-6-23">
        <title>Q1: (see Table 3)</title>
        <p>Biological question: Which proteins can act as chromatin
remodeling proteins and as modulators of transcription factor activity?
SPARQL query:
BASE &lt;http://www.semantic-systems-biology.org/&gt;
PREFIX rdfs: &lt;http://www.w3.org/2000/01/rdf-schema#&gt;
PREFIX ssb: &lt;SSB#&gt;
PREFIX taxon: &lt;SSB#NCBItx_9606&gt;
PREFIX graph1: &lt;ReXO&gt;
PREFIX graph2: &lt;ReTO-tc&gt;
SELECT distinct ?protein_id ?protein_name
WHERE {
GRAPH graph1: {
? protein_id ssb:is_a ssb:SSB_0001211 .
?b_process ssb:is_a ssb:GO_0040029 .
?b_process ssb:has_participant ? protein_id .
? protein_id ssb:has_source taxon: .
}
GRAPH graph2: {
ssb:GO_0034401 ssb:has_participant ? protein_id .
? protein_id rdfs:label ?protein_name .
}
}
LIMIT 4
Q2:
Biological question: Which proteins participate in both the
JAK/STAT signaling pathway and Apoptosis?
SPARQL query:
BASE &lt;http://www.semantic-systems-biology.org/&gt;
PREFIX rdfs: &lt;http://www.w3.org/2000/01/rdf-schema#&gt;
PREFIX ssb: &lt;SSB#&gt;
PREFIX taxon: &lt;SSB#NCBItx_9606&gt;
PREFIX pathway1: &lt;SSB#KEGG_ko04630&gt;
PREFIX pathway2: &lt;SSB#KEGG_ko04210&gt;
PREFIX graph: &lt;GeXO&gt;
SELECT distinct ?protein
WHERE {
GRAPH graph: {
?prot_id ssb:is_a ssb:SSB_0001211 .
?prot_id ssb:is_member_of ?cluster .
pathway1: ssb:has_agent ?cluster .
?prot_id ssb:has_source taxon: .
}
GRAPH graph: {
?prot_id ssb:is_member_of ?cluster .
pathway2: ssb:has_agent ?cluster .
?prot_id rdfs:label ?protein .
}
}
Q3:
Biological question: Which are the transcription factors (Human)
that are located in the cytoplasm?
SPARQL query:
BASE &lt;http://www.semantic-systems-biology.org/&gt;
PREFIX rdfs: &lt;http://www.w3.org/2000/01/rdf-schema#&gt;
PREFIX ssb: &lt;SSB#&gt;
PREFIX taxon: &lt;SSB#NCBItx_9606&gt;
PREFIX location: &lt;SSB#GO_0005737&gt;
PREFIX graph: &lt;ReTO-tc&gt;
SELECT distinct ?protein ?protein_name
WHERE {
GRAPH graph: {
?protein ssb:is_a ssb:SSB_0001211 .
?protein rdfs:label ?protein_name .
ssb:GO_0006355 ssb:has_participant ?protein .
?protein ssb:has_function ?function .
?function ssb:is_a ssb:GO_0003700 .
location: ssb:contains ?protein .
?protein ssb:has_source taxon: .
}
}
These queries offer just a glimpse of the repertoire of
biological question that can be addressed to the knowledge
system. In addition, users could also query the knowledge
base in combination with other complementary semantic
web resources to formulate advanced queries for hypothesis
generation. This could be performed through the query
federation features that are included in the latest version of</p>
      </sec>
      <sec id="sec-6-24">
        <title>SPARQL (ver. 1.1) and will be explored in the future.</title>
        <p>Protein ID
http://www.semantic-systemsbiology.org/SSB#UPKB_Q9NS37
http://www.semantic-systemsbiology.org/SSB#UPKB_P14373
http://www.semantic-systemsbiology.org/SSB#UPKB_Q62158
http://www.semantic-systemsbiology.org/SSB#UPKB_P17947
Protein Name
ZHANG_HUMAN
TRI27_HUMAN
TRI27_MOUSE
SPI1_HUMAN</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5 CONCLUSION</title>
      <p>The drastic increase in the amount of data generated in the
field of molecular biology and biomedicine requires
efficient knowledge management practices. Ontologies
certainly provide a robust method to integrate data and
efficiently represent specific (sub) domain knowledge. With the
creation of GeXKB, we have built a knowledge system that
specifically supports researchers focusing on various aspects
of gene expression. The three ontologies provide the user
with the flexibility of choosing an ontology depending on
the breadth and specificity of information needed. Further
flexibility is afforded by a range of available formats for
knowledge representation (OBO, RDF, OWL), data
exchange (XML), and visualisation (DOT).</p>
      <p>The presented examples demonstrate the utility of our
knowledge base with respect to answering realistic domain
specific questions, and this utility is expected to grow with
its further development. The primary goal will be to
augment the knowledge base with additional high quality,
curated sources of information with documented transcription
factor function and relations between transcription factors
and their target genes.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGEMENTS</title>
      <sec id="sec-8-1">
        <title>This work is funded by the Norwegian University of Sci</title>
        <p>ence and Technology (NTNU), Norway. AV was funded by</p>
      </sec>
      <sec id="sec-8-2">
        <title>Faculty of Natural Science and Technology and VM was funded by FUGE Mid-Norway.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Egaña</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Baets</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>ONTO-PERL: an API for supporting the development and analysis of bio-ontologies</article-title>
          .
          <source>Bioinformatics. Mar</source>
          <volume>15</volume>
          ;
          <issue>24</issue>
          (
          <issue>6</issue>
          ):
          <fpage>885</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Biological knowledge management: the emerging role of the Semantic Web technologies</article-title>
          . Brief Bioinform.,
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <fpage>392</fpage>
          -
          <lpage>407</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ball</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Botstein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Butler</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolinski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          et al., (
          <year>2000</year>
          ).
          <article-title>Gene ontology: tool for the unification of biology</article-title>
          .
          <source>Nature Genetics</source>
          ,
          <volume>25</volume>
          ,
          <fpage>25</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Barrell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dimmer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huntley</surname>
            ,
            <given-names>R.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Binns</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Donovan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            , and
            <surname>Apweiler</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>The GOA database in 2009--an integrated Gene Ontology Annotation resource</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>37</volume>
          :
          <fpage>D396</fpage>
          -
          <lpage>D403</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Beisswanger</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Splendiani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dameron</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Gene Regulation Ontology (GRO): design principles and use cases</article-title>
          .
          <source>Stud Health Technol Inform</source>
          .
          <year>2008</year>
          ;
          <volume>136</volume>
          :
          <fpage>9</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Hendler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>'Publishing on the semantic web'</article-title>
          .
          <source>Nature</source>
          ,
          <volume>410</volume>
          ,
          <fpage>1023</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Blondé</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkatesan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Baets</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kuiper</surname>
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Reasoning with bio-ontologies: using relational closure rules to enable practical querying</article-title>
          .
          <source>Bioinformatics, Jun</source>
          <volume>1</volume>
          ;
          <issue>27</issue>
          (
          <issue>11</issue>
          ):
          <fpage>1562</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Ekseth</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>TurboOrtho - a high performance alternative to OrthhoMCL</article-title>
          . European Conference on Computaional Biology:
          <year>September 2010</year>
          ; Ghent.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and Bizer, . (
          <year>2011</year>
          )
          <article-title>Linked Data: Evolving the Web into a Global Data Space (1st edition)</article-title>
          .
          <source>Synthesis Lectures on the Semantic Web: Theory and Technology</source>
          ,
          <volume>1</volume>
          :
          <issue>1</issue>
          ,
          <fpage>1</fpage>
          -
          <lpage>136</lpage>
          . Morgan &amp; Claypool.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Kanehisa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Goto</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>KEGG: Kyoto Encyclopedia of Genes and Genomes</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>28</volume>
          :
          <fpage>27</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Kerrien</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orchard</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montecchi-Palazzi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aranda</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quinn</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinod</surname>
            <given-names>N</given-names>
          </string-name>
          , Bader,
          <string-name>
            <given-names>G. D.</given-names>
            ,
            <surname>Xenarios</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          et al. (
          <year>2007</year>
          ).
          <article-title>Broadening the horizon--level 2.5 of the HUPO-PSI format for molecular interactions</article-title>
          .
          <source>BMC Biol</source>
          .
          <volume>9</volume>
          ;
          <issue>5</issue>
          :
          <fpage>44</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Kerrien</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alam-Faruque</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aranda</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bancarz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bridge</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derow</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dimmer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feuermann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al. (
          <year>2007</year>
          ).
          <article-title>IntAct - open source resource for molecular interaction data</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>35</volume>
          :
          <fpage>D561</fpage>
          -
          <lpage>565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoeckert</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <string-name>
            <surname>Jr.</surname>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>D. S.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>OrthoMCL: identification of ortholog groups for eukaryotic genomes</article-title>
          .
          <source>Genome Res</source>
          ,
          <volume>13</volume>
          :
          <fpage>2178</fpage>
          -
          <lpage>2189</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Magrane</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>and the UniProt consortium UniProt Knowledgebase: a hub of integrated protein data Database, 2011</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosse</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bug</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceusters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>L. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eilbeck</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          et al. (
          <year>2007</year>
          ).
          <article-title>The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration</article-title>
          .
          <source>Nat Biotechnol</source>
          ,
          <volume>25</volume>
          (
          <issue>11</issue>
          ),
          <fpage>1251</fpage>
          -
          <lpage>1255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Wheeler</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barrett</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benson</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bryant</surname>
            ,
            <given-names>S. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Canese</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Church</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DiCuccio</surname>
          </string-name>
          , M. et al. (
          <year>2005</year>
          ).
          <article-title>Database resources of the National Center for Biotechnology Information</article-title>
          .
          <source>Nucleic Acids Res</source>
          <year>2005</year>
          ,
          <volume>33</volume>
          :
          <fpage>D39</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>