<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Profiling of Semantically Annotated Proteins</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hollunder J</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mironov V</string-name>
          <email>vladimir.n.mironov@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antezana E</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hoehndorf R</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kuiper M</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Biology, Norwegian University of Science and Technology (NTNU)</institution>
          ,
          <addr-line>Høgskoleringen 5, N-7491 Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Physiology, Development and Neuroscience, University of Cambridge</institution>
          ,
          <addr-line>Downing Street, Cambridge CB2 3EG</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Plant Systems Biology, Flanders Institute for Biotechnology and Department of Plant Biotechnology and Bioinformatics, Ghent University</institution>
          ,
          <addr-line>Technologiepark 927, B-9052 Ghent</addr-line>
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We have exploited semantic annotations of biological entities to develop a novel approach to infer new knowledge. We demonstrate this in four use cases based on the Gene Expression Ontology, an applied ontology that we developed to serve the needs of researchers involved in the analysis of genes and proteins implicated in transcriptional control of pathways/diseases. We have found that semantic annotations associated with biological entities in various commonly used data sources support the identification of related entities, thereby emulating associations that can be inferred from sequence or other structural similarities between these entities. We demonstrate how those semantic annotations can be used to make inferences about the respective biological entities.</p>
      </abstract>
      <kwd-group>
        <kwd>ontology</kwd>
        <kwd>annotation</kwd>
        <kwd>semantic similarity</kwd>
        <kwd>gene expression</kwd>
        <kwd>pattern identification</kwd>
        <kwd>hypothesis generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Gene Expression Ontology (GeXO) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is an application ontology that integrates
fragments of GO and the Molecular Interaction ontology (MI) with data from GOA,
IntAct, KEGG, SwissProt, and NCBI Gene. It also includes information on predicted
orthology relations among the proteins. The knowledge in GeXO covers three
biological species: human, mouse, and rat. GeXO comprises 168,417 terms of which 39,680
correspond to proteins. In the present study we attempted to assess the global implicit
informational value contained in GeXO.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Semantic profiles of protein terms</title>
      <p>Protein features were extracted from GeXO in the form 'predicate-object.' The
features form a matrix with 39680 rows corresponding to proteins and 132360 columns
to features. The types of features we used are summarized in Table 1.</p>
      <p>
        This feature matrix was used to compute semantic similarity among all the proteins
in the data set on the basis of the Jaccard index weighted by the information content
as described in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We evaluated the quality of the computed semantic similarities
using ROC analysis.
      </p>
      <p>For classifying false and true positives we used KEGG clusters of orthologous
proteins as positive sets. The KEGG cluster annotations were removed from the data set
prior to the analysis. The results in Figure 1 demonstrate the very high predictive
value of semantic annotations (the results with the full data are given just as a reference).
To exclude an impact of sequence information on the analysis, we removed orthology
information from the data set, which affected the results only slightly. We concluded
that semantic annotations are able to reveal protein similarity even in the absence of
sequence information.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Patterns in semantic profiles</title>
      <p>
        To identify recurrent patterns in the data, we used the DASS tool [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which finds
closed sets in the data. Closed sets have the property that there is no superset that
occurs more frequently in the data set. A data set consists of a set of sets of elements
(referred to as host sets).
      </p>
      <p>Figure 2 provides the distribution of closed sets according to the size (number of
elements), or frequency (number of occurrence in the data set). The vast majority of
closed sets fall within a narrow range of the size and frequency. The sets of low
frequency are likely to be highly predictive due to their specificity. To have a more
precise view on the predictive value of the set we focused on KEGG clusters, which
define functionally distinct protein types, associated with closed sets. Figure 2 gives the
distribution of closed sets according to the number of associated clusters. The highest
number of sets was found to be associated with a single KEGG cluster, thus
confirming the high predictive value of the closed sets. It is worth noting the very high
number of closed sets without any associated KEGG cluster. In combination with the
results of the ROC analysis this suggests that the closed sets could be used to classify
the proteins associated to those closed sets.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Use cases</title>
      <p>To demonstrate the high predictive value of the closed sets, we extracted a subset
containing 106 transcription factors (TFs) of 40 distinct types known or suspected to
be involved in the response to the hormon gastrin.</p>
      <p>The 106 TFs were subjected to clustering with a number of approaches on the basis
of associated closed sets. The resulting clusters were used as templates for screening
the total GeXO data set to identify hypothetical TFs and target genes (TGs). For an
initial validation of the identified candidates, we downloaded additional information
from UniProtKB (lookup for the term: gastric in http://uniprot.org) and mapped it on
the identified candidates. We identified more than 1700 potential candidates including
more than 460 genes with transcriptional activity (non-deep analysis, automated
screening by extracting information from http://www.uniprot.org/uniprot/ with a
customized Python script). Furthermore, we identified 53 known TGs and 28 genes
linked to the term gastric, whereas 11 TGs and 7 gastric genes occur in more than two
clustering solutions and represent (partially) supported hypotheses. Thus, these results
show that we can use the closed sets concept for predicting TGs and regulators
involved in response to gastrin. Additionally, we identified more than 400 novel
candidates occurring in more than two clustering solutions. Evidently, not all of these
candidates are directly involved in this response, but they represent a good basis for
further (more detailed) analyses as well as possible wet-lab experiments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Venkatesan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Towards an integrated knowledge system for capturing gene expression events</article-title>
          .
          <source>ICBO, 3rd International Conference on Biomedical Ontology</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hoehndorf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonfield</surname>
            ,
            <given-names>P. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkoutos</surname>
            ,
            <given-names>G. V.</given-names>
          </string-name>
          :
          <article-title>PhenomeNET: a whole-phenome approach to disease gene discovery</article-title>
          .
          <source>Nucl. Acids Research</source>
          ,
          <volume>39</volume>
          , e119 (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hollunder</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beyer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Workman</surname>
            ,
            <given-names>C. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilhelm</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>DASS: efficient discovery and p-value calculation of substructures in unordered data</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>23</volume>
          ,
          <issue>77</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>