<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontological interpretation of biomedical database annotations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Filipe Santana da Silva</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ludger Jansen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fred Freitas</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Schulz</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Motivation: In general, the meaning of biological database records is not sufficiently specified from an ontological point of view. We explore the options for an ontology-based integration and interpretation of database content of individuals, defined classes, dispositions and a combination of these. Results: Four interpretation models are created, interpreting annotations in database records as referring to (i) individuals, (ii) defined classes, (iii) disposition universals, and (iv) a combination of these. Evaluation is done by using competency questions to test the retrieval capacities. Availability: Interpretation models and sample data are available at http://www.cin.ufpe.br/~integrativo. * Contact: fss3@cin.ufpe.br</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Biological databases (BIO-DBs) are used to store summarized
results of laboratory experiments. Apart from numeric and textual
entries, they include semantic annotations. E.g., the Unified Protein
Resource (UniProt)
        <xref ref-type="bibr" rid="ref8">(The UniProt Consortium, 2015)</xref>
        includes
annotations from the Protein Ontology (PR)
        <xref ref-type="bibr" rid="ref1 ref2 ref5">(Natale et al., 2014)</xref>
        and the
Gene Ontology (GO)
        <xref ref-type="bibr" rid="ref7">(The Gene Ontology Consortium, 2014)</xref>
        .
While these ontologies, in isolation, obey formal principles and
convey precise meaning, the meaning of the database record as a whole
remains vague and depends on implicit background assumptions.
What it means when, e.g., in an annotation the UniProt protein term
Methionine synthase is linked to the GO process term Methylation,
is left to the user. Hence, on the one hand, we have rich and
wellcurated BIO-DBs with highly structured tabular content, but limited
ontological explicitness. On the other hand, large bio-ontologies
provide formal descriptions of their content, enabling logic-based
reasoning. In order to use these features with BIO-BDs, we want to
make explicit what annotations exactly refer to and to express this
in a formal, computer-processable way.
      </p>
      <p>
        It has already been argued that there are benefits for content
retrieval, regarding correctness, completeness, and user-friendliness
given a seamless integration between BIO-DBs and ontologies, and
that such systems could accommodate large amounts of data from
BIO-DBs
        <xref ref-type="bibr" rid="ref3 ref3 ref5 ref5">(Hoehndorf et al., 2011; Santana et al., 2011)</xref>
        . It is,
however, still an open question (1) how implicit knowledge about the
entities and relationships described in the structure of a BIO-DB be
represented, (2) whether the content denoted by BIO-DBs (i.e. the
domain entities represented by the data elements and the way how
the former are connected) is fit to be represented, and, if this is the
case, (3) how it can be translated into axioms using appropriate
representational patterns, and finally, once database structure and
content are expressed by formal-ontological means, (4) how the existing
bio-ontologies can be plugged into this structure. Addressing these
questions, we demonstrate that there are feasible ways to express
implicit and explicit database content by formal-ontological means
and combine it with existing domain ontologies. We show how
annotation terms used in a typical BIO-DB entry can be interpreted as
referring to entities from different ontological categories. Each of
these interpretations requires different means like the introduction
of individuals, the addition of new axioms to existing classes or the
introduction of additional defined classes. The resulting OWL
models are tested under three aspects: (i) database content retrieval,
using ontologies as query vocabulary for data integration; (ii)
information completeness; and (iii) reasoning behaviour in Description
Logics (DL).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>METHODS</title>
      <p>
        For the analyses, we selected a typical example from biomedical
databases, generated by joining data from UniProt and Ensembl
        <xref ref-type="bibr" rid="ref1 ref2 ref5">(Cunningham et al., 2014)</xref>
        . Records in BIO-DBs are mainly
composed of (i) one protein term (e.g., CBS); (ii) one taxon term (e.g.,
Rattus norvegicus); (iii) one to many terms from GO for biological
processes (e.g., Methylation); (iv) one to many terms from GO for
cellular components (e.g., Cytoplasm); (v) zero to many phenotype
terms (e.g., Endocrine pancreas increased size); and (vi) one to
many small molecules (e.g. Homocysteine). We implement four
different interpretive strategies (IND, SUBC, DISP and HYB) in OWL
using the editor Protégé v.5 and the reasoner FACT++
        <xref ref-type="bibr" rid="ref9">(Tsarkov &amp;
Horrocks, 2006)</xref>
        to check for consistency and taxonomic
subsumption. We used BioTopLite2 (BTL2) as an upper-level ontology with
highly constrained classes and a small set of relations
        <xref ref-type="bibr" rid="ref6">(Schulz &amp;
Boeker, 2013)</xref>
        . To test each interpretation model, we created four
competency questions (CQs), first in natural language, and then
translated into DL queries.
      </p>
      <p>Individuals as the referents of annotations (IND)
The first interpretation rests on the fact that a database entry is about
the outcome of a concrete experiment. Accordingly, the annotations
that feature in such an entry can be interpreted as referring to the
individual molecules, objects and processes that belonged to that
particular experiment. Thus, the entry “Cystathionine gamma-lyase”
denotes a molecule or a collection of molecules of the class
‘Cystathionine gamma-lyase’. BIO-DB content is therefore represented as
a set of Abox-level class-membership assertions and relationships.
3.2</p>
      <p>Subclasses as the referents of annotations (SUBC)
Second, database content can be interpreted by means of a number
of maximally fine-grained defined classes, introduced by means of
equivalence axioms for each universal entity which the annotations
refer to. For instance, the annotations of a record combining the
protein term Methionine synthase and the species term Rattus
norvegicus are represented by a customized defined class combining the
information about a subclass of Methionine synthase, defined as
Methionine synthase that is part of an organism of the type Rattus
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>RESULTS</title>
      <p>norvegicus. Using OWL-EL expressiveness, we can formalize this
as follows:
‘Methionine synthase_in_Rattus Norvegicus’ equivalentTo</p>
      <p>Methionine_Synthase and (‘is part of’ some ‘Rattus norvegicus’)
3.3</p>
      <p>
        Dispositions as the referents of annotations (DISP)
Real world entities are often described scientifically in terms of
dispositions, i.e. tendencies to behave in a certain way under certain
circumstances. Biomedical observations yield statistical results
indicating that participants of an experiment (a protein Methionine
synthase) have dispositions to bear certain capabilities
        <xref ref-type="bibr" rid="ref4">(Jansen,
2007)</xref>
        , like being able to perform a Methylation process. Interpreting
database entries as statements about dispositions means that we
represent the database content regarding a disposition of organisms of
a certain species, e.g., that all instances of Homo sapiens have the
disposition to develop a pathological condition P. For this purpose,
we use General Class Inclusion (GCI) axioms that allow for
subclass assertions between two complex class expressions, e.g.:
‘Endochondral ossification’
and (‘is included in’ some ‘Bos taurus’)
      </p>
      <p>subClassOf ‘has participant’ some ‘Cysthationine beta-synthase’
The output of DISP is an ontology file representing the classes
referred to by the annotations together with a small set of GCIs, using
DL-SHI expressiveness.
3.4</p>
      <sec id="sec-3-1">
        <title>Hybrid interpretation (HYB)</title>
        <p>To avoid the complexity of GCI expressions, we combine SUBC
with DISP. HYB uses subclass statements like SUBC, enriched by
axioms about dispositions like in DISP. This combination reduces
the amount subclasses to be created. Disposition axioms are limited
to material objects like proteins and organisms, asserting that they
are capable of participating in specific biological processes. The
HYB output needs DL-SHI expressiveness.
3.5</p>
      </sec>
      <sec id="sec-3-2">
        <title>Fitness test</title>
        <p>The four ontology models were tested for consistency and the
following queries were used for retrieval evaluation: (Q1) Which
biological processes have proteins of the kind Prot1 as participants?
(Q2) In which cellular locations is Prot2 active in organisms of the
type Org1? (Q3) Which proteins are involved in processes of the type
BProc in organisms of the type Org1? (Q4) Which organisms are
able to exhibit a specific phenotype Phen1? – These queries were
translated into DL, which enabled the retrieval of content in
interpretations IND, SUBC and HYB. The model HYB was the only one
able to retrieve content for Q4. As DISP expresses everything in
GCIs, retrieval is not supported at all.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>DISCUSSION</title>
      <p>
        We proposed four interpretation strategies: IND, SUBC, DISP and
HYB. Of these, only IND is completely based on single individuals
(Abox entities).
        <xref ref-type="bibr" rid="ref1">Ceusters et al. (2014)</xref>
        use a similar approach for
applying relations between individuals in electronic health records.
      </p>
      <p>
        SUBC is based on generating customized definitions of classes.
This approach is not far from the work of Hoehndorf et al.
        <xref ref-type="bibr" rid="ref3 ref5">(Hoehndorf et al., 2011)</xref>
        . However, in SUBC an annotation does not
refer directly to the class matching to the annotation term, but to a
defined subclass of it. This requires a non-standard interpretation of
DL queries, targeting the existence of subclasses. On the downside,
SUBC involves an excessive number of subclasses. However, this
does not have severe consequences on performance because of the
good scaling behaviour of OWL-EL ontologies. This has also been
confirmed by our preliminary experiments.
      </p>
      <p>DISP alone is not helpful for most of the queries. It provides a
more compact representation, but it is also incomplete because not
all knowledge embedded within a database record can be sensibly
expressed by dispositions. The combination of SUBC and DISP in
HYB has finally the huge advantage that it enables querying whether
certain biological entities are capable of participating certain
processes, assuming that we agree that parts of the underlying
knowledge in BIO-DBs is about dispositions.
5</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSION</title>
      <p>We proposed four ontological representations of structure and
content of biological databases. The solutions we presented targeted
aspects of ontology-based database retrieval, expressiveness and
content retrieval based on DL reasoning. Only part of database content
is really of ontological nature in a strict sense, i.e., expressible by
axioms that hold universally for all instances of a class. We
addressed this limitation by three ways. Firstly, we interpreted the
denoted entities as (prototypical) individuals, which requires
representation and reasoning on an Abox level. Secondly, we expressed
contingent database content by creating defined subclasses for which
then universally valid statements could be made. Thirdly, we
interpreted part of the database content as reporting dispositions, which
was, however, not very helpful for the answering of our queries, in
contrast with the second modelling approach, when DL reasoning
was used to check for the existence of subclasses.</p>
      <p>Funding: This work was funded by Conselho Nacional de
Aperfeiçoamento de Pessoal de Nível Superior (CAPES) 3914/2014-03 and
Conselho Nacional de Desenvolvimento Científico e Tecnológico
(CNPq) 140698/2012-4.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ceusters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ).
          <article-title>Clinical Data Wrangling using Ontological Realism and Referent Tracking</article-title>
          . In W. R.
          <string-name>
            <surname>Hogan</surname>
          </string-name>
          , et al. (Eds.),
          <source>ICBO 2014</source>
          (pp.
          <fpage>27</fpage>
          -
          <lpage>32</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al.(
          <year>2014</year>
          ).
          <article-title>Ensembl 2015</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>43</volume>
          (
          <issue>D1</issue>
          ),
          <fpage>D662</fpage>
          -
          <lpage>D669</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Hoehndorf</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , et al. (
          <year>2011</year>
          ).
          <article-title>Integrating systems biology models and biomedical ontologies</article-title>
          .
          <source>BMC Systems Biology</source>
          ,
          <volume>5</volume>
          ,
          <fpage>124</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <source>Tendencies and other Realizables in Medical Information Sciences. The Monist</source>
          ,
          <volume>90</volume>
          (
          <issue>4</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Natale</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          , et al. (
          <year>2014</year>
          ).
          <article-title>Protein Ontology: A controlled structured network of protein entities</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>42</volume>
          (
          <issue>D1</issue>
          ),
          <fpage>D415</fpage>
          -D421
          <string-name>
            <surname>Santana</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al. (
          <year>2011</year>
          ).
          <article-title>Ontology patterns for tabular representations of biomedical knowledge on neglected tropical diseases</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>27</volume>
          (
          <issue>13</issue>
          ),
          <fpage>i349</fpage>
          -
          <lpage>i356</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Boeker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>BioTopLite: An Upper Level Ontology for the Life Sciences. Evolution, Design and Application</article-title>
          . In M. Horbach (Ed.),
          <source>Informatik</source>
          (pp.
          <fpage>1889</fpage>
          -
          <lpage>1899</lpage>
          ). GI.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>The</given-names>
            <surname>Gene Ontology Consortium.</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Gene Ontology Consortium: going forward</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>43</volume>
          (
          <issue>D1</issue>
          ),
          <fpage>D1049</fpage>
          -
          <lpage>D1056</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>The UniProt Consortium.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>UniProt: a hub for protein information</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>43</volume>
          (
          <issue>D1</issue>
          ),
          <fpage>D204</fpage>
          -
          <lpage>D212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Tsarkov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>FaCT ++ Description Logic Reasoner : System Description</article-title>
          . In LNCS (pp.
          <fpage>292</fpage>
          -
          <lpage>297</lpage>
          ). Springer: Berlin/Heidelberg.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>