<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards the semantic standardization of orthology content</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jesualdo Tomas Fernandez-Breis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mar a del Carmen Legaz-Garc a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hirokazu Chiba</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ikuo Uchiyama</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Informatica, Universidad de Murcia, IMIB-Arrixaca</institution>
          ,
          <addr-line>Murcia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute for Basic Biology, National Institute of Natural Sciences</institution>
          ,
          <addr-line>Okazaki</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The amount of resources and data about orthology, together with the increasing interest in orthologs in biomedical research has triggered the need for the shareability of the content generated by the di erent tools and stored in the di erent databases. In recent years OrthoXML permitted to advance in the standardization of the content for exchange within the orthology community, but the interest in exchanging content with other communities suggested the research on the use of Semantic Web languages like RDF and to use ontologies for making the meaning of the content explicit. This possibility was reinforced by the existence of initial e orts for using Semantic Web technologies for representing orthology content. In this work, we describe the advances done with the objective of the semantic standardization of orthology content. The need for a common ontology has permitted to obtain a draft of an orthology ontology built by reusing existing ontologies and following best practices in ontology construction. This ontology should serve as knowledge framework for the semantic standardization of orthology content. A mapping between OrthoXML and this orthology ontology has been designed, speci ed and applied to sample OrthoXML datasets. For this purpose, the Semantic Web Integration Tool (SWIT) has been used, which provides support for doing the previously described actions and permit to create an integrated repository from multiple orthology content sources.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In recent years, the number of genome sequences determined has signi cantly
increased and many on-going research projects will permit to know not only the
reference genome of many organisms but also the genome of individuals. In this
new era, being able to perform computational comparative analysis might bring
many opportunities to biomedical research. Homology information can play a
central role in integrating and comparing multiple genomes, because homology
permit establishing evolutionary relations between genes from multiple species.
Three basic concepts, namely, homologs, orthologs, and paralogs need to be
distinguished in this context[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: (1) homologous genes have diverged from an
ancestral gene; homology may hold for genes of the same species or from di erent
species; (2) orthologous genes have diverged by speciation from an ancestral gene,
and their biological functions are usually conserved; orthology may hold between
genes from di erent species; (3) paralogous genes have diverged by duplication
from an ancestral gene; paralogy may hold between genes of the same species or
from di erent species.
      </p>
      <p>Despite homology relations are usually obtained in particular studies from
a pairwise perspective, they are usually calculated and represented as clusters,
that is, groups of genes holding homologous relationship. Homolog clusters are
further divided into two types: ortholog groups consist of genes derived from a
speciation event and paralog groups consist of genes derived from a duplication
event. Thus, ortholog and paralog groups can be represented in a form of nested
hierarchy according to their evolutionary history, where each orthologous group
is associated with a taxonomic range that corresponds to a set of organisms
derived from a speci c speciation event.</p>
      <p>
        Ortholog information is a useful resource to link the corresponding genes
of di erent species and transfer the biological knowledge of model organisms
to organisms with newly sequenced genomes. In addition, ortholog groups are a
vital resource for the comparative analysis of multiple genomes, and they provide
a basis for the analysis of phylogenetic pro les [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        The Quest for Orthologs (QfO) Consortium has identi ed more than forty
resources about orthology (http://questfororthologs.org/orthology databases),
which represent di erent needs of information management in the orthology
eld. Many of these resources store information about prediction of gene
evolutionary relations and there is a diversity of objectives for these databases.
There is heterogeneity in how data are stored and shared by these resources.
For example, Inparanoid [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] stores orthology relations between two species,
whereas OrthoMCL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and MBGD [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] stores ortholog groups among
multiple genomes. OMA [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provides various types of orthology relations including
pairwise orthologs and hierarchical ortholog groups. Many resources use their
own representation format based on tabular les, despite this community has
developed the OrthoXML and SeqXML [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] formats to standardize the
representation of orthology data. OrthoXML permits the comparison and integration
of orthology data from di erent resources within the orthology community.
      </p>
      <p>
        In recent years, semantic web formats have been used for representing
orthology data. OGO [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] was created with the purpose of providing an integrated
resource of information about genetic human diseases and orthologous genes.
OGO integrated information from orthology databases such as Inparanoid [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
OrthoMCL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and OMIM [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This resource developed an OWL ontology for
representing the domain knowledge. More recently, RDF has been used by
sharing the content of the Microbial Genome Database for Comparative Analysis
(MBGD) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] database. This resource also developed an OWL ontology for
representing the domain knowledge, called OrthO, which had similar concepts to the
OGO one, despite being developed independently.
      </p>
      <p>
        The report of the 2013 QfO meeting [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] identi ed a series of aspects about
semantics that have been the key drivers of our activities: (1) the orthology
community should use shared ontologies to facilitate data sharing; (2) exploiting
automated reasoning should be bene cial for the QfO consortium. In this work,
we describe the development of a draft on an orthology ontology, built from
existing related ontologies and we explain how we are approaching the transformation
and exploitation of existing orthology datasets using semantic web technologies.
We believe this work provides a step forward towards the standardization in the
orthology community.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Orthology Ontology</title>
      <p>In this section, we describe the development of the orthology ontology. The
design of the ontology has followed two basic principles: (1) reusing content
from existing ontologies to facilitate interoperability across biomedical domains:
(2) modelling based on membership of genes to clusters of homologs, orthologs
or paralogs, so pairwise relations are obtained by analysing the structure and
content of the dataset.
2.1</p>
      <sec id="sec-2-1">
        <title>Related ontologies</title>
        <p>
          We studied how orthology related concepts and properties were covered in
potentially related ontologies. Our initial selection contained the following
ontologies: Comparative Data Analysis Ontology (CDAO) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], Relations Ontology
(RO) [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], Semanticscience Integrated Ontology (SIO) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], Homology Ontology
(HO) [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], National Cancer Institute Thesaurus (NCIt) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], and Clusters of
Orthologous Groups (COG) Analysis Ontology [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It should be noted that these
ontologies present some overlaps and, in such cases, we inspected the axioms
and the textual content in order to make our decisions. For example, HO, NCIt
and COG were not nally reused, since we found the other ontologies more
appropriate. The CDAO provides a series of classes of interest that we decided to
reuse for phylogenetic purposes. It contains a taxonomic module that de nes
types of Trees, which we found useful for de ning phylogenetic trees. It also
provides a taxonomic module of hereditary changes, which we can use in order to
de ne concepts like orthologs or paralogs. The RO and the SIO include a series
of properties of interest for our domain like has part, is part of, in taxon and
many evolutionary relations. It seems that the recent versions of the RO include
most classes of the HO as properties. The SIO also de nes types of genes and
concepts like databases in a way that can be e ectively re-used for our purpose,
so we also selected it. We also reused the NCBI taxonomy for representing the
species.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Description of the orthology ontology</title>
        <p>The orthology ontology has been implemented in OWL using Protege and it is
currently available at https://github.com/qfo/OrthologyOntology. Figure 1
represents an excerpt of this ontology. There, classes are represented as boxes and
properties as arrows. The entities without pre x are de ned in our orthology
ontology. The pre xes cdao, sio and ro represent entities reused from the
corresponding ontologies. We can see HomologsCluster is a subclass of GeneTreeNode,
and it has two descendants, namely, OrthologsCluster and ParalogsCluster, which
are associated with cdao:speciation and cdao:geneDuplication respectively.
GeneTreeNode is a subclass of cdao:Node, which is not shown in the gure. The
members of HomologsCluster are instances of GeneTreeNode, which means that
they can be clusters of homologs or sequence units. The class SequenceUnit
has three subclasses, Gene, Subgene and Protein. The membership property is
hasHomologous, which is a subproperty of sio:has part. We use this property
instead of two hasOrthologous and hasParalogous, because the pairwise relations
are obtained by analysing the tree. Finally, genes and proteins are linked to
ncbi:organisms through the property ro:in taxon.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Semantic Transformation of Orthology Data</title>
      <p>The availability of the orthology ontology enables the possibility of using it as
schema for the generation of RDF datasets from existing orthology resources.
As it has been previously mentioned, OrthoXML has been a format proposed
for the exchange of orthology content. Hence, we have de ned a mapping
between OrthoXML and the orthology ontology, whose main entities can be seen
in Table 1. The rst four rows show mappings corresponding to entities in
the ontology. The rest are examples of mappings to properties in the
ontology. For instance orthologGroup/paralogGroup means that a group of paralogs
has been de ned in an group of orthologs and in the ontology this is represented
through the property hasHomologous. The complete mapping le is available at
https://github.com/qfo/OrthologyOntology.</p>
      <p>
        One of the technical objectives of the work is to be able to use tools available
for supporting the di erent processes, so the managers of orthology resources
do not need to create their own transformation scripts into semantic formats.
To this end, we have used the SWIT tool (http://sele.inf.um.es/swit), which
was used in our research group to develop and maintain the OGO Linked Open
Dataset [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. SWIT provides a web interface through which the user is guided to
perform all the steps of the process:
1. Inputs and outputs: the user has to provide the OrthoXML schema and the
orthology ontology, as well as the corresponding datasets to be transformed.
SWIT is able to generate the dataset in OWL or RDF formats, which can
be downloaded or directly stored into a triple store like Virtuoso. This tool
also enables the user to de ne the structure of the URI for the RDF/OWL
instances (by default, the ontology base URI).
2. De nition of the mappings: the user can de ne the mappings between
OrthoXML and the orthology ontology. Alternatively, mappings created in
previous sessions can be uploaded and re-used. In this case, the le containing
the mappings between OrthoXML and the orthology ontology described in
the previous section was uploaded.
3. Once the mappings have been de ned, they can be executed to generate the
corresponding repository.
      </p>
      <p>
        SWIT applies the mapping rules to the data source to generate the
semantic content. Brie y speaking, the mapping rules associate entities and
attributes of the OrthoXML schema with owl:Class, owl:DatatypeProperty and
owl:ObjectProperty de ned in the ontology. SWIT also generates integrated
repositories from multiple sources. For supporting the integration process, SWIT
permits the de nition of identity rules, which aim to avoid the creation of
redundant individuals in the semantic repository. Automated reasoning ensures that
only logically consistent content is transformed. To this end, SWIT is currently
using Hermit [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] as reasoner.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Calculating Pairwise Relationships</title>
      <p>In this section we focus on how we can obtain pairwise relationships between
genes in a given tree. Figure 2 will be used to illustrate some concepts. In the next
subsections, we will describe how we can obtain them from the RDF repository.
We can de ne pairwise orthologs as pairs of genes from di erent species whose
common ancestor has a speciation event associated. If we inspect Figure 2, X1
has two ancestral nodes associated with speciation events: one is the common
ancestor of gene X1 and gene Y1, and the other is the common ancestor of gene
X1 and gene Z1. Therefore, we can say that Y1 and Z1 are orthologs of X1,
so the pairs (Y1, X1) and (Z1, X1). Table 2 shows the query for that purpose.
It can be explained as the search for ?common ancestor, which is a cluster of
orthologs (so it has a speciation event associated) and is the common ancestor
of two nodes of trees (?tree node1 and ?tree node2 ) to which two genes (?gene1
and ?gene2 ) from di erent species belong. We have omitted the taxonomic range
in the query for simplicity.
We can de ne pairwise paralogs as pairs of genes whose common ancestor has
a duplication event associated. This means that the genes may be found in the
same species or in di erent ones. If we inspect Figure 2, X2 and Y2 are paralogs to
X1, since their common ancestor have a duplication event associated. Therefore,
a similar query to the previous one for orthologs can work for this purpose
by introducing two changes: use of ParalogsCluster instead of OrthologsCluster
and not using the condition for species in the FILTER clause. This part is not
shown. However, this query cannot work for Inparanoid in which in-paralogs
are included in OrthologsCluster and not de ned as ParalogsCluster. Thus, in
this case, another query (see Table 3) is necessary to search for members of
OrthologsCluster from the same species. In fact, many orthology resources use
at structure of ortholog clusters where all member genes are directly associated
with the top-level speciation node. In such cases, although the same query as in
Table 3 can be used to list all pairwise intra-species inparalogy relationships, it
is impossible to distinguish between orthologs (e.g. X1 and Y1) and inter-species
inparalogs (e.g. X1 and Y2).</p>
    </sec>
    <sec id="sec-5">
      <title>Integrated Data Exploitation</title>
      <p>
        The mapping de ned between OrthoXML and the orthology ontology can be
applied over speci c datasets. In this work, we have transformed data from two
orthology resources, Inparanoid 8 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and OMA Sept 2014 hierarchical
orthologous groups [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. These two resources provide data in OrthoXML format, but they
structure orthology data in a di erent manner, so they constitute an interesting
exploratory use case. For example, OMA uses clusters of orthologs and clusters
of paralogs in a hierarchical manner but Inparanoid only considers clusters of
orthologs, which may contain paralogs generated by duplications after the
speciation of the two target species (termed in-paralogs [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). This means that all
paralogy relations have to be inferred from Inparanoid datasets. Besides, OMA
stores the taxonomic range associates with a certain cluster, but Inparanoid does
not. Provided that both resources use OrthoXML we are able to reuse the same
mapping le for transforming the data. SWIT is able to execute the rules that
can be applied to each dataset, which permits to skip in the case of Inparanoid,
for instance, the rule for transforming the taxonomic range. The transformed
RDF contains 8798758 genes from OMA, 1713180 genes for Homo sapiens
orthologs and 1367940 genes for Mus musculus orthologs from Inparanoid. Table 4
shows the SPARQL query for getting the orthologs of the human gene OR4D2 in
Mus Musculus, which would include the results from both Inparanoid and OMA.
The results of this query are shown in Table 5. In this use case, we have used
Virtuoso 7 as the triple store. Some sample queries exploiting this integrated
dataset are available at https://github.com/qfo/OrthologyOntology.
In this paper we have described how we can approach the standardization of
orthology content using semantic web technologies. We have been able to build
a draft of the orthology ontology by reusing existing ones. We have been able
gene species database
"Olfr463" "Mus musculus" "InParanoid"
"Olfr462" "Mus musculus" "InParanoid"
"MOUSE03761" "Mus musculus" "OMA"
"MOUSE03760" "Mus musculus" "OMA"
"MOUSE03762" "Mus musculus" "OMA"
to use state-of-the-art tools for all the processes involved: construction of the
ontology, de nition of mappings, transformation and exploitation of the data.
Our e ort would permit the systematic application of the process to any
OrthoXML database to generate an integrated knowledge base and to carry out
evaluation and testing processes in both usefulness and performance. There are
also remaining tasks and challenges. First, the draft of the ontology needs to be
improved in terms of classes, properties and documentation of what has already
been produced. We have started by de ning queries of pairwise paralogs and
orthologs, but other relations and concepts like in-paralogs, out-paralogs, etc. could
be de ned and encoded as queries. So far, there is no friendly way to query the
transformed content, only through SPARQL endpoints, which are for machines
rather than for humans. E orts in that sense should be made, taking into account
that there will be distributed SPARQL endpoints of orthology-related content.
The future work also includes the development of a Linked Data API over the
triple store that has method calls enabling the exploitation of the datasets.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This work was supported by the Ministerio de Econom a y Competitividad and
the FEDER programme through grant TIN2014-53749-C2-2-R2, the Fundacion
Seneca through grant 15295/PI/10, and the National Bioscience Database
Center, Japan Science Technology Agency. We would like to thank the organizers of
the BioHackathon meeting (http://www.biohackathon.org) for providing us an
invaluable opportunity to accomplish this work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mackey</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoeckert</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          :
          <article-title>Orthomcl-db: querying a comprehensive multi-species collection of ortholog groups</article-title>
          .
          <source>Nucleic acids research 34(suppl 1)</source>
          ,
          <source>D363{D368</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chiba</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nishide</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uchiyama</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Construction of an ortholog database using the semantic web technology for integrative analysis of genomic data</article-title>
          .
          <source>PloS one 10(4)</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dessimoz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cannarozzi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Margadant</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonnet</surname>
            ,
            <given-names>G.H.</given-names>
          </string-name>
          :
          <article-title>Oma, a comprehensive, automated project for the identi cation of orthologs from complete genome data: introduction and rst achievements</article-title>
          .
          <source>In: Comparative Genomics</source>
          , pp.
          <volume>61</volume>
          {
          <fpage>72</fpage>
          . Springer (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baran</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callahan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chepelev</surname>
            ,
            <given-names>L.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz-Toledo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicholas</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rio</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duck</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furlong</surname>
            ,
            <given-names>L.I.</given-names>
          </string-name>
          , et al.:
          <article-title>The semanticscience integrated ontology (sio) for biomedical research and knowledge discovery</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>5</volume>
          ,
          <issue>14</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Koonin</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          :
          <article-title>Orthologs, paralogs, and evolutionary genomics 1</article-title>
          .
          <source>Annu. Rev. Genet</source>
          .
          <volume>39</volume>
          ,
          <issue>309</issue>
          {
          <fpage>338</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Towards a semantic web application: Ontology-driven ortholog clustering analysis</article-title>
          .
          <source>In: ICBO</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>McKusick</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          :
          <article-title>Mendelian inheritance in man: a catalog of human genes and genetic disorders</article-title>
          ,
          <source>vol. 1</source>
          . JHU Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Min</surname>
          </string-name>
          <article-title>~arro-</article-title>
          <string-name>
            <surname>Gimenez</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          , Madrid,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Fernandez-Breis</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.T.</surname>
          </string-name>
          :
          <article-title>Ogo: an ontological approach for integrating knowledge about orthology</article-title>
          .
          <source>BMC bioinformatics 10(Suppl</source>
          <volume>10</volume>
          ),
          <source>S13</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Min</surname>
          </string-name>
          <article-title>~arro-</article-title>
          <string-name>
            <surname>Gimenez</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <source>Egan~a Aranguren</source>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Villazon-Terrazas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Fernandez</surname>
          </string-name>
          <string-name>
            <surname>Breis</surname>
          </string-name>
          , J.T.:
          <article-title>Translational research combining orthologous genes and human diseases with the ogolod dataset</article-title>
          .
          <source>Semantic Web</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <volume>145</volume>
          {
          <fpage>149</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>O</given-names>
            <surname>'Brien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.P.</given-names>
            ,
            <surname>Remm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sonnhammer</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.L.</surname>
          </string-name>
          :
          <article-title>Inparanoid: a comprehensive database of eukaryotic orthologs</article-title>
          .
          <source>Nucleic acids research 33(suppl 1)</source>
          ,
          <source>D476{D480</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pellegrini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcotte</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeates</surname>
            ,
            <given-names>T.O.</given-names>
          </string-name>
          :
          <article-title>Assigning protein functions by comparative genome analysis: protein phylogenetic pro les</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>96</volume>
          (
          <issue>8</issue>
          ),
          <volume>4285</volume>
          {
          <fpage>4288</fpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Prosdocimi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chisham</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pontelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoltzfus</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Initial implementation of a comparative data analysis ontology</article-title>
          .
          <source>Evolutionary bioinformatics online 5</source>
          ,
          <issue>47</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Remm</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Storm</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sonnhammer</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Automatic clustering of orthologs and in-paralogs from pairwise species comparisons</article-title>
          .
          <source>Journal of molecular biology 314(5)</source>
          ,
          <volume>1041</volume>
          {
          <fpage>1052</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Roux</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson-Rechavi</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An ontology to clarify homology-related concepts</article-title>
          .
          <source>Trends in Genetics</source>
          <volume>26</volume>
          (
          <issue>3</issue>
          ),
          <volume>99</volume>
          {
          <fpage>102</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Schmitt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Messina</surname>
            ,
            <given-names>D.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreiber</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sonnhammer</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Seqxml and orthoxml: standards for sequence and orthology information</article-title>
          . Brie ngs in bioinformatics p.
          <source>bbr025</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Shearer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
          </string-name>
          , I.:
          <article-title>HermiT: A highly-e cient OWL reasoner</article-title>
          .
          <source>In: Proceedings of the 5th International Workshop on OWL: Experiences and Directions (OWLED</source>
          <year>2008</year>
          ). pp.
          <volume>26</volume>
          {
          <issue>27</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sioutos</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>de Coronado</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haber</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hartel</surname>
            ,
            <given-names>F.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaiu</surname>
            ,
            <given-names>W.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>L.W.</given-names>
          </string-name>
          :
          <article-title>Nci thesaurus: a semantic model integrating cancer-related clinical and molecular information</article-title>
          .
          <source>Journal of biomedical informatics 40(1)</source>
          ,
          <volume>30</volume>
          {
          <fpage>43</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceusters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klagges</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Kohler, J.,
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lomax</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neuhaus</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rector</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosse</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Relations in biomedical ontologies</article-title>
          .
          <source>Genome biology</source>
          <volume>6</volume>
          (
          <issue>5</issue>
          ),
          <source>R46</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Sonnhammer</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabaldon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>da Silva</surname>
            ,
            <given-names>A.W.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson-Rechavi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boeckmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dessimoz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , et al.:
          <article-title>Big data and other challenges in the quest for orthologs</article-title>
          . Bioinformatics p.
          <year>btu492</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Uchiyama</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mihara</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nishide</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiba</surname>
          </string-name>
          , H.:
          <article-title>Mbgd update 2015: microbial genome database for exible ortholog analysis utilizing a diverse set of genomic data</article-title>
          .
          <source>Nucleic acids research</source>
          <volume>43</volume>
          (
          <issue>D1</issue>
          ),
          <source>D270{D276</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>