<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Similarity between semantic description sets: addressing needs beyond data integration</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Todd Vision</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Judith Blake</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hilmar Lapp</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paula Mabee</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monte Wester eld</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jackson Laboratory</institution>
          ,
          <addr-line>Bar Harbor, ME</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Evolutionary Synthesis Center</institution>
          ,
          <addr-line>Durham, NC</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of North Carolina at Chapel Hill</institution>
          ,
          <addr-line>Chapel Hill, NC</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Oregon</institution>
          ,
          <addr-line>Eugene, OR</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of South Dakota</institution>
          ,
          <addr-line>Vermillion, SD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Descriptive information is easy to understand and communicate in natural language. Examples in the biological realm include the cellular functions of proteins and the phenotypes exhibited by organisms. Large latent stores of such descriptive data are stored in databases that can be mined, but even more still reside only in the scienti c literature. Although such information has traditionally been opaque to computers, in recent years signi cant e orts have gone into exposing descriptive information to computation through the development of ontologies and associated tools. A host of software applications now employ simple reasoning over Gene Ontology annotated data to help interpret experimental ndings in genomics in terms of protein function. In the domain of biological phenotypes, the combination of entity terms from taxonspeci c anatomy ontologies with quality terms from generic ontologies such as PATO have been used to construct semantically precise and contextualized descriptions. It is natural for multiple semantic descriptions to pertain to single instances in the real world, as in the case of both protein functions and organismal phenotypes. However, applications for ontology-based annotations that go beyond simple knowledge organization, and that exploit sets of semantic descriptions, are puzzlingly rare. In particular, we argue that there is wide applicability, and a sore need, for tools that can satisfy the simple, common use case of identifying statistically improbable similarity between sets of semantic descriptions. Several metrics have been proposed for this task in the literature, but not yet fully evaluated, explored, and adopted. The requirements for semantic similarity tools tailored to sets of semantic descriptions would include speed, scalability to large numbers of sets, demonstrated statistical and biological validity, and ease of use.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Ontologies are a foundational technology for a semantic web of linked data.
As a key element for data discovery, reuse, and integration, they allow the
standardization and relation of concepts across documents, databases, and
communities of practice. Ontologies also allow the semantics of concepts to be exposed
to formal machine reasoning, and thus provide the opportunity for the linked
data web to be used for more sophisticated logical operations than are currently
possible.</p>
      <p>
        Here, we de ne descriptive data broadly as information about the qualitative
features of objects in the world,e.g.\albatrosses have long wing". A large and
diverse universe of descriptive data is known to science, but such statements
are typically expressed in natural language, as above. These are then typically
transformed to quantitative form (e.g. word frequency) for the purposes of
computation, and loss of meaning accompanies this transformation. What if, instead,
we had tools to exploit the semantic content of diverse collections of descriptive
data directly [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]?
      </p>
      <p>
        Biological discourse is particularly rich in descriptive data, though it is often
expressed within text and not managed within data collections. The Gene
Ontology (GO) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] rst introduced the biological community to the use of ontologies
for standardization of descriptive data, in this case the function and location of
gene products, across a broad community of practice. The years since have
witnessed steady growth in the diversity of specialized knowledge domains within
biology, particularly biomedicine, for which ontologies have been developed, as
illustrated by the current breadth of the NCBO BioPortal [11] and the Open
Biological and Biomedical Ontologies [15]. The popularity of ontologies among
biologists is in large part due to their suitability as controlled vocabularies that
aid in the harmonization and integration of terminology among di erent
communities of practice [13]. Secondarily, they are increasingly used as a classi cation
aid that enables richer navigation of information resources [
        <xref ref-type="bibr" rid="ref7">14, 7</xref>
        ], and to identify
those concepts that are more frequently associated with documents in a corpus,
either directly or by inference, than would be expected by chance alone [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Although tremendously useful, these relatively straightforward knowledge
organization tasks fail to exploit the potential of ontologies as tools for scienti c
knowledge discovery. The e ort required to produce an ontology for a knowledge
domain can be signi cant, and engaging the community of domain experts to
ensure its tness for wide adoption is challenging. Knowledge discovery
applications o er an additional route by which ontologies can deliver powerful scienti c
returns to bench biologists, and in so doing incentivize them to contribute to
the building of more comprehensive and useful community ontologies.</p>
      <p>
        One such application that has recently been shown to harbor great
potential, particularly in biology [16] and drug discovery [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ], is semantic similarity
search. Brie y, a semantic similarity search application takes as input a set of
one or more ontological statements of descriptive data, compares the set against
a database of such sets, and returns those sets from the database that have
greater semantic similarity to the query set than would be expected by chance.
Each set corresponds to a number of descriptive data statements, in the form
of ontological classes, made about a common object. The semantic similarity
between two classes may re ect the ontology graph distances between concepts
and/or the information content of common subsuming ontology concepts (see
[12] for an overview of metrics used in bioinformatics applications). A common
application of semantic similarity search is to identify objects in a database that
are semantically similar to a query object. For example, Washington et al. [16]
used classes expressing the semantics of heritable human disease phenotypes
to query a database of mutant phenotypes from genetically well-characterized
model organisms. The query returned model organism genes that have mutant
phenotypes with semantics similar to the phenotypes of human heritable
diseases. The genes obtained in this way suggested testable hypotheses for the
genetic causes of those diseases, which were previously unknown.
      </p>
      <p>
        The semantic similarity metrics employed so far are relatively
straightforward to calculate if they are applied to ontology classes that, aside from being
placed in a subsumption hierarchy, are not axiomatically de ned, for example
when assessing the semantic similarity between genes based on the GO terms
associated with each. However, the inferences possible from such simple
hierarchies are limited. Conversely, axiomatically de ning the classes by combining
several orthogonal, modular domain ontologies as class expressions, such as using
intersections of property restrictions in OWL, greatly increases expressivity of
the ontological expressions, as well as the inferences a reasoner can make from
them [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This in turn can increase sensitivity as well as speci city of nding
semantically similar matches. However, as a consequence of the increased
expressivity, enumerating all subsumers of a complex, possibly nested, class expression,
which the currently best-performing metrics require, can quickly become time
and memory consuming with large ontologies and large databases of class
expressions. As an example, calculating the similarity statistics for the Washington
et al. study took several days on a relatively small database with only several
thousand sets of class expressions.
      </p>
      <p>
        The use of ontologies to reason about qualitative phenotypes is now being
explored by a number of di erent groups. Among these e orts is Phenoscape, a
project with which we are all involved and which aims to enable computation
across phenotypic information from di erent biological disciplines (e.g. genetics
and biodiversity) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our application of semantic similarity is over sets of
phenotypic data, where a phenotype is de ned as the set of observable traits present
in an individual organism as a result of the interaction of heredity,
environmental in uences, and the developmental process, e.g. the elongated wings of
albatrosses. The qualitative phenotype descriptions are central to and investigated
in meticulous detail in many di erent areas of biology, and such phenotypes
have traditionally been reported and communicated in expressive, but highly
discipline-speci c natural language.
      </p>
      <p>
        To express qualitative phenotypes with computable semantics, we use the
emerging standard of an Entity-Quality (EQ) formalism [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which decomposes
phenotypes into three main components: a quality (e.g., \elongate" shape), the
entity bearing that quality (e.g., a \wing"); and the class of organism
expressing the phenotype (e.g., a taxon or a genotype). Each of the components of
an EQ expression, and the relations between them, are expressed using terms
and properties from appropriate ontologies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. When represented in OWL, EQ
expressions are conjunctive class expressions, and thus axiomatically de ned
classes. Phenotypes may be somewhat complex, such as qualities borne only by
particular regions or parts of an anatomical entity, or described as spatial
relationships between anatomical structures, and therefore such class expressions
can consist of multiple and recursively nested property restrictions.
      </p>
      <p>The Phenoscape Knowledgebase (http://kb.phenoscape.org) currently
contains over half a million semantic phenotype descriptions in the form of EQ class
expressions for more than 5,000 di erent biological taxa that are linked to more
than 4,000 candidate genes through EQ phenotype descriptions from mutants of
a single model organism. Thus, there are many sets of descriptions with a
cardinality on the order of 100 descriptions. Although this is a large data store, it has
been compiled from only a restricted branch of the tree of life (ostariophysan
shes), and thus represents only a small fraction of the amount of latent
phenotype information for all organisms. Although restricted in taxonomic scope, the
Phenoscape Knowledgebase already contains information that would be su
cient to generate thousands of hypotheses about the genetic basis of phenotypic
transitions in evolution, were we in a position to compute the semantic similarity
between the sets of semantic phenotype expressions associated with each of the
thousands of taxa and genotypes!</p>
      <p>Calculating the best performing of the currently available metrics for
semantic similarity [12] requires the enumeration of all subsumers, or identifying
the Least Common Subsumer, of the classes for which the similarity is being
evaluated. Available DL reasoners can return only subsuming classes that are
actually present in a knowledgebase. Although this is su cient if the classes
being evaluated are terms drawn from a subsumption hierarchy, in the case of class
expressions there is typically a large number of possible combinations of
property restrictions that subsume a given class expression, most of which will not
be present in the knowledgebase. A custom-written algorithm could be used to
enumerate all possible subsumers and add them to the knowledgebase so that a
DL reasoner can subsequently return them, but the number of subsumers grows
combinatorially with the number of property restrictions, and thus quickly
becomes too large and too time-consuming to compute at query time. For example,
for a conjunctive class expression with n property restrictions, where object ci
in restriction i has Ni asserted or inferred superclasses (in the knowledgebase),
there are (Q Ni)=n! possible subsuming class expressions. Based on our results
so far, semantic similarity search engines will require a speedup of two to three
orders of magnitude to enable a user to launch multiple searches over even a few
thousand sets of semantic descriptions (the current scale of Phenoscape), if the
results are to be returned within a single user session. To be t for large-scale
reasoning over a web of linked data, algorithms will need to scale to substantially
larger problem instances.</p>
      <p>Additionally, although there have been some performance comparisons among
metrics, and important biological demonstrations of the validity of some of the
patterns found [12, 16], the available measures have not been statistically
justied in terms of consistency and bias, nor benchmarked against error. It is likely
that there is considerable room for improvement in the metrics that are available.</p>
      <p>Finally, to enable the broad use of semantic similarity search technology
in bioinformatics and beyond, applications will need to be embedded within
web applications that are accessible to the average information-seeking scientist.
Complex dependencies on external software such as a database or reasoner, and
the need to convert the format of source ontologies or data, will consign use of
such tools to the realm of the specialist.</p>
      <p>The paucity of tools available for computation over qualitative data in
bioinformatics is a striking contrast to the vast array of tools that operate on other
forms of non-numeric data, such as the character strings found in nucleotide and
protein sequences. This is truly unfortunate given the importance and volume
of descriptive data within biological discourse, and given the many applications
that await to be developed for nding similar objects on the web of linked data.
11. Noy, N.F., Shah, N.H., Whetzel, P.L., Dai, B., Dorf, M., Gri th, N., Jonquet, C.,
Rubin, D.L., Storey, M.A., Chute, C.G., Musen, M.A.: BioPortal: ontologies and
integrated data resources at the click of a mouse. Nucleic Acids Research 37(Web
Server issue), W170{3 (Jul 2009)
12. Pesquita, C., Faria, D., Falca~o, A.O., Lord, P., Couto, F.M.: Semantic similarity
in biomedical ontologies. PLoS Computational Biology 5(7), e1000443 (Jul 2009)
13. Schuurman, N., Leszczynski, A.: Ontologies for bioinformatics. Bioinformatics and</p>
      <p>Biology Insights 2, 187{200 (Jan 2008)
14. Shah, N.H., Jonquet, C., Chiang, A.P., Butte, A.J., Chen, R., Musen, M.a.:
Ontology-driven indexing of public datasets for translational bioinformatics. BMC
Bioinformatics 10 Suppl 2, S1 (Jan 2009)
15. Smith, B., Ashburner, M., Rosse, C., Bard, J., Bug, W., Ceusters, W., Goldberg,
L., Eilbeck, K., Ireland, A., Mungall, C., Leontis, N., Rocca-Serra, P., Ruttenberg,
A., Sansone, S.A., Scheuermann, R., Shah, N., Whetzel, P., Lewis, S.: The OBO
Foundry: coordinated evolution of ontologies to support biomedical data
integration. Nature Biotechnology 25(11), 1251{1255 (Nov 2007)
16. Washington, N.L., Haendel, M.A., Mungall, C.J., Ashburner, M., Wester eld, M.,
Lewis, S.E.: Linking human diseases to animal models using ontology-based
phenotype annotation. PLoS Biology 7(11), e1000247 (Nov 2009)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ball</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Botstein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Butler</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolinski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dwight</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eppig</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>IsselTarver</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasarskis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matese</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richardson</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ringwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherlock</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Gene ontology: tool for the uni cation of biology. The Gene Ontology Consortium</article-title>
          .
          <source>Nature Genetics</source>
          <volume>25</volume>
          (
          <issue>1</issue>
          ),
          <volume>25</volume>
          {9 (May
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Campillos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gavin</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bork</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Drug target identication using side-e ect similarity</article-title>
          .
          <source>Science</source>
          (New York, N.Y.)
          <volume>321</volume>
          (
          <issue>5886</issue>
          ),
          <volume>263</volume>
          {266 (Jul
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dahdul</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balho</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engeman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grande</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hilton</surname>
            ,
            <given-names>E.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kothari</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapp</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lundberg</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Midford</surname>
            ,
            <given-names>P.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vision</surname>
          </string-name>
          , T.J., Wester eld, M.,
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          :
          <article-title>Evolutionary Characters, Phenotypes and Ontologies: Curating Data from the Systematic Biology Literature</article-title>
          .
          <source>PLoS ONE</source>
          <volume>5</volume>
          (
          <issue>5</issue>
          ), e10708 (May
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>J.a.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Couto</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          :
          <article-title>Semantic Similarity for Automatic Classi cation of Chemical Compounds</article-title>
          .
          <source>PLoS Computational Biology</source>
          <volume>6</volume>
          (
          <issue>9</issue>
          ),
          <source>e1000937 (Sep</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherman</surname>
            ,
            <given-names>B.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lempicki</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          :
          <article-title>Bioinformatics enrichment tools: paths toward the comprehensive functional analysis of large gene lists</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>13</fpage>
          (Jan
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bork</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Ontologies in quantitative biology: a basis for comparison, integration, and discovery</article-title>
          .
          <source>PLoS Biology</source>
          <volume>8</volume>
          (
          <issue>5</issue>
          ), e1000374 (May
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kapushesky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Emam</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holloway</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurnosov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zorin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustici</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parkinson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brazma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gene expression atlas at the European bioinformatics institute</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>38</volume>
          (
          <issue>Database issue</issue>
          ),
          <source>D690{8 (Jan</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cronk</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkoutos</surname>
            ,
            <given-names>G.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haendel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Segerdell</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wester</surname>
            <given-names>eld</given-names>
          </string-name>
          , M.:
          <article-title>Phenotype ontologies: the bridge between genomics and evolution</article-title>
          .
          <source>Trends in Ecology &amp; Evolution</source>
          <volume>22</volume>
          (
          <issue>7</issue>
          ),
          <volume>345</volume>
          {50 (Jul
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berardini</surname>
            ,
            <given-names>T.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deegan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ireland</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lomax</surname>
          </string-name>
          , J.:
          <article-title>Cross-product extensions of the Gene Ontology</article-title>
          .
          <source>Journal of Biomedical Informatics (Feb</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkoutos</surname>
            ,
            <given-names>G.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haendel</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Integrating phenotype ontologies across multiple species</article-title>
          .
          <source>Genome Biology</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <source>R2 (Jan</source>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>