<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unifying Phenotypes to Support Semantic Descriptions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eduardo Miranda</string-name>
          <email>eduardo.miranda@students.ic.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andre´ Santanche`</string-name>
          <email>santanche@ic.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computing - State University of Campinas Av. Albert Einstein</institution>
          ,
          <addr-line>1251 - Cidade Universita ́ria, Campinas</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>154</fpage>
      <lpage>165</lpage>
      <abstract>
        <p>In life sciences, there are several biological datasets shared through the web. All this abundance of data carries a great opportunity to explore complex relationships among the diversity of species. However, their physical format varies from independent data files to databases, which are heterogeneous in model and representation, hampering their integration. Ontologies are one of the promising choices to address this challenge. However, the existing digital phenotypic descriptions are stored in semi-structured formats, making extensive use of natural language. If on one hand, this patrimony is highly relevant, on the other hand, converting it in ontologies is not a straightforward task. The present article addresses this problem adding an intermediate step between semi-structured phenotypic descriptions and ontologies. It remodels semi-structured descriptions to a graph abstraction in which the data are linked. Graph transformations subsidize the transition from semi-structured data representation to a more formalized representation through ontologies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Bioinformatics is the science of integrating, managing, mining and interpreting
information from biological data [
        <xref ref-type="bibr" rid="ref7">Gibas and Jambeck 2001</xref>
        ]. In the life science field, there are a
large number of distributed biological datasets freely available and ready to use. However,
this wealth of information has hardly been tapped even today due its distributed nature,
heterogeneity and complex data types and representation [
        <xref ref-type="bibr" rid="ref16">Parr et al. 2012</xref>
        ]. In this
scenario, their combination and interconnection are barely feasible [
        <xref ref-type="bibr" rid="ref18">Quan 2007</xref>
        ]. A massive
amount of relevant information is hidden in the potential connection of unrelated files.
      </p>
      <p>
        In this work we are interested in a specific biology context, in which biologists
apply computational tools to build and share digital descriptions of living beings as
phenotypes. These descriptions are a fundamental starting point for several biology tasks,
like living beings identification and tools for phylogenetic tree analysis. Even though the
last generation of these tools is based on open standards (e.g., XML), the descriptions are
still based on textual sentences in natural language [
        <xref ref-type="bibr" rid="ref2">Balhoff et al. 2010</xref>
        ].
      </p>
      <p>
        Semantic integration in this context is one of the main challenges. Besides
ontologies to support phenotype description, there are tools to annotate descriptions by
associating ontology concepts to textual descriptions [
        <xref ref-type="bibr" rid="ref2">Balhoff et al. 2010</xref>
        ]. This distinction
between description and their annotations based on ontologies does not consider that
descriptions can conversely contribute to ontology expansion and revision. The challenge in
this work is to establish a model to represent a common denominator among phenotipical
description standards, which will support findings in the latent semantics implicit in
relations in a strategy inspired by folksonomies. These semantics can guide the interaction
between textual descriptions and ontologies.
      </p>
      <p>
        In a previous work [
        <xref ref-type="bibr" rid="ref1">Alves and Santanche` 2013</xref>
        ], we showed that the latent
semantics presented in tags and their correlations, as a product of an organic work collectively
produced by a community on the web (the folksonomies), can be exploited to expand and
review ontologies. While the model behind folksonomies is based on the correlation of
three elements – tags, resources and users – descriptions in the biological context present
a more complex and specialized structures. Co-occurrence is a strong principle we
considered to extract latent semantics. The main idea is that the set of tags put together in a
given resource can provide a “context” to interpret each tag. Consider a tag cell, which
can have a distinct interpretation according to the context. The co-occurrence with the
tags cytoplasm or organelle will put it in the biology context. Moreover, the compilation
of data concerning the occurrence and co-occurrence of millions of tags can support the
analysis of similarity among terms – see more details in [
        <xref ref-type="bibr" rid="ref1">Alves and Santanche` 2013</xref>
        ]. We
consider that we can apply an equivalent technique to put terms of phenotype descriptions
in a context, to improve their interpretation and correlation.
      </p>
      <p>The present paper addresses this problem in exploiting existing biology assets
related to phenotypic descriptions, and the latent semantics resulting from their
interconnection, to support their development towards a richer semantical representation, as part
of ontologies. It implies promoting relations among concepts to first class citizens.
Accordingly, we designed a three layered method illustrated in Figure 1, in which graph
databases intermediate this evolvement process from fragmentary data sources to
accomplish full integration descriptions as ontologies.</p>
      <p>Our approach remodels semi-structured descriptions to a graph abstraction, in
which the data can be integrated more easily. Graph transformations are applied for the
transition from a semi-structured data representation to a more formalized
representation through ontologies. As we will further explain, this graph representation will also
support an analytical tool to compare data across studies, wherein it will help
evolutionary biologists to answer evolutionary questions. This paper presents a work in progress
concerning the first step of this method, focusing in the integration of data from the
semistructured data layer and their transition to the graph data abstraction layer. Our proposed
graph-based model is derived from a comparative analysis among four standards related
to phenotype description, plus a practical experiment.</p>
      <sec id="sec-1-1">
        <title>Ontology concepts</title>
      </sec>
      <sec id="sec-1-2">
        <title>Graph data abstraction</title>
      </sec>
      <sec id="sec-1-3">
        <title>Semi-structured data</title>
        <p>This paper is organized as follows: Section 2 summarizes the related work;
Section 3 presents the comparative analysis which subsidizes our minimal common
denominator model; Section 4 presents out graph-based model; Section 5 shows a practical
experiment of unifying phenotypes; Section 6 presents concluding remarks.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Integration is a key point as humans are progressively unable of handling the sheer volume
of data presented [
        <xref ref-type="bibr" rid="ref3">Bell et al. 2009</xref>
        ]. It is an important step towards knowledge discovery
[
        <xref ref-type="bibr" rid="ref10">Lenzerini 2002</xref>
        ]. The integration of digital phenotype descriptions is a relevant challenge
in this context since they support fundamental biology tasks as the building of
identification keys for living beings and can support the creation of a complete evolutionary Tree of
Life [
        <xref ref-type="bibr" rid="ref16">Parr et al. 2012</xref>
        ] assembling genomic and morphological data so as to congregate the
phylogenetic relationships among all living or extinct organisms [
        <xref ref-type="bibr" rid="ref4">Ciccarelli et al. 2006</xref>
        ].
Likewise, integrating these data may contribute to better understanding of how a
morphological trait became organized and evolved over time [
        <xref ref-type="bibr" rid="ref11">Mabee 2006</xref>
        ].
      </p>
      <p>
        Recent approaches enrich descriptions via ontology annotations, using the
Entity-Quality (EQ) formalism for phenotype modeling. EQ is a representation
[
        <xref ref-type="bibr" rid="ref2">Balhoff et al. 2010</xref>
        ] which associates ontology entity terms (E) – e.g., bone or vertebra
from Teleost Anatomy Ontology (TAO) – with quality terms (Q) – e.g., triangular,
horizontal, smooth from the Phenotype and Trait Ontology (PATO) [
        <xref ref-type="bibr" rid="ref6">Dahdul et al. 2010</xref>
        ].
Ontologies have gained wide acceptance in biology due to their ability of representing
knowledge and also the advantage of querying and reasoning information [
        <xref ref-type="bibr" rid="ref8">Gkoutos et al. 2004</xref>
        ].
Furthermore, semantic web standards to represent ontology concepts with unique
identifiers facilitates interoperability across databases [
        <xref ref-type="bibr" rid="ref12">Mabee et al. 2007</xref>
        ]. Recently, several
tools have emerged to support annotation of biological phenotypes using ontologies,
e.g., Phenex (http://phenoscape.org/wiki/Phenex) and Phenote (http://www.phenote.org/ ),
both curation tools designed for annotation of phenotypic characters with ontology
concepts using EQ formalism [
        <xref ref-type="bibr" rid="ref2">Balhoff et al. 2010</xref>
        ].
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref6">Dahdul et al. 2010</xref>
        ] developed a workflow for curation of phenotypic characters
extracted from scientific publications. It is important to note the limitations of this
curation process, considering that it is very time-consuming since it is manually carried out
by domain experts.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Common Denominator</title>
      <p>There is a wide variety of representation formats for phenotype description, adopted by
information systems and open standards, which represent differently the same
information. In this section, we analyze four of them – Xper2, SDD, Nexus and NeXML – looking
for a minimal common denominator, which is the foundation for our graph-based model,
to be used to link related information.</p>
      <p>
        SDD, Nexus and NeXML are widely adopted open standards further detailed.
Xper2 (http://lis-upmc.snv.jussieu.fr/lis/ ) is a management system adopted by the
systematist community, for the storing, editing and analyzing of phenotype descriptive data. It
focuses mainly on taxonomic descriptions, allowing creation, sharing and comparison
of identification keys [
        <xref ref-type="bibr" rid="ref21 ref22">Ung et al. 2010</xref>
        a,
        <xref ref-type="bibr" rid="ref21 ref22">Ung et al. 2010</xref>
        b]. Xper2 was developed in the
Laboratoire Informatique &amp; Syste´matique of the University Pierre et Marie Curie and
this work is part of a bigger project in collaboration with this lab. Therefore, Xper2 was
adopted for our practical experiments.
      </p>
      <p>
        In order to illustrate our analysis, let us consider a practical case, in which a
biologist is building a phenotype description of monitor lizards (genus Varanus). The process
starts with the biologist collecting observations of lizards, organized as characters and
character states (C, CS). [
        <xref ref-type="bibr" rid="ref17">Pimentcl and Riggins 1987</xref>
        ] defined character as “a feature of
organisms that can be evaluated as a variable with two or more mutually exclusive and
ordered states”. The observations involved the species Varanus albiguralis and Varanus
brevicauda. The final result is the character-by-taxon matrix illustrated in Figure 2.
      </p>
      <p>Varanus albiguralis
Varanus brevicauda
2
1
m
fr
o
'
s
l
it
r
s
o
n
l
it
a
e
h
lsa fto
r
e n
svn ito
tra sce
1
2
2
1
s
e
l
a
c
s
l
a
h
c
u
n</p>
      <sec id="sec-3-1">
        <title>Nostrils' form</title>
        <p>1 – well round
2 – oval or split-like</p>
      </sec>
      <sec id="sec-3-2">
        <title>Transversal section of the tail</title>
        <p>1 – laterally compressed
2 – roundish</p>
      </sec>
      <sec id="sec-3-3">
        <title>Nuchal scales</title>
        <p>1 – same size than head scales
2 – bigger than head scales</p>
        <p>
          In order to transform these observations to digital records and generalize them –
e.g., devising general characters and states observed in a genre of monitor lizards – the
biologist will use a tool like Xper2. Phenotypes descriptions can be stored in the Xper2
native format or can be exported to the SDD open format. The Structure Descriptive Data
(SDD) (http://wiki.tdwg.org/SDD) is a platform and application-independent XML-based
standard developed by the Biodiversity Information Standards (historic acronym: TDWG)
for recording and exchanging descriptions of biological and biodiversity data of any type
[
          <xref ref-type="bibr" rid="ref9">Hagedorn 2007</xref>
          ]. SDD is adopted by several other phenotype description tools – e.g.,
Lucid Central (http://www.lucidcentral.org) and Linnaeus II (http://www.eti.uva.nl/ ).
        </p>
        <p>We further introduce some key elements of the SDD format, which are recurrent
in the formats confronted in this section. A SDD description comprises, in a single file,
a domain schema and its instances. Figure 3 shows a diagram with a fragment of a SDD
file containing the description of a varanus lizard. A (C,CS) description in SDD has two
main blocks: (i) defines the characters involved and their possible states – Figure 3 top;
(ii) describes an Operational Taxonomic Unit (OTU) using the characters defined in (i) –
Figure 3 bottom. OTU is a biology term which refers to a given entity in sampling level
adopted to the study – e.g., a specimen, a gender etc.</p>
        <p>
          &lt;CategoricalCharacter&gt;s and their &lt;States&gt; (shown in Figure 3 top) are
primitives to describe an OTU [
          <xref ref-type="bibr" rid="ref9">Hagedorn 2007</xref>
          ]. Each &lt;CategoricalCharacter&gt; has its
&lt;Representation&gt; – comprising a label and a description as plain texts – and a set of
&lt;StateDefinition&gt; elements with their possible states. &lt;CategoricalCharacter&gt; and
&lt;StateDefinition&gt; elements defined here will be referred throughout the XML document
by their ids.
        </p>
        <p>The &lt;CodedDescription&gt; (Figure 3 bottom) links the OTU being described
to States of each &lt;CategoricalCharacter&gt;. It has two essential items: (i) the OTU
Datasets
Dataset</p>
        <p>id=“c6”
CategoricalCharacter</p>
        <p>id=“D1”
CodedDescription</p>
        <p>States
SummaryData
Representation</p>
        <p>ref=“c6”
Categorical</p>
        <p>Label</p>
        <p>“V. albiguralis”
Detail
“Monitors' nostrils may have different forms...”</p>
        <p>id=“s12”
StateDefinition</p>
        <p>id=“s13”
StateDefinition</p>
        <p>Label</p>
        <p>“well round”
Detail</p>
        <p>“Nostrils look like a quite per...”
Label</p>
        <p>“oval or split-like”
Detail</p>
        <p>“Nostrils are not perfectly rou...”
State</p>
        <p>ref=“s13”
Detail</p>
        <p>“White-throated monitor. Distribution: Africa (West...”
being described, where its name and description are listed in natural language under
&lt;Representation&gt;; (ii) a set of character and values (&lt;Categorical&gt; and &lt;State&gt;),
which address the characters defined in the previous section through the ref attribute.
It is possible and usual to define multiple states for a character of a given OTU. A first
integration, problem observed here is that each character or OTU described does not have
a global unique identification among documents. Therefore, the description can only be
used by the document where it was declared and it is not possible to guarantee the
equivalence of two or more &lt;CategoricalCharacters&gt;.</p>
        <p>In Figure 5 we expand our analysis to the Xper2 native format, Nexus and NeXML.
Our study addresses mainly morphological character descriptions. Figure 5 provides
simplified diagrams focusing on the elements to record descriptions, which will be confronted
here. Figure 4 presents the symbols adopted in the diagram. All the formats adopt XML
and the symbols represent the relations among elements and their respective cardinality.
Five types of elements, which are focus of our analysis, receive special symbols: the
Entity being described, which can be a taxon or a specimen; the Character
definition and its respective association with entities (Character instance); the State
definition and its respective association with entities (State instance).</p>
        <p>
          Nexus [
          <xref ref-type="bibr" rid="ref13">Maddison et al. 1997</xref>
          ] is an extensively used file format developed for
storage and exchange of phylogenetic data, including morphological and molecular
characters, taxa distances, genetic codes, phylogenetic trees etc. It was designed in 1987 and it is
still used by many popular software as Xper2 (http://lis-upmc.snv.jussieu.fr/lis/ ), Mesquite
(http://mesquiteproject.org/ ), MrBayes (http://mrbayes.sourceforge.net/ ) and data
repositories, like TreeBASE(http://treebase.org/ ) and Dryad (http://datadryad.org/ ). Nexus
gathers together (C,CS) based descriptions and related trees [
          <xref ref-type="bibr" rid="ref23">Vos et al. 2012</xref>
          ].
        </p>
        <p>1Knowledge base of the genus Varanus from
http://lis-upmc.snv.jussieu.fr/xper2/infosXper2Bases/listebases-recherche.php
Element types
structural
element
exclusive
option</p>
        <p>Relationship types
one to one
one to many
(zero or more)
Structural element specializations
one to one
(zero or one)
one to many
(one or more)
Entity</p>
        <p>
          NeXML (http://www.nexml.org) [
          <xref ref-type="bibr" rid="ref23">Vos et al. 2012</xref>
          ] is a standard inspired by the
Nexus. It supports and extends Nexus functionalities and addresses some Nexus
limitations – e.g., connects objects with ontology concepts, supports citations and annotations
[
          <xref ref-type="bibr" rid="ref23">Vos et al. 2012</xref>
          ]. In order to accomplish full compatibility and interoperability among
different environments, NeXML defines a formalized XSD grammar and enables
semantic annotations of any element in a NeXML document, which goes towards to a
“Minimum Information About a Phylogenetic Analysis” (MIAPA) standard.
        </p>
        <p>
          These comparative diagrams show that even if the structures are arranged
differently, they address the same key elements. All formats organize data in accordance
with the (C,CS) data model that, in practice, is an entity-attribute-value (EAV) model,
in which entities are OTUs, attributes are characters and values are character-states
[
          <xref ref-type="bibr" rid="ref23">Vos et al. 2012</xref>
          ]. Nexus and NeXML formats define a matrix, in which OTUs are listed in
rows, characters are columns and the cells contain a numeric code for a specific
characterstate (see Figure 2). Although Xper2 and SDD do not define a matrix, both formats have
a similar structure to describe OTUs with their (C, CS) records.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. From XML Structures to Graphs</title>
      <p>
        The next step in our Three Tier Method is designing a graph model. In a previous
work [
        <xref ref-type="bibr" rid="ref1">Alves and Santanche` 2013</xref>
        ], we have compared several approaches to capture
latent relations+semantics among tags produced collaboratively. Graph models to represent
and analyze data were a common denominator. The role of the graph is not to reflect all
details of the original model. The central challenge is how to abstract key elements, for
which we are looking for potential relations to be discovered. It is a movement from the
latent semantics to an explicit semantics expressed as links.
      </p>
      <p>
        On one hand, we devised in the previous section the common denominator we
are looking for: OTUs, character and character states. On the other hand, a second
important ingredient is devising what is our target in ontologies. As mentioned in
Section 2, a predominant ontology model for phenotype descriptions is the Entity-Quality
(EQ) [
        <xref ref-type="bibr" rid="ref2">Balhoff et al. 2010</xref>
        ]. An Entity refers to the “part” of the OTU being described,
which is related to one or more Qualities. In a comparison with the (C, CS) approach,
a Character comprises an Entity plus the Quality involved in the description in a single
textual sentence. A State is a complementary part of the Quality. Even though it is not a
trivial task to split Characters into their components of Entity and Quality, a first step will
SDD
      </p>
      <p>Xper2
Characters
charStateLabel
matrix</p>
      <p>row</p>
      <p>CategoricalCharacter
CodedDescription
taxon-­‐name</p>
      <p>entry
(a) Nexus</p>
      <p>Representation</p>
      <p>States
Representation
SummaryData
(c) SDD</p>
      <p>OTUs
Format
matrix
Variables</p>
      <p>Individuals
charLabels
stateLabels
be linking disperse elements referring to the same semantic concept.</p>
      <p>Departing from the key elements identified in the previous section, we can devise
the following linking discovery challenges:</p>
      <p>Which OTUs in the graph refer to the same real world OTU (link OTU-OTU)?
Which characters can be applied to each OTU (link OTU-character)?
Which states for each character can be observed in each OTU (link
OTUcharacter-state)? Conversely, which OTUs have a given character+state?
The answer to these questions will enable to integrate, summarize and compare
data concerning each OTU and each character. Therefore, it becomes possible to answer
queries like:
What are the possible colors of a Varanus tongue?
Which animals present an oval nostrils form?</p>
      <p>The discovery process is carried by graph transformations. As graphs are crucial
for our modeling approach, our method was built over graph databases. These databases
reduce the gap between how data is modeled (as graphs) and how it is stored. It is capable
of representing data structures with high abidance. Compared with relational databases,
graph databases do not require join operations because it is done implicitly traversing the
graph from node to node. Graph databases are less schema-dependent and for this reason,
they can scale more easily in size and complexity as the application evolves.</p>
      <p>
        The questions stated before were the basis to conceive the model presented in
Figure 6. We adopted the property graph model, in which nodes and relationships can
maintain extra metadata as a set of key/value pairs. Moreover, relationships are typed,
enabling to create multi-relational networks with heterogeneous sets of edges. Different
from single-relational networks, in which edges are of the same type, multi-relational
networks are more appropriate to represent complex domain models, due the variety of
relationship types in the same graph [
        <xref ref-type="bibr" rid="ref20">Rodriguez and Shinavier 2010</xref>
        ].
      </p>
      <p>In our graph model, OTUs and character-states are nodes connected by characters
(edges). Therefore the statement “V. albiguralis has a well round tail shape” becomes V.
albiguralis (node) ! tail shape (edge) ! well round (node).</p>
      <p>Type
Label
Detail</p>
      <p>OTU</p>
      <p>OTU</p>
      <p>Character
Type
Detail</p>
      <p>Character</p>
      <p>Character-State</p>
      <p>State
Type
Label
Detail</p>
    </sec>
    <sec id="sec-5">
      <title>5. Practical Experiment of Unifying Phenotypes</title>
      <p>We have implemented an automatic process to ingest SDD files into a graph database,
in order to show the linking possibilities raised by our model. In our experiments, we
use the Neo4j (http://www.neo4j.org/ ), an open-source graph database. Our data
integration processing flow is divided into the main stages: preprocessing, data ingestion, data
linkage.</p>
      <p>
        One of the problems faced in bioinformatics is related to the identification of
objects within and across repositories [
        <xref ref-type="bibr" rid="ref15">Page 2008</xref>
        ]. More precisely, an object may refer
to a taxon, gene, anatomical feature, phenotypic description, geographical location etc.
Uniquely identifying those objects is undoubtedly a key point for the success of our
proposed solution.
      </p>
      <p>
        In order to address this issue, some organizations – e.g., Universal Biological
Indexer and Organizer (uBio), Integrated Taxonomic Information System (ITIS),
Catalogue of Life (CoL), The International Plant Names Index (IPNI), National Center for
Biotechnology Information (NCBI) etc. – incorporated into their projects the Life
Science Identifiers (LSIDs), which was proposed by the Object Management Group (OMG)
(http://www.omg.org/ ). LSID is a persistent, location-independent resource identifier,
whose purpose is to uniquely identify biological resources [
        <xref ref-type="bibr" rid="ref5">Clark et al. 2004</xref>
        ]. The
persistent property refers to the fact that LSID identifiers are unique, can be assigned to only
one object forever and they never expire. The location-independent property specifies
that each authority locally creates LSIDs and they are the responsible to guaranteeing the
uniqueness of LSIDs.
      </p>
      <p>We applied LSIDs to unify OTUs in the graph referring to the same real world
object. In order to find a valid LSID, we adopted the Global Names Resolver (GNR) web
service (http://resolver.globalnames.org/ ) that executes exact or fuzzy matching against
canonical forms of scientific names in 170 distinct data sources. The Canonical form (cf)
is the simplest, most complete and unambiguous form of a name. The Canonical form of
scientific names consists of the genus and species – when applied – with no authorship,
rank, nomenclatural annotation or subgenus.</p>
      <p>
        Our system used three of the six types of matching offered by the GNR resolver:
(i) exact matching; (ii) exact matching of canonical forms – this process reduce a given
name to its canonical form and checks it with an exact match; (iii) fuzzy matching of
canonical forms – uses a modified version of the TaxaMatch algorithm [
        <xref ref-type="bibr" rid="ref19">Rees 2008</xref>
        ] and it
intends to work around misspellings errors. It does a fuzzy match of the canonical form
of a given name – even with mistakes – against spellings considered correct. The GNR
resolver reports the matching quality (“confidence score”) for each match.
      </p>
      <p>The matching module of the system is still a work in progress, but we already
have obtained some relevant results to show the viability of our approach. From the
LIS knowledge base we collected 7 distinct morphological descriptions: genus Varanus;
species Varanus gouldii, Varanus timorensis, Varanus auffenbergi and Varanus scalaris;
species groups Varanus indicus, Varanus prasinus, Varanus salvator; and Autralian
spinytailed monitor lizards. Through Xper2 those morphological descriptions were exported to
the SDD format and imported into the graph database, with no preprocessing. Figure 7(a)
shows an overview of the resulting graph without labels. We can note the
disconnectedness of the graph (7-partite graph). On the other hand, Figure 7(b) shows the same
knowledge after employing the LSID unification. The graphs became connected. Before
applying the LSID unification the graph had 74 distinct taxonomic units (TUs). After
performing the LSID unification its total reduced to 44 TUs, i.e., 30 taxonomic units (40%)
were recurring and were integrated in a single node.</p>
      <p>The next step is to link equivalent characters of the same OTU, enabling
integration of states of the same character. In the present stage of this research we apply a simple
matching algorithm. One example of our preliminary results is presented in the diagram
of Figure 8. As can be seen, our algorithm was able to unify all “nuchal scales”
characters, by defining the same type to the edges. Moreover, we unified and congregated the
possible states observed for this character across different description files.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>Several initiatives propose to relate phenotype descriptions with ontologies to enable a
semantic integration. The challenge is how to expand and revise the ontology while new
descriptions were created. Tools which annotate descriptions with ontologies address them
as an external artifact crafted apart, disregarding the synergy between building an
ontolVaranus bogerti
Varanus beccarii</p>
      <p>Varanus prasinus
Varanus komodoensis
nuchal scales
nuchal scales
nuchal scales
nuchal scales
nuchal scales
nuchal scales
nuchal scales</p>
      <p>Prasinus.sdd
smooth, unkeeled
granular to slightly keeled</p>
      <p>strongly keeled
triangular keeled, hull-shaped
same size than head scales
bigger than head scales</p>
      <p>Varanus.sdd
(a) Graph 7-partite
(b) Connected graph
ogy and using it. [Shirky 2005] emphasizes the importance of the semantics organically
built by a community, where a binary categorization approach – in which a concept A “is”
or “is not” part of a category B – to a probabilistic approach – in which a percentage of
people relates A to B. This work contributes in this direction. Inspired by previous work,
which explores latent semantics in folksonomies, this work analyzes standards to describe
phenotypes to find a common denominator, which is the bases to link descriptions.</p>
      <p>The main contribution of this work is to create the basis to exploit the latent
semantics in the descriptions. The viability and the potential of our approach were tested by
experiments. These experiments are the first steps to exploit a bigger latent semantics
scenario. Moreover, having the capability of integrating knowledge around taxonomic units
will enable, for instance, evolutionary biologists to generate new research questions, gain
predictive insight or confront evolutionary hypotheses. More complete answers might be
provided as new data sources are integrated.</p>
      <p>
        Our representation in a graph database is aligned with the
RDF [
        <xref ref-type="bibr" rid="ref14">Manola and Miller 2004</xref>
        ] graph-based representation, which will be the next
step to achieve the third layer. The challenge will be to map labels of
character/characterstates in RDF properties/values. The unification of characters and states, as shown on
this preliminary work, is a first and high relevant step for this mapping. Since several
ontologies related to phenotype descriptions are in OWL, the relations discovered in
our graph can subsidize a better matching of labels and concepts in OWL ontologies by
confronting relations. For example, to enhance the match of a character label (in the
graph database) with an OWL property, it is possible to consider the states allowed by
the character, confronting them with the property range (values allowed by the property).
      </p>
      <p>There are several possible ways to extend this work. One possible way is to
incorporate morphological descriptions stored in other knowledge bases, e.g., MorphoBank
(http://morphobank.org/ ) or Dryad (http://datadryad.org/ ). Another direction is to
investigate correlations between State nodes and ontology terms.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgment</title>
      <p>Work partially financed by (CNPq 138197/2011-3), the Microsoft Research FAPESP
Virtual Institute (NavScales project), CNPq (MuZOO Project and PRONEX-FAPESP),
INCT in Web Science(CNPq 557.128/2009-9) and CAPES, as well as individual grants
from CNPq.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alves</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          and Santanche`,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Folksonomized Ontology and the 3E Steps Technique to Support Ontology Evolvement</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>18</volume>
          (
          <issue>1</issue>
          ):
          <fpage>19</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Balhoff</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dahdul</surname>
            ,
            <given-names>W. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kothari</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapp</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lundberg</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Midford</surname>
            ,
            <given-names>P. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Westerfield</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vision</surname>
          </string-name>
          , T. J. (
          <year>2010</year>
          ).
          <article-title>Phenex: Ontological annotation of phenotypic diversity</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <volume>5</volume>
          (
          <issue>5</issue>
          ):
          <fpage>e10500</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bell</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hey</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Szalay</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <volume>323</volume>
          (
          <issue>5919</issue>
          ):
          <fpage>1297</fpage>
          -
          <lpage>1298</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Ciccarelli</surname>
            ,
            <given-names>F. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doerks</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Von Mering</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creevey</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bork</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Toward automatic reconstruction of a highly resolved tree of life</article-title>
          .
          <source>Science</source>
          ,
          <volume>311</volume>
          (
          <issue>5765</issue>
          ):
          <fpage>1283</fpage>
          -
          <lpage>1287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Liefeld</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Globally distributed object identification for biological knowledgebases</article-title>
          . Briefings in bioinformatics,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>59</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Dahdul</surname>
            ,
            <given-names>W. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balhoff</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engeman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grande</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hilton</surname>
            ,
            <given-names>E. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kothari</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapp</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lundberg</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Midford</surname>
            ,
            <given-names>P. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vision</surname>
          </string-name>
          , T. J.,
          <string-name>
            <surname>Westerfield</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P. M.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Evolutionary characters, phenotypes and ontologies: Curating data from the systematic biology literature</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <volume>5</volume>
          (
          <issue>5</issue>
          ):
          <fpage>e10708</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Gibas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Jambeck</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Developing bioinformatics computer skills.</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          , Inc.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Gkoutos</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mallon</surname>
            ,
            <given-names>A.-M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hancock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Using ontologies to describe mouse phenotypes</article-title>
          .
          <source>Genome Biology</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>R8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Hagedorn</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Structuring Descriptive Data of Organisms - Requirement Analysis and Information Models</article-title>
          .
          <source>PhD thesis</source>
          , Universita¨t Bayreuth,Fakulta¨t f u¨r Biologie,
          <source>Chemie und Geowissenschaften.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Lenzerini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Data integration: A theoretical perspective</article-title>
          .
          <source>In Proceedings of the twenty-first ACM SIGMOD-SIGACT-SIGART symposium on Principles of database systems</source>
          , pages
          <fpage>233</fpage>
          -
          <lpage>246</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P. M.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Integrating evolution and development: the need for bioinformatics in evo-devo</article-title>
          .
          <source>BioScience</source>
          ,
          <volume>56</volume>
          (
          <issue>4</issue>
          ):
          <fpage>301</fpage>
          -
          <lpage>309</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Mabee</surname>
            ,
            <given-names>P. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cronk</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkoutos</surname>
            ,
            <given-names>G. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haendel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Segerdell</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Westerfield</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Phenotype ontologies: the bridge between genomics and evolution</article-title>
          .
          <source>Trends in ecology &amp; evolution, 22</source>
          (
          <issue>7</issue>
          ):
          <fpage>345</fpage>
          -
          <lpage>350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Maddison</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swofford</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Maddison</surname>
            ,
            <given-names>W. P.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>Nexus: An extensible file format for systematic information</article-title>
          .
          <source>Systematic Biology</source>
          ,
          <volume>46</volume>
          (
          <issue>4</issue>
          ):
          <fpage>590</fpage>
          -
          <lpage>621</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Manola</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2004</year>
          ). RDF Primer - W3C
          <string-name>
            <surname>Recommendation</surname>
          </string-name>
          .
          <source>Technical report, W3C.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Biodiversity informatics: the challenge of linking data and the role of shared identifiers</article-title>
          .
          <source>Briefings in Bioinformatics</source>
          ,
          <volume>9</volume>
          (
          <issue>5</issue>
          ):
          <fpage>345</fpage>
          -
          <lpage>354</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Parr</surname>
            ,
            <given-names>C. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guralnick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cellinese</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>R. D.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Evolutionary informatics: unifying knowledge about the diversity of life</article-title>
          .
          <source>Trends in ecology &amp; evolution, 27</source>
          (
          <issue>2</issue>
          ):
          <fpage>94</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Pimentcl</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Riggins</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>1987</year>
          ).
          <article-title>The nature of cladistic data</article-title>
          .
          <source>Cladistics</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>201</fpage>
          -
          <lpage>209</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Quan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Improving life sciences information retrieval using semantic web technology</article-title>
          .
          <source>Briefings in bioinformatics</source>
          ,
          <volume>8</volume>
          (
          <issue>3</issue>
          ):
          <fpage>172</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Rees</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Taxamatch, a ”fuzzy” matching algorithm for taxon names, and potential applications in taxonomic databases</article-title>
          . In Weitzman, A. and
          <string-name>
            <surname>Belbin</surname>
          </string-name>
          , L., editors,
          <source>Provisional Abstracts of the 2008 Annual Conference of the Taxonomic</source>
          Databases Working Group, Fremantle,
          <string-name>
            <given-names>Australia. Biodiversity</given-names>
            <surname>Information</surname>
          </string-name>
          <article-title>Standards (TDWG) and the Missouri Botanical Garden</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Rodriguez</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Shinavier</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Exposing multi-relational networks to singlerelational network analysis algorithms</article-title>
          .
          <source>Journal of Informetrics</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <fpage>29</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Ung</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Causse</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>and Vignes</given-names>
            <surname>Lebbe</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2010a</year>
          ).
          <article-title>Xper2: managing descriptive data from their collection to e-monographs.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Ung</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubus</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>Zaragu¨ eta-</article-title>
          <string-name>
            <surname>Bagils</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Vignes-Lebbe</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2010b</year>
          ).
          <article-title>Xper2: introducing e-taxonomy</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>26</volume>
          (
          <issue>5</issue>
          ):
          <fpage>703</fpage>
          -
          <lpage>704</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Vos</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balhoff</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caravas</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holder</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapp</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maddison</surname>
            ,
            <given-names>W. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Midford</surname>
            ,
            <given-names>P. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Priyam</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sukumaran</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , et al. (
          <year>2012</year>
          ).
          <article-title>Nexml: rich, extensible, and verifiable representation of comparative data and metadata</article-title>
          .
          <source>Systematic Biology</source>
          ,
          <volume>61</volume>
          (
          <issue>4</issue>
          ):
          <fpage>675</fpage>
          -
          <lpage>689</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>