<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bottom-up taxon characterisations with shared knowledge: describing specimens in a semantic context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Patrick Plitzner</string-name>
          <email>p.plitzner@bgbm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tilo Henning</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Müller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton Güntsch</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naouel Karam</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norbert Kilian</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Botanic Garden and Botanical Museum Berlin, Freie Universität Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics and Computer Science, Freie Universität Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Using the angiosperm order Caryophyllales, we will provide an exemplar use case on optimizing the taxonomic research process with respect to delimitation and characterisation (“description”) of taxa using the the European Distributed Institute of Taxonomy (EDIT) Platform for Cybertaxonomy. The workflow for sample data handling of the EDIT platform will be extended: Character data (data on genotypic and phenotypic characters of any type, here focusing on morphology) will be captured and stored in structured form. The structure consists of character and character state matrices for individual specimens instead of taxa, which shall allow to generate taxon characterisations by aggregating the data sets for the individual specimens included. To ensure data integrity, especially for the aggregation process, semantic web technologies will be used to establish and continuously elaborate expert community-coordinated exemplar vocabularies with term ontologies and explanations for characters and states. In cooperation with the "German Federation for Biological Data" (GFBio), the GFBio Terminology Service is used for publishing the ontologies via a public API. The EDIT platform will be extended to use and integrate the GFBio Terminology Service in order to work with the latest version of the ontology used for specimen respective taxon descriptions.</p>
      </abstract>
      <kwd-group>
        <kwd>descriptive data</kwd>
        <kwd>e-taxonomy</kwd>
        <kwd>terminology management</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In a precursor project [1, 2], we have implemented a workflow for processing
specimen-related metadata on the European Distributed Institute of Taxonomy (EDIT)
Platform for Cybertaxonomy [3], a comprehensive taxonomic data management and
publication environment that offers a collection of tools and services and works as a service
provider to support taxonomic workflows, publishing, data storage and exchange, etc.
The aim was to organise the links between (a) samples of individual organisms
collected, (b) research data obtained from them, (c) specimens of these individuals
deposited in research collections, and (d) taxon assignments (“identifications”) of the
investigated individuals.</p>
      <p>On this basis, the current project will optimise the taxonomic research process with
respect to delimitation and characterisation (“description”) of taxa.</p>
      <p>Working on the angiosperm order Caryophyllales [4], character data (mainly
morphological data) of individual specimens will be recorded and stored in the underlying
“Common Data Model” (CDM) [5] compliant data store of the platform. For specimen
descriptions, a community-developed expert ontology backed by the GFBio
terminology service for ontology management is being developed and used to ensure data
integrity. In a final step, data aggregation of the individual character data sets assisted by
the terminology service will generate automated descriptions on taxon level.</p>
      <p>This project combines two major scientific areas, semantic descriptions and taxon
characterization both of which are crucial for sustainable scientific work. Taxon
characterizations on specimen level allow for generated taxonomic delimitation. However,
this is partly a subjective work leading to different definitions for certain features (leaf
colour is “reddish green” vs “greenish red”). To align different characterizations the
combination with semantically defined terms will relate existing definitions and also
unify newly created ones by proposing existing terms.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Terminology service</title>
      <p>One of the project goals is to create an ontology for specimen descriptions which should
be used and developed collaboratively. This ontology should be made publicly
available to increase the reach and usage of the semantic concepts developed for it. The
GFBio terminology service [6], which is simultaneously being implemented, supports
working with formal ontologies, taxonomies or other Semantic Web compliant
collections of terms. It will be used to store and publish the aforementioned ontology. The
service, as seen in Fig 1, provides a web service interface to support various requests
related to retrieving semantic information from the stored ontologies. Another
important feature is the mapping of internal and external terminological resources which
promotes even more the collaborative work on ontologies.</p>
      <p>Fig 1 Overview of the GFBio terminology service architecture</p>
    </sec>
    <sec id="sec-3">
      <title>Specimen Description workflow</title>
      <p>Ontologies backed by the terminology service will be created, managed, used and
extended during the entire workflow for specimen based data acquisition and taxon
descriptions. Three main applications can be identified, all of which will be integrated
into the EDIT platform as part of the current project (see Fig 2)</p>
      <p>Fig 2 The EDIT platform uses the API of the terminology service to integrate the
terminology services into three applications: 1) the term editor which allows editing on a
synced copy of the ontology, 2) the character editor where the user defines taxon specific term
hierarchies for structures, properties and their corresponding states and 3) the character
matrix which serves for the character-based description of single specimens.
3.1</p>
      <sec id="sec-3-1">
        <title>Ontology Management</title>
        <p>Ontology editing facilities are implemented into the EDIT platform using the API of
the terminology service. The platform itself provides a user and rights management
which will serve for collaborative work on the ontology preparation and maintenance.
Additionally, the CDM as the storage model adds more fine-grained meta information
to the development process. It allows tracking changes i.e. allowing a versioning
mechanism and also an extended documentation via annotations and notes is possible.</p>
        <p>Working on the ontology within the platform will be done on a synced copy of the
data. The CDM will be extended to support the linkage of terms and their relations as
well as their semantic concept in the remote ontology provided by the terminology
service.</p>
        <p>A term editor based on the EDIT platform is used to visualise and edit the synced
copy.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Creating the descriptive data set/Character editor</title>
        <p>For a comprehensive morphological analysis of a taxon in general as well as
specimenwise, a well-defined, established terminology is essential that has already been widely
used in the respective plant group. The individual botanist must be able to choose the
necessary terms from a vocabulary that is persistently embedded in or linked to a stable
term-ontology (e.g. The Plant Ontology [7]).</p>
        <p>
          To describe the morphological characters observed, composite terms are used
following the tripartite principle proposed by Diederich [8] and realised in the Prometheus
model [9, 10]. That means that characters are composed of three single terms that
belong to different categories: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) plant structures, defining the morphological structure
of a plant organism from root to flower, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) properties, describing the morphological
aspects of the plant structures, (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) states for setting the quantitative or categorical space
of the properties.
        </p>
        <p>Structures and properties will be stored in tree structures into CDM-based data
stores. The tree structure allows for designing taxonomic group specific hierarchies and
dependencies between the single terms. The compilation of structure tree, property tree
and states connected to a taxon is called a descriptive data set.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Character matrix and aggregation</title>
        <p>The first two steps dealt with the conceptual creation of the descriptive data set by
evaluating what terms of the ontology are needed, how they are ordered and how their
boundaries are defined. The final step is the actual description of specimens including
the creation of characters and measuring their states.</p>
        <p>As pointed out in the previous chapter, data triplets based on the Prometheus model
are used. Every single character that describes a certain feature of the specimen is built
up from a structure term and a property term. The range of the property term itself is
limited by the states assigned to it.</p>
        <p>The specimen descriptions are edited in a character matrix combining all specimens
associated with the taxonomic group of the current descriptive data set with the
characters created to describe the morphological features. The matrix can be seen as a table
with ordered rows which will be built up by the characters that were previously created
to describe the taxon. The columns will be the specimens belonging to that certain
taxon. The order of the characters also provides semantic information. There are, for
example, character that cannot exist because the overall structure to which they belong
does not exist as well as a more general character may already define the boundaries of
a sub character.</p>
        <p>The editing process will be enriched with the semantic knowledge about the terms.
This enables rules for value hierarchies, data entry assistance through semantic
documentation, data validation, etc.</p>
        <p>The ordering of state information into a character matrix enables the procedure of
generating taxon descriptions via an aggregation algorithm. Specimen descriptions will
be comparable to each other because of structured character data organization. Single
characters and their states are semantically defined by the underlying ontology
describing what they are and how to interpret their values. The semantic knowledge also assists
when comparing or merging character data from different sources.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>The EDIT platform in combination with the GFBio terminology service creates a
capable environment for the process of a specimen-based and dynamic description of taxa
using character data. The descriptive data set as a data structure connects the “raw”
specimen character data to a taxonomic group, making data aggregation possible which
allows the generation of automated taxon descriptions. Each application of the
workflow is based on the platform and the CDM so that the user rights and roles management
system can be set up specifically for each task by granting access only to those users
that are authorized.</p>
      <p>In any step of the workflow it is common that requests to change or edit the ontology
will come up. The CDM provides the link to the synced copy of the ontology but
anyway, in a future step, change and versioning strategies should be discussed in more
detail as there are still no established solutions to this problem.</p>
      <p>Another advantage of working with semantic technology is reasoning. This will
especially be of interest during the aggregation process when dealing with conflicting data
or generated taxon descriptions vs. descriptions from literature.
4. Borsch T, Hernandez-Ledesma P, Berendsohn WG, Flores-Olvera H, Ochoterena H,
Zuloaga FO, v. Mering S, Kilian N (2015) An integrative and dynamic approach for
monographing species-rich plant groups—building the global synthesis of the angio-sperm order
Caryophyllales. Perspect Plant Ecol Evol Syst 17: 84–300.
doi.org/10.1016/j.ppees.2015.05.003
5. Anonymous. (2008) Common Data Model.
http://dev.e-taxonomy.eu/trac/wiki/Common</p>
      <p>
        DataModel (25 July 2017, date last accessed).
6. Naouel Karam, Claudia Müller-Birn, Maren Gleisberg, David Fichtmüller, Robert
Tolksdorf, Anton Güntsch: A Terminology Service Supporting Semantic Annotation, Integration,
Discovery and Analysis of Interdisciplinary Research Data.
DatenbankSpektrum 16(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ): 195-205 (2016)
7. The Plant Ontology. http://planteome.org/
8. Diederich J (1997) Basic properties for biological databases: character development
and support. Math Computer Model 25: 109–127.
9. Pullan MR, Watson MF, Kennedy JB, Raguenaud C, Hyam R (2000) The Prometheus
Taxonomic Model: A Practical Approach to Representing Multiple Classifications. Taxon
49(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ): 55–75.
10. Pullan MR, Armstrong KE, Paterson T, Cannon A, Kennedy JB, Watson MF, McDonald S,
Raguenaud C (2005) The Prometheus Description Model: an examination of the
taxonomic description-building process and its representation. Taxon 54(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ): 751–765.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kilian</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henning</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plitzner</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Güntsch</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stöver</surname>
            <given-names>BC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Müller</surname>
            <given-names>KF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berendsohn</surname>
            <given-names>WG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borsch</surname>
            <given-names>T</given-names>
          </string-name>
          (
          <year>2015</year>
          )
          <article-title>Sample data processing in an additive and reproducible taxonomic workflow by using character data persistently linked to preserved individual specimens</article-title>
          .
          <source>Database</source>
          <year>2015</year>
          :
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          . doi:
          <volume>10</volume>
          .1093/database/bav094
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Campanula</given-names>
            <surname>Data</surname>
          </string-name>
          <article-title>Portal</article-title>
          . http://campanula.e-taxonomy.net/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Berendsohn</surname>
            <given-names>WG</given-names>
          </string-name>
          (
          <year>2010</year>
          )
          <article-title>Devising the EDIT Platform for Cybertaxonomy</article-title>
          . In: Nimis L,
          <string-name>
            <surname>Vignes-Lebbe R</surname>
          </string-name>
          (eds).
          <article-title>Tools for Identifying Biodiversity: Progress and Problems</article-title>
          . roceedings of the International Congress, Paris,
          <fpage>20</fpage>
          -
          <issue>22</issue>
          <year>September 2010</year>
          .
          <article-title>EUT Edizioni niversita` di Trieste</article-title>
          , Trieste, pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>