<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Representing bioinformatics datatypes using the OntoDT ontology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pan cˇe Panov</string-name>
          <email>pance.panov@ijs.si</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Larisa Soldatova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sasˇ o Dzˇ eroski</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Brunel University London, Department of Computer Science</institution>
          ,
          <addr-line>Kingston Lane, UB8 3PH, Uxbridge</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Knowledge Technologies, Jozˇ ef Stefan Institute</institution>
          ,
          <addr-line>Ljubljana</addr-line>
          ,
          <country country="SI">Slovenia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Fig. 1. Representation of datatypes in OntoDT</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Data processing is at the heart of science. Hence, the problem of data typing is an important problem that has been addressed from different aspects and in different forms. For example, the Research Data Alliance1 (RDA), whose major goal is to speed up the international data-driven innovation and discovery by facilitating research data sharing and exchange, has identified that the problem of data typing is an important problem that deserves attention. For this purpose, the RDA formed a Data Type Registry (DTR) working group with the goal to: compile a set of use cases for datatype use and management, formulate a data model and expression for datatypes, design a functional specification for type registries, and propose a federation strategy among multiple type registries. In data mining research it is impossible to efficiently connect parts of workflows (semi-) automatically, such as data pre-processing and data mining, perform analysis of the research results and communicate the research outputs, without machine-processable representation of datatypes and their properties. Hence, there is a need for a standardized semantically-defined and machine amenable representation of scientific datatypes to support crossdomain applications. Unfortunately, the existing representations of datatypes do not fully address such a need. To address this gap, we built an generic ontology for the representation of scientific knowledge about datatypes, named OntoDT (Panov et al., 2015). 1 http://rd-alliance.org/ 2 http://tinyurl.com/qdua9f7 3 http://tinyurl.com/nmjnlw2</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
    </sec>
    <sec id="sec-2">
      <title>ONTODT: ONTOLOGY OF DATATYPES</title>
      <p>OntoDT defines the meaning of the key entities and represents the
knowledge about datatypes in a machine friendly way. The OntoDT
ontology is based on the latest revised version of the ISO/IEC
114042 standard for datatypes. The design of the OntoDT ontology
follows best practices in ontology engineering, such as the OBO
Foundry principles. We used the Information Artifact Ontology3
(IAO) to define the upper level classes and re-used existing
ontological resources, such as Open Biomedical Ontologies.</p>
      <p>The OntoDT ontology defines the basic entities (see Fig. 1), such
as datatype, properties of datatypes, value space, and characterizing
operations. We also define a taxonomy of datatypes. The top-level
ontology classes include primitive datatypes, generated datatypes,
and user defined datatypes. Primitive datatypes are defined by
explicit specification and are independent of other datatypes.
Generated datatypes are syntactically and semantically dependent
on other datatypes, and are specified implicitly with datatype
generators. User defined datatypes are defined by a datatype
declaration and allow defining additional identifiers and refinements
to both primitive and generated datatypes. At the lower levels, the
datatypes are distinguished with respect to their datatype properties.</p>
      <p>
        OntoDT was used within an Ontology of core data mining
entities for constructing taxonomies of datasets, data mining tasks,
generalizations and data mining algorithms
        <xref ref-type="bibr" rid="ref3">(Panov et al., 2014)</xref>
        .
Furthermore, OntoDT can be used for annotation and querying
machine learning dataset repositories. OntoDT can also improve
the representation of datatypes in the BioXSD exchange format for
basic bio-informatics types of data. The generic nature of OntoDT
enables it to support a wide range of other applications, especially in
combination with other domain specific ontologies: the construction
of data mining workflows, annotation of software and algorithms,
semantic annotation of scientific articles, etc. OntoDT is open
source and is available at http://www.ontodt.com/ and at
BioPortal (http://bioportal.bioontology.org/).
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>BIOINFORMATICS DATATYPES</title>
      <p>OntoDT is a generic ontology and it allows easy extensions
to represent domain specific datatypes. This can be done by
directly extending the OntoDT datatype taxonomy and defining the
semantic meaning of the domain datatypes by linking them to the
corresponding entities in other domain ontologies. For example, we
can define an amino-acid sequence datatype as a subclass of the
character sequence datatype class (which is a sequence datatype
having characters as its base type). Its semantic meaning can be
defined via the IS-ABOUT relation to the amino acid sequence entity
Panov et al
provided by the National Cancer Institute Thesaurus4. In this way,
OntoDT can be used for representation of bioinformatics datatypes.</p>
      <p>
        Currently, BioXSD is used to define the basic bio-informatics
types of data (
        <xref ref-type="bibr" rid="ref2">Kalasˇ et al., 2010</xref>
        ). BioXSD does not support
arbitrary datatypes and it does not provide a clear framework
for the representation of the semantic meaning of the data. We
propose to enhance the representation of bioinformatics datatypes
by exploiting the rigorous taxonomy of datatypes defined in OntoDT
and the framework for the representation of semantic meanings
adopted by OntoDT. OntoDT is fully interoperable with OBO
bioontologies because it was developed by following the OBO Foundry
recommendations and therefore it fully supports the representation
of the semantic meaning of the data by the corresponding entities
defined in domain-specific bio-ontologies.
      </p>
      <p>For example, the BioXSD datatype sequence represents a string
of 1-letter coded nucleotides or amino-acids. A sequence record is
a datatype containing a sequence, and optionally some metadata
about the sequence (for the purpose of identification). The semantic
meanings of the terms sequence and nucleotide are curtail for the
capturing of the semantic meaning of the data of the datatype
sequence. However, the sequence datatype is not explicitly linked
to the classes nucleotide and amino-acid defined in the ChEBI
ontology5, recommended by OBO Foundry as a reference ontology.</p>
      <p>In Fig. 2, we present the extension of OntoDT to represent the
biosequence datatype from BioXSD. We represent the bio-sequence
datatype class as a subclass of the character sequence datatype
class with the defined semantic meaning in the NCI Thesaurus
and the EDAM ontology6. In order to define the nucleotide and
amino acid sequences datatypes, we define two subclasses of the
character datatype class: nucleotide character datatype and amino
acid character datatype. In order to define their semantic meaning,
we explicitly link them to the nucleotide and amino acid classes
from the ChEBI ontology. Consequently, the bio-sequence datatype
class has two subclasses: nucleotide sequence datatype and amino
4 http://ncit.nci.nih.gov/
5 http://www.ebi.ac.uk/chebi/
6 http://edamontology.org/
acid sequence datatype. Furthermore, both datatypes have two
subclasses, depending on whether they include ambiguous bases (in
the case of nucleotides) or ambiguous and additional residues (in the
case of amino acids). For example, the nucleotide sequence datatype
class has two subclasses: nucleotide sequence with ambiguous bases
(general nucleotide sequence in BioXSD) and nucleotide sequence
without ambiguous bases (nucleotide sequence in BioXSD).</p>
      <p>In a simmilar way, we represent the bio-sequence record datatype
class as a subclass of the record datatype class. This datatype
is defined by a record generator and the bio-sequence-field-list.
As defined in BioXSD, the datatype contains a bio-sequence as a
mandatory component and a set of metadata (such as name, note,
species, translationalData, reference, inlineBaseQuality) as
nonmandatory components. In OntoDT, we model the bio-sequence
field component class is as a role of the bio-sequence datatype.</p>
      <p>
        BioXSD uses a combined approach of a pure XML Schema
annotated by a data-type ontology using Semantic Annotations for
Web Services Description Language7 (WSDL) and XML Schema.
SAWSDL defines a set of extension attributes for the WSDL
and XML Schema definition languages. Application of attributes
allows the description of additional semantics by using references
to conceptual semantic models, e.g., ontologies. BioXSD datatypes
are annotated with terms from the EDAM ontology
        <xref ref-type="bibr" rid="ref1">Ison et al.
(2013)</xref>
        using SAWSDL. In the same way, BioXSD datatypes can
be annotated with OntoDT terms. For example, by annotating the
datatype bio-sequence record from BioXSD with terms from the
OntoDT ontology, the web services using this format would have
the information that bio-sequence record is in fact a record datatype
that is heterogeneous and has components, its values are unordered,
it has fixed size, and each component can be accessed by keying.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSION</title>
      <p>The use case presented in this extended abstract demonstrates
that OntoDT provides logically consistent representation of
bioinformatics datatypes from BioXSD and enables an accurate
representation of the semantic meanings of the data of specified
datatypes. OntoDT has been designed as a generic and
comprehensive ontology of datatypes and consequently any
datatype from other resources can also be represented by OntoDT.
We suggest that OntoDT can serve as a reference model for
the consistent representation of datatypes used within biomedical
domains and wider.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonassen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolser</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uludag</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McWilliam</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Rice</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>29</volume>
          (
          <issue>10</issue>
          ),
          <fpage>1325</fpage>
          -
          <lpage>1332</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Kalasˇ</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puntervoll</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joseph</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bartasˇevicˇiute</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Topfer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkataraman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bryne</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blanchet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rapacki</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Jonassen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>BioXSD: the common data-exchange format for everyday bioinformatics web services</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>26</volume>
          (
          <issue>18</issue>
          ),
          <fpage>i540</fpage>
          -
          <lpage>i546</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Panov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soldatova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and Dzˇeroski,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Ontology of core data mining entities</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          ,
          <volume>28</volume>
          (
          <issue>5-6</issue>
          ),
          <fpage>1222</fpage>
          -
          <lpage>1265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Panov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soldatova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and Dzˇeroski,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Generic ontology of datatypes</article-title>
          .
          <source>Information Sciences</source>
          .
          <article-title>(accepted for publication)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>7 http://www.w3.org/TR/sawsdl</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>