<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NGBO: The Introduction of -omics Data to Biobanking Ontology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dalia Alghamdi</string-name>
          <email>Dalia.Alghamdi@bccdc.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Damion M. Dooley</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gurinder Gosal</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emma J. Griffiths</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William W.L. Hsiao</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BC Centre for Disease Control Public Health Laboratory</institution>
          ,
          <addr-line>Vancouver, BC V5Z 4R4</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>King Fahad Medical City</institution>
          ,
          <addr-line>Riyadh 59046</addr-line>
          ,
          <country country="SA">Saudi Arabia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Simon Fraser University</institution>
          ,
          <addr-line>Burnaby, BC V5A 1S6</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of British Columbia</institution>
          ,
          <addr-line>Vancouver BC V6T 1Z4</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A biobank contains a collection of biological samples, along with associated medical information of sample donors, which can be used for different types of studies. Given the wealth of information that can be derived from stored information and biological materials, there is a pressing need for structuring biobank data for more computeramenable analyses. The utility of first generation biobanks was originally evaluated simply based on the number of samples that they contained. Currently, the value of biobank data lies in how it can linked with other molecular and clinical data (“-omics data”), to provide new insights into health and disease. Linking data has thus far, however, proven challenging due to unstructured and incompatible data types. Here, we describe the development of a Next-Generation biobanking ontology (NGBO) (https://github.com/Dalalghamdi/NGBO) that is capable of supporting both Biospecimen processing, management, storage and retrieval infrastructure, and acting as a knowledge hub for an integrated clinical and translational research ecosystem integrating omics data. NGBO harmonizes the instrumentation and procedures used to prepare and process specimens, and also covers terminology used to describe computational biology algorithms, analytical tools, electroniccommunication protocols, in vitro assays. Laboratories, investigators, and other biobanks would also benefit from the knowledge contained in the ontology, by the means of using NGBO a biobank data catalogue that can be used to map any existing unstructured data.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>Biobank</kwd>
        <kwd>Next generation sequencing</kwd>
        <kwd>data harmonization</kwd>
        <kwd>data integration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        biomedical research began centuries ago, but has since undergone dramatic change
due to technology advancements in storage techniques, sample information retrieval,
as well analysis of specimen material. As such, the informatics needs of modern
biobanks are far more complex than past repositories that often captured only the date
or location of a sample collection. Besides sample metadata, modern biobanks cover
the storage and management of more complex data generated from high-throughput
biological studies such as proteomic, genomics and other –omics studies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
International and national collaborations can improve the value of biobanks, but this
requires harmonizing the data fields and values across biobanking applications. There
have been several efforts to achieve collaboration and data sharing among various
national biobanks, for instance, the public population project in genomics (P3G)
(http://www.p3g.org) has previously tackled building biobanking resources as well as
data cataloguing and harmonization for data integration [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Still, biobanks remain
heterogeneous when it comes to their design, usage, size and types of the samples. It
is possible to link the samples to data records from expansive epidemiological
collections and family histories. However, it is laborious to manually harmonize the
terms across different biobanks. Furthermore, if data harmonization is conducted
individually, inconsistencies often arise. A key development in facilitating data
standardization is the application of ontology, a semantic web technology [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Semantic Web is the best practices and sets of standards used to share data and
meanings (semantics) of data over the web. The formal and machine-readable
definitions and axioms made it is possible to come up with automated querying
systems to facilitate faster, easier, and more accurate ways to share and reuse data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Semantic Web OWL ontology is a popular technology choice for representing
terminology and data structure relationships. If one is to use ontology, it becomes
possible to establish vocabularies necessary to model a problem or activity domain. In
the model, there are objects and concepts contained in specific areas and relations
describing how they are related [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Ontologies play a role in promoting the
realization of Semantic Web.
      </p>
      <p>
        Brochhausen et al. proposed the ontology for biobanking (OBIB) in 2016. The OBIB
was created through the merging of two biobank technologies including the Biobank
Ontology (BO) and the Ontologized Minimum Information About Biobank Data
Sharing (OMIABIS) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. BO and OMIABIS focus on specimen description and
biobanking administration respectively. The biomedical and biological research has
progressed to a level where the quantity and the types of samples kept are no longer
used in measuring the prowess of biobanks. Instead, measurements are based on the
extent that the samples and metadata are used. Biobanks fall into the category of
infrastructure used in research. Thus, their aim is to support scientific processes.
The integration of more knowledge domains into the biobanking ontology is
necessary for advancing the use of biobanks. We aim to integrate –omics data, with
the creation of Next-Generation Biobanking Ontology (NGBO) to deal with various
scientific research and personalized medicine requirements. To incorporate -omics
knowledge, we model the processes and entities related to diagnostic molecular
pathology procedures, including sample handling, phenotype characterization,
computational biology algorithms and analytical tools, in vitro assays, electronic
communication protocols and data coding. In addition to the technical data
provenance, sample data provenance such as patient phenotypic information, along
with the genetic data will provide the biobank users sufficient data and knowledge to
characterize the functional and pathogenic significance of genetic variants [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
NGBO is being built based on the Open Biological and Biomedical Ontologies
(OBO) foundry principles. For example, one of the OBO foundry principles is the
reuse of existing ontology to prevent re-inventing the wheel and creating multiple
representations of the same term. OWL (W3C Web Ontology Language) is used to
provide the means for data sharing and reusing between different resources in the
form of semantic application. NGBO will provide standard identifiers for classes and
relations within biobanking domain as well as definitions for all the vocabularies in
NGBO in human and machine-readable formats. Consider the class ‘input data’ as an
example, the definition of input data that can be read by human is “computer file that
has specific format and contains data that serve as input to a device or program”.
However, it could be defined using OWL language as:
      </p>
      <p>
        'is about' some entity and 'has format' exactly 1'file format'
This is one of the expressions possible for the class ‘input data’, which states that it is
about an entity and has only one file format from file format subclasses. As shown in
figure 1, NGBO re-use many terms from existing ontologies such as Bio-Assay
Ontology (BAO) and the OBO edition of the National Cancer Institute Thesaurus
(NCIT) ontology. NGBO depends primarily on the (is-a) relation between classes and
subclasses, thereby providing a hierarchy of classes that also enables inheritance of
the properties. For example, a concept “planned process” (OBI:0000011) in the
Ontology for Biomedical Investigations (OBI) is defined as “a processual entity that
realizes a plan which is the concretization of a plan specification”. Therefore, all
subclasses of planned process must inherit the definition of the process. In addition to
(isa), pre-existing relations such as (is_version_of) is used when needed. Protégé
(version 5.2.0) is used to build NGBO that is compatible with the Basic Formal
Ontology (BFO), a small upper level ontology mostly used to support information
retrieval, analysis and integration in biomedical and biological domains [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Figure 1
shows an example of selection of NGBO classes and their sub-classes.
In conclusion, by building NGBO, a next generation biobanking ontology, we are
providing a semantics infrastructure to support externalized biomedical collaborative
research by harmonizing biospecimens with their molecular makeup. It provides a
framework for reusing clear consistent terminology (classes) with their relationships
and the metadata that describe the intended meaning of these classes and
relationships.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Elger</surname>
          </string-name>
          , Bernice S., and
          <string-name>
            <surname>Arthur</surname>
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Caplan</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>“Consent and Anonymization in Research Involving Biobanks: Differing Terms and Norms Present Serious Barriers to an International Framework</article-title>
          .
          <source>” EMBO Reports</source>
          <volume>7</volume>
          (
          <issue>7</issue>
          ):
          <fpage>661</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Riegman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter H. J.</surname>
          </string-name>
          ,
          <string-name>
            <surname>Manuel</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Morente</surname>
          </string-name>
          , Fay Betsou, Pasquale de Blasio,
          <string-name>
            <surname>Peter Geary</surname>
          </string-name>
          , and Marble Arch International Working Group on Biobanking for Biomedical Research.
          <year>2008</year>
          . “
          <article-title>Biobanking for Better Healthcare</article-title>
          .”
          <source>Molecular Oncology</source>
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <fpage>213</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>De</surname>
            <given-names>Souza</given-names>
          </string-name>
          , Yvonne G.,
          <string-name>
            <given-names>and John S.</given-names>
            <surname>Greenspan</surname>
          </string-name>
          .
          <year>2013</year>
          . “Biobanking Past,
          <source>Present and Future: Responsibilities and Benefits.” AIDS</source>
          <volume>27</volume>
          (
          <issue>3</issue>
          ):
          <fpage>303</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ouellette</surname>
          </string-name>
          , Sylvie, and Anne Marie Tassé.
          <year>2014</year>
          .
          <article-title>“P(3)G - 10 Years of Toolbuilding: From the Population Biobank to the Clinic</article-title>
          .
          <source>” Applied &amp; Translational Genomics</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ):
          <fpage>36</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Andrade</surname>
            ,
            <given-names>André Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markus</surname>
            <given-names>Kreuzthaler</given-names>
          </string-name>
          , Janna Hastings, Maria Krestyaninova, and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Schulz</surname>
          </string-name>
          .
          <year>2012</year>
          . “
          <article-title>Requirements for Semantic Biobanks</article-title>
          .
          <source>” Studies in Health Technology and Informatics</source>
          <volume>180</volume>
          :
          <fpage>569</fpage>
          -
          <lpage>73</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          , et al, “
          <source>The World Wide Web,”Communications of the ACM</source>
          ,
          <year>August</year>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>T.</given-names>
            <surname>Gruber</surname>
          </string-name>
          .
          <year>1993</year>
          “
          <article-title>A translation approach to portable ontology specification”. “Knowledge Acquisition”</article-title>
          , pp.
          <fpage>199</fpage>
          -
          <lpage>220</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Brochhausen</surname>
            , Mathias, Jie Zheng, David Birtwell,
            <given-names>Heather</given-names>
          </string-name>
          <string-name>
            <surname>Williams</surname>
            , Anna Maria Masci, Helena Judge Ellis, and
            <given-names>Christian J. Stoeckert</given-names>
          </string-name>
          <string-name>
            <surname>Jr</surname>
          </string-name>
          .
          <year>2016</year>
          . “
          <article-title>OBIB-a Novel Ontology for Biobanking</article-title>
          .
          <source>” Journal of Biomedical Semantics</source>
          <volume>7</volume>
          (May):
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Johnston</surname>
            <given-names>JJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Biesecker</surname>
            <given-names>LG</given-names>
          </string-name>
          .
          <article-title>Databases of genomic variation and phenotypes: existing resources and future needs</article-title>
          .
          <source>Human Molecular Genetics</source>
          .
          <year>2013</year>
          ;
          <volume>22</volume>
          (
          <issue>R1</issue>
          ):
          <fpage>R27</fpage>
          -
          <lpage>R31</lpage>
          . doi:
          <volume>10</volume>
          .1093/hmg/ddt384.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Arp</surname>
            , Robert,
            <given-names>Barry</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            , and
            <given-names>Andrew D.</given-names>
          </string-name>
          <string-name>
            <surname>Spear</surname>
          </string-name>
          .
          <year>2015</year>
          . “
          <article-title>Building Ontologies with Basic Formal Ontology</article-title>
          .” The MIT Press. https://dl.acm.org/citation.cfm?id=
          <fpage>2846229</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>