<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Registries of domain-relevant semantic reference models help bootstrap interoperability in domains with fragmented data resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Roos</string-name>
          <email>m.roos@lumc.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark D Wilkinson</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rajaram Kaliyaperumal</string-name>
          <email>r.kaliyaperumal@lumc.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Thompson</string-name>
          <email>m.thompson@lumc.nl</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Carta</string-name>
          <email>claudio.carta@iss.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ronald Cornet</string-name>
          <email>r.cornet@amc.uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David van Enckevort</string-name>
          <email>david.van.enckevort@umcg.nl</email>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luiz Bonino</string-name>
          <email>luiz.bonino@dtls.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academic Medical Center University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dutch Techcentre for Life Sciences</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Istituto Superiore di Sanita</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Leiden University Medical Centre</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Linkoping University</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Universidad Politcnica de Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University Medical Center Groningen</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The specialist eld of rare diseases must connect its vast array of globally distributed disease and patient registries to maximise their value. Unfortunately, many registries are \boutique", with few or no sta with formal informatics training. At a series of Bring Your Own Data workshops, we helped registry owners transform their data into formally structured triple stores following the Linked Data principles and demonstrated the potential of data linkage. We documented several useful approaches that we believe could be followed independently by other registry owners worldwide, including: that the transformation to Linked Data could be considered as passing through layers of increasing semantic complexity; that only a subset of ontologies are relevant at each layer; and that certain data transformation processes could be modelled as an \archetype", and presented to registry sta to ll-in with their data. We propose that formally capturing these ontological layers and archetypes, and registering them as a reference and teaching resource will facilitate the wider community of non-expert data owners self-direct their own data transformations.</p>
      </abstract>
      <kwd-group>
        <kwd>semantic reference model</kwd>
        <kwd>linked data</kwd>
        <kwd>ontologies</kwd>
        <kwd>rare disease</kwd>
        <kwd>archetype</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Making data linkable at the source has become a key ambition for the
development of robust infrastructure that supports data integration in the rare disease
community. There are over 6000 rare diseases, each with multiple data resources
across the globe, ranging from biobanks, patient or disease registries, and omics
data sources. We must accept the challenge of implementing solutions that can
scale-up to be adopted by thousands of resources, with the knowledge that
maintaining a centralised warehouse at this scale, and with this kind of sensitive data,
is neither feasible nor ethically or legally acceptable.</p>
      <p>Perhaps more than in other domains, progress in the rare disease eld
depends on combining data, given that the disease-speci c data is so sparse. It
is, however, well established that biomedical data integration is an extremely
error-prone process that requires a deep understanding of both biology and
data/knowledge management to reconcile data from di erent sources. To
improve the data ecosystem for this important target community, we are working
on a standard set of procedures and lightweight technologies that will make rare
disease data Findable, Accessible, Interoperable and Reusable for both humans
and computers (FAIR) at the source.</p>
      <p>
        In this position paper we discuss requirements and subsequent design
decisions that we have chosen to pursue during a still ongoing plan to make rare
disease biobanks and registries linkable at the source. The plan also includes
a study of the steps to make data FAIR at the molecular level (e.g. genetic
variants, metabolites, molecular pathways) in order to link to information in
registries and biobanks. The plan is guided by experiences gained from a
number of Bring Your Own Data workshops (BYODs) in the rare disease domain
[
        <xref ref-type="bibr" rid="ref13 ref14">14, 13</xref>
        ]. While not all components described here have been built or tested, we
take the position that our early successes in the early stages of this approach
suggest that the future extensions - currently in development and based on the
same layered design - will exhibit similar successes.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Backbone: Linkable Data and Ontologies</title>
      <p>
        Choosing Linked Data principles and Ontologies to make rare disease data
linkable at the source was our rst design decision, as RDF was designed with the
objective of creating quali ed networks of data, upon which increasingly
complex domain models can be overlaid to assist with interpretation of that data.
For instance, the Human Phenotype Ontology and the Orphanet Rare Disease
Ontology are obvious choices to denote human phenotypes and diseases in rare
disease resources within this Linked Data. We therefore considered this the best
way to facilitate integrative biological and translational research across rare
disease resources. Other tools in this general domain that use ontologies include the
exomizer, matchmaker exchange tools, and Monarch, providing examples of the
power of using phenotype annotations and cross-species phenotype mappings
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. RDF is capable of representing disease specimen identi ers, patient/disease
personal and clinical information, and molecular data, thus the choice of this
singular technological framework helps reduce the overall cost of downstream
data integration for rare disease resources. As such, we received wide support
for putting this rst design decision into practice from many sources, such as
RD-Connect, Elixir, BBMRI, ODEX4All, FAIRDict, academic hospitals, and an
increasing number of patient organisations.
2.1
      </p>
      <p>Composite semantic models as reference for preparing data for
integration
Our position is that, by making an explicit set of increasingly rich semantic
layers (diagrammed as a set of "Modules" in Figure 1), where each Module may
be taught and undertaken in-isolation from the others, focuses trainees on the
speci c subset of tasks and ontologies required to achieve success in that layer.
We will now elaborate on that position.</p>
      <p>Linked Data with strong ontological underpinnings, and a clear model for
achieving proper access control, was our rst ambition for preparing the
relatively small, but numerous and disparate, rare disease data sets for wide-scale
data integration. However, an immediate and major bottleneck was the sparsity
of expertise in the community to make informed decisions about which
ontological concepts to use for their data annotations. Searching for a concept, e.g. in
NCBOs bioportal or EBIs ontology lookup service, typically returns too many
hits for a non-ontologists to choose from. Speci c ontologies may be advised by
experts, but the breadth of data types across data sets is large. For example,
working with rare disease patient registry managers, we easily listed at least 10
ontologies relevant for even a small a subset of their registry s data, and not
all of these are included in the BioPortal or EBI search services. Providing our
target community with too many choices will be confusing. At the same time,
investigating individual ontologies for each rare disease resource that we prepare
for analysis across data sets is time consuming and ine cient.</p>
      <p>Ideally, therefore, we should attempt to record and reuse previous
ontologyassessments every time we go through the process of making a rare disease
resource linkable, such that we consistently advise only one or at most a small
number of ontologies for any given class or type of data/observation. The key
objective of Module 1, therefore, was to create a searchable subset of
domainrelevant ontological resources or ontology-slices, rather than asking our
community to do an open ontology search for every term. This provides a way for
non-experts to start to become good Linked Data publishers and reduces
confusion and frustration for our rare disease registry community. Moreover, because
the constraints are only on what we present to the data publisher, the power
of the full ontology remains available to machines that consume or query that
data. To date, we have successfully used these approaches to assist a number of
rare disease data registries - many of them with little or no formal training in
data or knowledge management, to create Linked Data from their data that is
of su ciently high quality that it can be used as a source for federated SPARQL
queries. We now wish to scale-up these e orts, such that registry owners require
ever-fewer formal contacts with Linked Data or ontological experts.</p>
      <p>
        The next layer in our stack, Module 2, guided again by our experience
working with this community, derives from our observation that many of the core
data models from patient registries and biobanks have near-identical structures,
particularly for similar 'types' of data (for example, clinical observations are
similar in structure to each other, but distinct from coded phenotypic
observations). We have therefore started to compose semantic reference models for
rare disease data integration (closely related to Archetypes in health
information systems, e.g. see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) that our community can simply copy and populate
with their speci c data. These, too, will be published in a searchable registry
of such models, and will be 'tagged' with keywords related to the kinds of data
they are capable of representing. Note that these models are data structures, not
novel semantic models of diseases or phenotypes - we do not intend these models
to be new conceptualisations of a domain - in the sense that, for instance, the
Human Phenotype Ontology is a distinct conceptualisation in the phenotypic
domain; rather, these are meant as artefacts to further our data integration
goals, which, together with the constrained ontological choices, suggest/limit
both structure and semantics, reducing freedom-of-choice, but enhancing
interoperability through capturing what we believe are the best-practises de ned by
data publishing experts. Archetypes are initially being designed through a
collaboration between rare disease domain experts and Linked Data experts until
a mutually-acceptable model is created. This model will then be published as
a reference for individual data owners to build Linked Data within their
domain/scope. Through our ongoing pursuit this approach, we intend to gradually
build-up a clearly-de ned set of starting points for all of the various data-types in
the rare disease domain, allowing us to rapidly scale-up to absorb new resources
into our integrated community through their own individual e orts.
      </p>
      <p>
        This stack is currently being extended further (Module 3), where we are
planning to create resources of archetypes with greater semantic complexity
that may be used to combine individual local observations, as was done in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ],
or link local observations with remote observations. As we create these reference
models, we propose to distinguish between the three distinct outcomes that
the models should support. These are, in order of complexity: the need to (i)
harmonize and simplify annotation of source data, (ii) query across resources and
enable statistical analysis of knowledge graphs, and (iii) enable logical reasoning
to facilitate discovery by revealing\unknown unknowns". We propose to follow
the modelling suggestions of the SemanticScience Integrated Ontology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that
semantic models de ned using OWL axioms allow a modular, layered approach,
resulting in a composite model that can address each of these needs. Thus each
of our proposed modules is an independent OWL le that can be utilised in
isolation, depending on the expertise of the publisher.
      </p>
      <p>
        Each module in the stack serves a speci c, and increasingly more complex
integrative purpose; the full spectrum of requirements - up to and including
semantic reasoning - will only be achieved when all of the modules have been used
to annotate/represent the data. For instance, Module 1 (green, in Figure 1) for
core identi er annotation primarily recommends ontological classes that allow
explicit typing of the identi ers commonly seen in rare disease databases, but
excludes any deeper properties such as those that would facilitate faceted data
integration or querying based on properties or their values. In our current model
we use the EMBRACE Data and Methods model (EDAM [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) as the source of
identi er-type semantics, and we encourage the use of the identi ers.org URI
schemes to harmonise the identi er structures, where appropriate. The
application of Module 1, therefore, facilitates simple queries to identify repositories
that contain data of a particular nature, but are insu cient for the more
complex integrative behaviours. Such richer behaviours are enabled by applying the
archetypes and ontologies recommended by Modules 2 and 3. What is
important to note is that the layers separate and stratify the tasks of semantic data
migration. Module 1 starts with the most core question \what data do I have",
which is in-itself an important piece of semantic information. The layers make
it clear that this basic task can and should be clearly separated from the other,
more complex tasks required to support full integrative queries.
      </p>
      <p>We argue that the task undertaken in Module 1 is su ciently comprehensible
and self-evident that it provides non-ontologists an relatively easy way to pursue
semantic transformations unaided by a data linking expert. We take the position
that the \shallow" semantic transformation undertaken by Module 1 not only
provide useful integrative behaviours, but do not in any way compromise the
later addition of greater semantic expressivity, guided by Modules 2 and 3. We
further hold the position that both Module 1 and Module 2, when presented to
the community as a limited set of choices, provide a level of expectation
wellwithin the capabilities of our target, non-expert data publising community. While
the complexity of ontologies is a bottleneck for many who approach semantic
transformations on their own, we propose that this layered approach lowers
the bar for participation, and will stimulate more registries to undertake these
preliminary transformations \at-source", on their own initiative.</p>
      <p>
        As mentioned, Modules 2 and 3 are still under-development. We believe
that these Modules will contain recommendations that can be reused to
stimulate interoperability between resources. They provide predicates that de ne
the relations between individuals and their observed or measured clinical
features/phenotypes, and archetypes for how to assemble these observations
without loss of data. For example, using the Semanticscience Integrated Ontology
model (SIO; [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) for recording measurements ensures that all clinical
observations must include a value, a measurement unit, and an ontological type (for
example, systolic blood pressure). Tables and relational database models rarely
explicitly express such semantics, and thus by providing these simple, but
rigorous archetypes, we provide a clear path forward for those who wish to further
transform their data. Moreover, by agreeing on archetypes, we are able to
create registry software that uses these as data-capture templates, ensuring that
newly generated data lls these richer models without requiring extensive
training of the registry owners. Indeed, we are working with patient registry software
providers to incorporate support for this directly in their tools. For existing data,
we can apply tools such as OpenRe ne with the RDF plugin to add URIs for
data values (e.g. HPO URIs for phenotypes) and data types, their interrelations,
and their links to the reference model.
2.2
      </p>
      <p>
        A prototype semantic reference model for enabling questions
across rare disease resources
We have created a rst version of a semantic reference model in the rare disease
domain. Its main purpose is to enable questions across rare disease biobanks
and registries. This is re ected by separate modules that comprise our
reference model ( gure 1), an example is given in gure 2. Each module is available
as a separate owl le (8). The model refers to EDAM [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for identi ers, OBIB
(Ontology for Biobanking [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) for biological specimen, ORE (the Object Reuse
and Exchange model [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) for aggregating research materials (a decision inspired
by the research object model [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]), and assumed the use of HPO (the Human
Phenotype Ontology [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) for phenotype identi ers and ORDO (Orphanet Rare
Disease Ontology [
        <xref ref-type="bibr" rid="ref17 ref8">17, 8</xref>
        ]) for rare disease identi ers. At this time, we make no
further assumptions as to which ontologies to recommend in the rare disease
domain, but this is anticipated with support from RD-Connect. We included some
initial mappings to for instance SNOMEDCT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. With our collaborators, we
have also started work on mappings to MIABIS (Minimum Information About
      </p>
      <sec id="sec-2-1">
        <title>8 https://github.com/LUMC-BioSemantics/Rare-Disease-Semantic-Model</title>
        <p>
          BIobank data Sharing [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]). In a next revision, we also aim to address the issue
of implicit reinterpretation of rare disease data by changes in the ontologies that
we use in the reference model. Assuming that ontology versioning is still
imperfect, this may require an additional layer in our stacked approach. The current
model is stored in github, and we have created an entry in BioSharing9. We have
not (yet) decided on uploading the model to NCBOs BioPortal.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Registries for domain-speci c semantic reference models</title>
      <p>Composite semantic reference models provide a useful service for data
integration. In our experience however, these models could be better supported.
Considering the FAIR paradigm, they are currently hard to make ndable, accessible
and reusable. Model registries and lookup services could make them more
ndable as data integration artefacts, and advocate them as standard schemas for
data annotation in speci c domains (in our case the rare disease domain). We
envision searching for models for data integration by linked data graphs, such as
a search for models that prepare data for linking genes to diseases (`Which model
can make genes in my data set linkable to diseases in other people's data?' ). We
have uploaded our alpha version to BioSharing9, which may be the
appropriate platform. It aims to be a central point to nd standards, and it will help
reuse because we can add example annotated data. We envision that disease and
sample registry software will use semantic archetype registries to optimize data
entry towards generating interoperable data. However, at this time it has no
special features for searching semantic models that prepare a resource for data
integration. Tools such as EBIs Zooma and Ontology Lookup Service help users
nd a wealth of possibilities with great precision, but do not yet allow ltering</p>
      <sec id="sec-3-1">
        <title>9 https://biosharing.org/bsg-s000676</title>
        <p>on semantic archetypes to limit the search results. For example Zooma returns
a list of concepts for the search terms and shows which other sources use that
particular ontological term, but not the archetype of the source itself.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>
        The concept of building composite models from existing ontologies for speci c
applications is not new, and they often help data integration. For example, EBIs
Experimental Factor Ontology was developed as an application ontology for
linking EBI [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] resources, and the Just Enough Results Model [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] may be
considered a reference model for linking systems biology data within the FAIRdom
initiative10. The need to simplify the ontological landscape for applied
ontologists also seems a strong motivation for developing SIO [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Our suggestion is
to mitigate this need by reusing the work of ontologists through published data
integration models that apply state of the art axiomatized ontologies.
      </p>
      <p>We take the position that it is useful to create a registry that supports
nding, accessing and reusing domain-relevant subsets of ontological classes and
Linked Data models. When designing this approach, we took the position that
the task can be cast into three independent Modules that address distinct levels
of semantic complexity; we hope that semantic model tool builders will now
investigate how well their tools support this. Finally, we take the position that,
when cast in this way, the tasks represented by Modules 1 and 2 become tractable
to non-experts in data publishing, due to the enhanced clarity and simpli ed,
task-speci c, search results.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We are creating a reusable semantic reference model to speed up the process of
making rare disease data resources FAIR and linkable at the source. This
pertains to the many patient/disease registries, biospecimen collections (biobanks),
and omics data resources that we need to be able to query across in order to
speed up rare disease research in healthcare and life science. We have observed in
previous BYODs that nding recommendable concepts in existing ontologies is
the main bottleneck for rare disease stakeholders, and a redundant time
investment for Linked Data experts. We take the position that we can mitigate this by
improving support speci cally for semantic models that are made to facilitate
data integration downstream of semantic data encoding and annotation. These
semantic models should be easy to nd, access, and reuse. For interoperability
use cases beyond ndability, we advocate a modular approach, providing
appropriate modules for data annotation, enabling simple manual queries, and big
data analytics and reasoning. We propose that BioSharing could be a target for
extending support for semantic data integration models, for instance by allowing
searches for linked data patterns.
10 http://fairdom.org
Acknowledgments. We thank all domain experts and linked data experts
who contributed to previous Bring Your Own Data workshops for rare disease
registries and biobanks. We thank the colleagues at ISS (Istituto Superiore di
Sanita, Rome, Italy), particularly Sabina Gainotti and Domenica Taruscio, and
Mascha Jansen (Dutch Techcentre for Life Sciences) for their preparatory work
in organising these workshops. We thank Andrew Gibson and Katy Wolstencroft
for fruitful discussions. The work leading to this paper is supported by grants
from RD-Connect (FP7/20072013, grant agreement No. 305,444), Elixir
infrastructure for life science data, Elixir-Excelerate (H2020-INFRADEV-1-2015-1),
and BBMRI-NL2 (NWO National Roadmap for Large-Scale Research
Facilities). MDW is supported by the Fundacion BBVA and the UPM Isaac Peral
programme, and the Spanish Ministerio de Economa y Competitividad grant
number TIN2014-55993-R.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roure</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gamble</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchan</surname>
          </string-name>
          , I.: Research Objects:
          <article-title>Towards Exchange and Reuse of Digital Knowledge (feb</article-title>
          <year>2010</year>
          ), http://eprints. ecs.soton.ac.uk/18555/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Belhajjame</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garijo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gamble</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hettne</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palma</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mina</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klyne</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Using a suite of ontologies for preserving work ow-centric research objects</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>32</volume>
          ,
          <issue>16</issue>
          {42 (may
          <year>2015</year>
          ), http: //linkinghub.elsevier.com/retrieve/pii/S1570826815000049
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brochhausen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birtwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masci</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ellis</surname>
            ,
            <given-names>H.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoeckert</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , Jr.:
          <article-title>OBIB-a novel ontology for biobanking</article-title>
          .
          <source>Journal of biomedical semantics 7</source>
          ,
          <issue>23</issue>
          (
          <year>2016</year>
          ), http://www.ncbi.nlm.nih.gov/pubmed/27148435http: //www.pubmedcentral.nih.gov/articlerender.fcgi?artid=PMC4855778
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dumontier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baran</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callahan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chepelev</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz-Toledo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Rio</surname>
            ,
            <given-names>N.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duck</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furlong</surname>
            ,
            <given-names>L.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keath</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klassen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Queralt-Rosinach</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villanueva-Rosales</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <article-title>Hoehndorf: The Semanticscience Integrated Ontology (SIO) for biomedical research and knowledge discovery</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <volume>14</volume>
          (
          <year>2014</year>
          ), http://jbiomedsem.biomedcentral.com/articles/10.1186/2041-1480-5-14
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ellouze</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouaziz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghorbel</surname>
          </string-name>
          , H.:
          <article-title>Integrating semantic dimension into openEHR archetypes for the management of cerebral palsy electronic medical records</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>63</volume>
          ,
          <issue>307</issue>
          {
          <fpage>324</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Haendel</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasilevsky</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brush</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hochheiser</surname>
            ,
            <given-names>H.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobsen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oellrich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , Washington, N., Kohler,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.E.</given-names>
            ,
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.N.</given-names>
            ,
            <surname>Smedley</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Disease insights through cross-species phenotype comparisons</article-title>
          .
          <source>Mammalian Genome</source>
          <volume>26</volume>
          (
          <issue>9-10</issue>
          ),
          <volume>548</volume>
          {555 (oct
          <year>2015</year>
          ), http://link.springer.
          <source>com/10. 1007/s00335-015-9577-8</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ison</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonassen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolser</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uludag</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McWilliam</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rice</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>EDAM: an ontology of bioinformatics operations, types of data and identi ers, topics and formats</article-title>
          .
          <source>Bioinformatics</source>
          (Oxford, England)
          <volume>29</volume>
          (
          <issue>10</issue>
          ),
          <volume>1325</volume>
          {32 (may
          <year>2013</year>
          ), http://www.ncbi.nlm.nih.gov/pubmed/23479348http://www.pubmedcentral. nih.gov/articlerender.fcgi?artid=PMC3654706
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kahn</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          :
          <article-title>Integrating ontologies of rare diseases and radiological diagnosis</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>22</volume>
          (
          <issue>6</issue>
          ) (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lagoze</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , Van de Sompel, H.,
          <string-name>
            <surname>Nelson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnston</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>A Web-based resource model for scholarship 2.0: object reuse &amp; exchange</article-title>
          .
          <source>Concurrency and Computation: Practice and Experience</source>
          <volume>24</volume>
          (
          <issue>18</issue>
          ),
          <volume>2221</volume>
          {2240 (dec
          <year>2012</year>
          ), http://doi.wiley.
          <source>com/10</source>
          .1002/cpe.1594
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holloway</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adamusiak</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kapushesky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolesnikov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhukova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brazma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parkinson</surname>
          </string-name>
          , H.:
          <article-title>Modeling sample variables with an experimental factor ontology</article-title>
          .
          <source>Bioinformatics</source>
          (Oxford, England)
          <volume>26</volume>
          (
          <issue>8</issue>
          ),
          <volume>1112</volume>
          { 1118 (apr
          <year>2010</year>
          ), http://www.ncbi.nlm.nih.gov/pubmed/20200009http://www. pubmedcentral.nih.gov/articlerender.fcgi?artid=PMC2853691
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Norlin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fransson</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eriksson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merino-Martinez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anderberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurtovic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litton</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          :
          <article-title>A Minimum Data Set for Sharing Biobank Samples, Information, and Data: MIABIS. Biopreservation and biobanking 10(4</article-title>
          ),
          <volume>343</volume>
          {8 (aug
          <year>2012</year>
          ), http://www.ncbi.nlm.nih.gov/pubmed/24849882
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          , Kohler,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Bauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Seelow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Horn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Mundlos</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>The Human Phenotype Ontology: a tool for annotating and analyzing human hereditary disease</article-title>
          .
          <source>American journal of human genetics 83(5)</source>
          ,
          <volume>610</volume>
          {5 (nov
          <year>2008</year>
          ), http://www.ncbi.nlm.nih.gov/pubmed/18950739http://www. pubmedcentral.nih.gov/articlerender.fcgi?artid=PMC2668030
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>A.J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waagmeester</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaliyaperumal</surname>
          </string-name>
          , R., van der Horst, E.,
          <string-name>
            <surname>Mons</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkinson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Bring Your Own Data Workshops: A Mechanism to Aid Data Owners to Comply with Linked Data Best Practices</article-title>
          . In: Paschke,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Burger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Romano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Marshall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Splendiani</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the 7th International Workshop on Semantic Web Applications and Tools for Life Sciences</source>
          , Berlin, Germany, December 9-
          <issue>11</issue>
          ,
          <year>2014</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>1320</volume>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1320</volume>
          / paper_36.pdf
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Roos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopes</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Bring your own data parties and beyond: make your data linkable to speed up rare disease research</article-title>
          . In: Vittozzi,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Salvatore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Taruscio</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.) Abstracts presented to the
          <source>EPIRARE International Workshop 24-25 November</source>
          <year>2014</year>
          . pp.
          <volume>21</volume>
          {
          <fpage>24</fpage>
          .
          <string-name>
            <surname>Rome</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://rarejournal.org/rarejournal/ article/viewFile/69/93
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Samadian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McManus</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkinson</surname>
          </string-name>
          , M.D.:
          <article-title>Extending and encoding existing biological terminologies and datasets for use in the reasoned semantic web</article-title>
          .
          <source>Journal of biomedical semantics 3(1)</source>
          ,
          <volume>6</volume>
          (jul
          <year>2012</year>
          ), http://www.ncbi.nlm.nih.gov/ pubmed/22818710
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Spackman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Snomed rt and snomedct. promise of an international clinical terminology. M.D. computing : computers in medical practice 17(6</article-title>
          ), 29, http: //www.ncbi.nlm.nih.gov/pubmed/11189756
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Vasant</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chanas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanauer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jupp</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parkinson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rath</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ordo: An ontology connecting rare disease, epidemiology and genetic data</article-title>
          .
          <source>In: Phenotype data at ISMB2014</source>
          (
          <year>2014</year>
          ), http://phenoday2014.bio-lark.org/pdf/9.pdf
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wolstencroft</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Owen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krebs</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stanford</surname>
            ,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Golebiewski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidemann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bittkowski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>An</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shockley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoep</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goble</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Seek: a systems biology data and model management platform</article-title>
          .
          <source>BMC Systems Biology</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <volume>33</volume>
          (dec
          <year>2015</year>
          ), http://www.biomedcentral.com/ 1752-0509/9/33
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>