<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bridging Clinical and Genomic Knowledge: An Extension of the SPHN RDF Schema for Seamless Integration and FAIRification of Omics Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eelke van der Horst</string-name>
          <email>eelke@thehyve.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deepak Unni</string-name>
          <email>deepak.unni@sib.swiss</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Femke C. Kopmels</string-name>
          <email>femke@thehyve.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Armida</string-name>
          <email>jan.armida@sib.swiss</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasundra Touré</string-name>
          <email>vasundra.toure@sib.swiss</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wouter Franke</string-name>
          <email>wouter@thehyve.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katrin Crameri</string-name>
          <email>katrin.crameri@sib.swiss</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisa Cirillo</string-name>
          <email>elisa@thehyve.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sabine Österle</string-name>
          <email>sabine.oesterle@sib.swiss</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Personalized Health Informatics Group, SIB Swiss Institute of Bioinformatics</institution>
          ,
          <addr-line>Basel</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Hyve B.V.</institution>
          ,
          <addr-line>Utrecht</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Swiss Personalized Health Network (SPHN) is a Swiss research infrastructure initiative that aims to facilitate the exchange of health-related data in a FAIR manner. The SPHN Dataset and SPHN RDF Schema form an essential part of the SPHN Semantic Interoperability Framework, which currently covers mostly clinical routine data. To facilitate the integration of omics data produced by the SPHN National Data Streams, a genomics extension was developed. This was done in close collaboration with clinicians, researchers, bioinformaticians, and data managers, from Swiss university hospitals, academic research groups and the omics platforms. Here, we present the genomics extension of the SPHN RDF Schema, which can be used to semantically describe genomics experiments and covers both clinical and research domains. The schema centers around the general omics process flow, with concepts that denote the individual steps, such as sample processing, assay, and data processing. Genomics-specific specializations are provided, such as library preparation, sequencing assay, and sequencing analysis. The schema also facilitates in capturing other important omics metadata, such as information about the sequencing instrument, standard operating procedure, and quality control metrics. The extension aligns with existing semantic data models and reuses common biomedical vocabularies, such as EDAM, OBI and FAIR genomes, as value sets, thereby facilitating semantic interoperability. It will be used to FAIRify data that is produced within the Swiss network and to facilitate sharing this data as one knowledge graph for reuse among its participants.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Genomics</kwd>
        <kwd>FAIR</kwd>
        <kwd>SPHN</kwd>
        <kwd>RDF</kwd>
        <kwd>NGS</kwd>
        <kwd>clinical</kwd>
        <kwd>data model</kwd>
        <kwd>Semantic web1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The Swiss Personalized Health Network (SPHN) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a research infrastructure initiative focused
on establishing an enabling framework to support the sharing of health-related data in
accordance with the FAIR principles [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. At the heart of this initiative is the SPHN Semantic
Interoperability Framework, a comprehensive system that provides semantic artifacts for the
representation, validation, and statistical analysis of health data. Additionally, a tool stack is in
place to support researchers and hospitals in effectively utilizing this technology. Built on the
Semantic Web stack, the framework leverages Resource Description Framework (RDF), Web
Ontology Language (OWL), Shape Constraints Language (SHACL), and SPARQL Protocol and RDF
Query Language (SPARQL) for its formalization. This foundation ensures a seamless and
standardized approach to handling health-related information.
0000-0002-8777-5612 (E. van der Horst); 0000-0002-3583-7340 (D. Unni)); 0000-0003-4194-0771 (F. C.
Kopmels); 0000-0003-4639-4431 (V. Touré); 0000-0001-5058-3767 (W. Franke); 0000-0003-3656-3457 (K.
Crameri); 0000-0002-0241-7833 (E. Cirillo); 0000-0003-3248-7899 (S. Österle)
© 2024 Copyright for this paper by its authors.
      </p>
      <p>Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR Workshop Proceedings (CEUR-WS.org)</p>
      <p>
        The 2023.2 release of the SPHN RDF Schema facilitates the semantic representation of clinical
data such as diagnoses, routine clinical measurements, and standard lab tests [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and already
provided partial coverage for clinical omics data tailored to a restricted set of specific use-cases.
For instance, genomic information is addressed through overarching concepts designed for
representing basic genomic variations, including single nucleotide polymorphisms (SNPs).
Additionally, it incorporates simple concepts for representing a chromosome, genes, transcripts,
and proteins. However, these existing concepts only cover a fraction of the vast genomic, and
more broadly the omics landscape. To comprehensively address the diverse needs of SPHN
National Data Streams (NDS) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], particularly in domains such as oncology, pediatric care, and
infectious diseases, there is a necessity to expand the SPHN RDF Schema. This expansion will be
a key focus in the upcoming 2024.1 release (which will be released in Jan 2024), ensuring that
the SPHN RDF Schema aligns with the evolving demands in this critical area.
      </p>
      <p>Here, we present a genomics concept model that covers all aspects of the (clinical) NGS
workflow and metadata. This model builds upon the existing SPHN RDF Schema and follows a
generalized design to allow for extension to other omics fields. It aligns with existing ontologies
and vocabularies, and mirrors the design of common domain data models where applicable.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Data models for capturing (gen)omics data, for instance those of nucleotide repositories such as
European Nucleotide Archive (ENA) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], European Genome-Phenome Archive (EGA) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and
Genomic Data Commons [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], are already well known, but have their own drawbacks. They are
either simple catch-all data models with little semantics, too application-centric, or more focused
on study and attribution metadata. Ontologies such as the Ontology for Biomedical investigations
(OBI) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Semanticscience Integrated Ontology (SIO) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] offer high expressivity and
ontological rigor. However, they leave room for different ways of representing information and
require a solid ontology engineering background. FAIR Genomes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] broadly fits our use case,
but it is too specific for genomics, tailored more towards the data capture side than the optimal
data structure, and designed to fit a broad range of data capture applications thereby staying
generic. Other approaches are contrasted in section 5. Discussion.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>First, a workshop was organized to capture the needs and requirements of stakeholders and map
the (gen)omics data domain. Participants from 12 institutions including clinicians, researchers,
bioinformaticians, data managers, and leadership, from Swiss university hospitals, technology
institutes, and the genome center, collaboratively identified the most relevant concepts, relations,
and attributes, as well as the process flow for omics experiments with a focus on genomics. This
served as input for follow-up interviews with each NDS, where stakeholders clarified their use
cases and domain, and presented example data for each use case. All acquired information about
the data domain and (gen)omics workflows was compiled into a ‘statements document’, i.e. a
document that lists what was assumed to be true about the domain as simple statements,
organized per topic. Stakeholders iteratively provided refinements until consensus was reached.</p>
      <p>A review was performed of common (gen)omics data models from nucleotide repositories and
platforms, as well as biomedical ontologies and vocabularies, to evaluate whether these could be
(partly) reused for intended use cases, requirements, and data.</p>
      <p>The final model was validated at a second workshop where participants scrutinized it by
applying it to example data of their use cases.</p>
      <p>
        The concept set was formalized using the SPHN Dataset template [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]; SPHN Schema Forge
[12] was used to validate the resulting dataset, as well as to generate the corresponding RDF
Schema, SHACL validation rules, and SPARQL queries [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>From literature research and the initial workflow that was created in the first workshop, it
became clear that omics research could be modeled as a series of consecutive steps, where the
output of one step may serve as the input of another. For example, a sample or data may be input
for sample or data processing steps. In Figure 1, the order of individual steps is indicated by
relating steps that directly precede each other with a ‘predecessor’ relation, forming a chain or
sequence of steps. The material or data that is produced and subsequently produced between
preceding steps is indicated using an output or input relation, respectively, but is optional, since
this information is not always available or relevant (it is implied). This pattern forms the
backbone of the model.</p>
      <p>Each step in the omics workflow is a process concept that is composed of essential metadata
concepts about that process. The three top-level concepts for representing the omics workflow
are ‘Sample Processing’, ‘Assay’, and ‘Data Processing’. Here, the name ‘Assay’ was favored over
‘Experiment’ since the latter has ambiguous interpretation, and is used differently in several data
models in the same domain. The ‘Sample Processing’ concept is composed of zero or more input
and/or output ‘Sample’ concepts, while the ‘Data Processing’ concept is composed of zero or more
input and/or output ‘Data File’ concepts. The ‘Assay’ concept is composed of zero or more input
‘Sample’ concepts and zero or more output ‘Data File’ concepts. Note that both top level concepts
‘Sample Processing’ and ‘Data Processing’ and their descendants may be repeated to express a
sequence of processing steps. The documents describing these concepts are available at the SPHN
Interoperability Framework Gitlab repository [13].</p>
      <p>For the genomics field, special concepts are introduced that derive from these generic
concepts: ‘Library Preparation’ derives from ‘Sample Processing’, ‘Sequencing Assay’ from
‘Assay’, and ‘Sequencing Analysis’ from ‘Data Processing’. Derived concepts inherit every
‘composedOf’ from their more generic counterpart, and may be consecutively repeated in a
similar manner. This portion of the SPHN schema is illustrated in Figure 2 and described in the
paragraphs below.</p>
      <p>Central to the genomics workflow is the sequencing assay. The ‘Sequencing Assay’ concept is
composed of essential metadata concepts, representing the sequencer (‘Sequencing Instrument’),
library preparation (‘Library Preparation’), intended read length and depth, and zero or more
runs (‘Sequencing Run’). The ‘Sequencing Run’ concept represents the actual execution of the
assay, and holds information that may vary per run, such as read count, average insert size,
average read length, and optional quality control metrics (represented via the ‘Quality Control
Metric’ concept).</p>
      <p>The ‘Library Preparation’ concept is a special type of ‘Sample Processing’ that is part of a
‘Sequencing Assay’. It holds information on the library preparation kit, target enrichment kit, and
intended insert size, and, in case a gene panel kit is used as target enrichment, information on the
gene panel’s focus genes. Any other processing steps that precede an assay’s library preparation
may be registered using the ‘Sample Processing’ concept.</p>
      <p>Following the sequencing assay, there are one or more data processing steps that manipulate
the output of the assay. These are represented by the ‘Sequencing Analysis’ concept, a
specialization of ‘Data Processing’, that is composed of an optional reference genome.</p>
      <p>The model further provides utility concepts with general applicability. For instance, ‘Standard
Operating Procedure’ concept was created to provide information about the prescribed
step-bystep procedure that was followed to conduct experimental procedures such as sample processing,
assays, and data processing. In addition, the ‘Quality Control Metric’ concept holds information
about the value of certain quality control metrics that are relevant for these processes, such as
the ‘Phred quality score’ for DNA sequencing.</p>
      <p>The ‘Isolate’ concept is a specialization of the ‘Sample’, with a property to indicate the isolated
organism, which is relevant for pathogen surveillance research.</p>
      <p>Where applicable, concepts are aligned to terms and classes from common public domain
terminologies, such as SNOMED CT and OBI, thereby facilitating semantic interoperability. For
instance, the ‘Assay’ concept has a meaning binding to OBI’s ‘assay’ class (OBI:0000070), while
‘Isolate’ has a meaning binding to SNOMED CT’s ‘Microbial isolate specimen (specimen)’ concept
(SNOMED:119303007). In addition, selected subsets or branches from common public
terminologies are imposed as value sets for most nominal attributes. For instance, the type of
sequence analysis is indicated by descendants of EDAM’s [14] ‘Analysis’ operation, or similar.</p>
      <p>The diagram in Figure 3 visualizes an example instance of a sequencing assay and related
metadata. Listing 1 gives an RDF representation of the same example; the full RDF in Turtle syntax
is available at [15]. Note that, as with other procedure concepts, many of the composite concepts
are optional and may be omitted in case this data is not known or not relevant. For instance, a
‘Standard Operating Procedure’ may not be known or shared, or may be trivial (i.e. the operating
instructions by the vendor of a platform). Note that each run produces its own data file(s), which
may be selected or discarded as input for data analysis, based on quality metrics of the run.</p>
      <p>The genomics extension, including corresponding SPHN RDF Schema, SHACL shapes, and
SPARQL queries, is distributed as part of the 2024.1 release, and is available for download at the
SPHN Interoperability Framework Gitlab repository [13].</p>
      <p>Listing 1: Example instantiation in RDF of a whole genome sequencing assay and related
metadata.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>A genomics semantic data model was developed that builds on the SPHN schema and follows a
design that makes it applicable to other omics domains. The model covers clinical use cases from
the SPHN NDSs including omics data from oncology, pediatric care, and pathogen surveillance
and others at Swiss University Hospitals and research institutions. It aligns with common
biomedical terminologies such as SNOMED CT, OBI, and GENEPIO, and mirrors the structure of
existing domain models such as FAIR Genomes where possible. For instance, the pattern of
expressing the investigative workflow as a sequence of processes, where material or data
produced by one process serves as input for the next process, is similar to that of OBI and SIO.
Also, the intended use of the genomics concepts mirrors that of FAIR Genomes, where FAIR
Genomes’ ‘Sample Preparation’ module broadly corresponds to the ‘Library Preparation’ concept,
the ‘Sequencing’ module with the ‘Sequencing Assay’ concept, and the ‘Analysis’ module with the
‘Sequencing Analysis’ concept, respectively. The SPHN omics extension also reuses the value set
for library kits from FAIR Genomes. In contrast to FAIR Genomes, we aimed for more
normalization, for instance by introducing a separate ‘Standard Operating Procedure’ concept
and reference it, while this information is part of the ‘Material’ and ‘Analysis’ modules as protocol
attributes in the FAIR Genomes model. The ‘Run’ concept also served to factor out run-specific
information, and keep all information that is the same for every run in a sequencing experiment
within the ‘Sequencing Assay’ concept. In addition, we aim to use concept references over text or
string attributes, such as with the ‘Standard Operating Procedure’, ‘Software’ (algorithm), or
‘Quality Control Metric’ concepts, all of which are expressed as text attributes in FAIR Genomes.
Lastly, the SPHN omics extension for the SPHN RDF Schema offers several utility concepts, such
as ‘Isolate’ and ‘Gene Panel’, that may be seamlessly combined with other concepts to express the
clinical case metadata, has extension points to add additional concepts for any step in the omics
workflow, and allows material or data processing concepts to be chained which allows for more
fine grained expression of the experimental processing. Note that these differences are not
shortcomings, but reflect the differences in application: where FAIR Genomes provides a generic
content model for data capture applications for clinical NGS, the SPHN (gen)omics extension
mainly focuses on semantic data exchange with a common framework to fit a broader range of
omics fields. Data expressed using the SPHN RDF Schema may easily be transformed to the FAIR
Genomes model, whereas the inverse is harder.</p>
      <p>The model allows to describe metadata on the process, and, in combination with other SPHN
concepts, the outcome of the genomics workflow. Bulk transcriptomics and, to a lesser extent,
omics research in general can be represented with the model that is presented in this paper. Since
the model is purely meant for exchange of experimental metadata and data, catalog-level
metadata, such as study, funding, or attribution, was left out. Although there are some general
purpose high-level concepts such as ‘Sample Processing’, ‘Assay’, and ‘Data Processing’, we
refrained from introducing a complete concept hierarchy; only concepts that have direct
applicability were introduced. Also, while it would be useful to allow for partonomy, for instance
by allowing sample or data processing concepts to be composed of arbitrary part processes, it
was deliberately not introduced in order to restrict to only one way to apply the concepts.</p>
      <p>The 2023.2 SPHN RDF Schema release incorporated various external terminologies to
enhance data description. For genomics, the following terminologies are provided on the DCC
Terminology Service [16] Genotype Ontology (GENO), HUGO Gene Nomenclature Committee
(HGNC), and Sequence Ontology (SO) for comprehensive variant representation and human gene
naming. In the 2024.1 release, as genomics concepts expanded, additional terminologies such as
EMBRACE Data and Methods (EDAM) ontology, Experimental Factor Ontology (EFO), Genomic
Epidemiology Ontology (GENEPIO), and Ontology for Biomedical Investigations (OBI) will be
integrated.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>We developed a genomics extension of the SPHN RDF Schema, in close collaboration with
stakeholders of the SPHN, to facilitate their data sharing use cases. This omics extension aligns
with common biomedical ontologies and offers extension points for other omics research. This
set of concepts forms the basis for the further concepts developed in the NDSs for other omics
data, driving the vision of holistic view on the personalized health data in one single knowledge
graph.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We wish to thank Aitana Neves, Alexander Leichtle, Andre Kahles, Andrea Agostini, Assaf
Sternberg, Cédric Howald, Cyril Matthey-Doret, Daniel Damian, Daphné Chopard, Dylan Lawless,
Gaëtan De Frapoint, Irene Keller, Jossena Lype, Katrin Männik, Leonardus Bosch, Lorenzo Cerutti,
Manuel Schweighofer, Marc Zimmerli, Michael P. Schmid, Natalia Chicherova, Nicolas Edouard
Martin Freundler, Nicolas Freundler, Nora Toussaint, Oksana Riba Grognuz, Patrick Buhlman,
Patrick Hirschi, Patrick Pedrioli, Peter Fritsch, Phil Cheng, Reinhard A. Dietrich, Ruben Casanova,
Stefan Neuenschwander, Stefan Nicolet, Sylvain Pradervand, Tess Brody, Thomas Müller-Focke,
Tural Yarahmadov, Vito Zanotelli and Walid Gharib for their participation to the Omics data
modeling workshops and their contributions to this extension of the SPHN RDF Schema. This
work was funded by the Swiss State Secretariat for Education, Research and Innovation (SERI)
through the SPHN initiative.
[12] SPHN Schema Forge - SIB Swiss Institute of Bioinformatics, 2023, URL
https://schemaforge.dcc.sib.swiss/. Accessed: 2023-11-27.
[13] SPHN Interoperability Framework Release-candidate-2024-1, 2023. URL:
https://git.dcc.sib.swiss/sphn-semantic-framework/sphn-schema/-/tree/releasecandidate-2024-1.
[14] J. Ison, M. Kalas, I. Jonassen, D. Bolser, M. Uludag, H. McWilliam, J. Malone, R. Lopez, S. Pettifer,
P. Rice, EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics
and formats, Bioinformatics 29, 10 (2013). doi:10.1093/bioinformatics/btt113.
[15] Example instantiation of genomic concepts from SPHN RDF Schema 2024.1. URL:
https://gist.github.com/deepakunni3/1a1324fcddb82fb0c1064f8025be5852.
[16] P. Krauss, V. Touré, K. Gnodtke, K. Crameri, S. Österle, DCC Terminology Service—An
Automated CI/CD Pipeline for Converting Clinical and Biomedical Terminologies in Graph
Format for the Swiss Personalized Health Network, Applied Sciences 11.23 (2021) 11311.
doi:10.3390/app112311311.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Touré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kraus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gnodtke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Buchhorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Unni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Horki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Raisaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kalt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Teixeira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Crameri</surname>
          </string-name>
          , S. Österle,
          <article-title>FAIRification of health-related data using semantic web technologies in the Swiss Personalized Health Network</article-title>
          .
          <source>Scientific Data 10.1</source>
          (
          <year>2023</year>
          )
          <article-title>127</article-title>
          . doi:
          <volume>10</volume>
          .1038/s41597-023-02028-y.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>M. D.</surname>
          </string-name>
          Wilkinson et al.,
          <article-title>The FAIR Guiding Principles for scientific data management and stewardship</article-title>
          .
          <source>Scientific Data 3.1</source>
          (
          <year>2016</year>
          )
          <article-title>160018</article-title>
          . doi:
          <volume>10</volume>
          .1038/sdata.
          <year>2016</year>
          .
          <volume>18</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>SPHN</given-names>
            <surname>Interoperability</surname>
          </string-name>
          <article-title>Framework version</article-title>
          <year>2023</year>
          -2 release,
          <year>2023</year>
          . URL: https://git.dcc.sib.swiss/sphn-semantic-framework/sphn-schema/-/releases/2023-2.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>SPHN</given-names>
            <surname>National Data</surname>
          </string-name>
          <article-title>Streams</article-title>
          . URL: https://sphn.ch/services/funding_old/nds/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yuan</surname>
          </string-name>
          et al.,
          <source>The European Nucleotide Archive in 2023, Nucleic Acids Research</source>
          (
          <year>2023</year>
          )
          <article-title>gkad1067</article-title>
          . doi:
          <volume>10</volume>
          .1093/nar/gkad1067.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Freeberg</surname>
          </string-name>
          et al.,
          <source>The European Genome-phenome Archive in 2021, Nucleic Acids Research</source>
          <volume>50</volume>
          ,
          <issue>D1</issue>
          (
          <year>2022</year>
          )
          <fpage>D980</fpage>
          -
          <lpage>D987</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkab1059.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ferretti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Grossman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Staudt</surname>
          </string-name>
          ,
          <article-title>The NCI Genomic Data Commons as an engine for precision medicine</article-title>
          ,
          <source>Blood 130.4</source>
          (
          <year>2017</year>
          )
          <fpage>453</fpage>
          -
          <lpage>459</lpage>
          . doi:
          <volume>10</volume>
          .1182/blood-2017
          <source>-03- 735654.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bandrowski</surname>
          </string-name>
          et al.,
          <article-title>The Ontology for Biomedical Investigations</article-title>
          ,
          <source>PLOS ONE 11.4</source>
          (
          <issue>2016</issue>
          ),
          <year>e0154556</year>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0154556</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          et al.,
          <article-title>The Semanticscience Integrated Ontology (SIO) for biomedical research and knowledge discovery</article-title>
          ,
          <source>Journal of Biomedical Semantics 5.1</source>
          (
          <issue>2014</issue>
          ),
          <volume>14</volume>
          . doi:
          <volume>10</volume>
          .1186/2041-1480-5-14.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>K. J. van der Velde</surname>
          </string-name>
          et al.,
          <article-title>FAIR Genomes metadata schema promoting Next Generation Sequencing data reuse in Dutch healthcare and research</article-title>
          .
          <source>Scientific Data 9.1</source>
          (
          <year>2022</year>
          )
          <article-title>169</article-title>
          . doi:
          <volume>10</volume>
          .1038/s41597-022-01265-x.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>SPHN</given-names>
            <surname>Interoperability</surname>
          </string-name>
          <article-title>Framework dataset template</article-title>
          ,
          <year>2023</year>
          . URL: https://git.dcc.sib.swiss/sphn-semantic-framework/sphn-schema/- /tree/master/templates/dataset_template. Accessed:
          <fpage>2023</fpage>
          -11-28.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>