<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>FAIR Data Cube, a FAIR data infrastructure for integrated multi-omics data analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xiaofeng Liao</string-name>
          <email>XiaoFeng.Liao@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuliia Orlova</string-name>
          <email>yuliia.orlova@surf.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cenna Doornbos</string-name>
          <email>Cenna.Doornbos@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Niehues</string-name>
          <email>annaniehues@eatris.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Casper de Visser</string-name>
          <email>Casper.deVisser@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junda Huang</string-name>
          <email>Junda.Huang@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas H.A. Ederveen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Purva Kulkarni</string-name>
          <email>Purva.Kulkarni@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K. Joeri van der Velde</string-name>
          <email>joeriv@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Morris A. Swertz</string-name>
          <email>m.a.swertz@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Brandt</string-name>
          <email>martin.brandt@surf.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alain J. van Gool</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter A.C. 't Ho</string-name>
          <email>Peter-Bram.tHoen@radboudumc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CEUR Workshop Proceedings</institution>
          ,
          <addr-line>CEUR-WS.org</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Human Genetics, Radboud university medical center</institution>
          ,
          <addr-line>Nijmegen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Genomics Coordination Center, University of Groningen and University Medical Center Groningen</institution>
          ,
          <addr-line>Groningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Medical BioSciences department, Radboud university medical center</institution>
          ,
          <addr-line>Nijmegen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>SURF</institution>
          ,
          <addr-line>Science Park 140, Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Translational Metabolic Laboratory, Department of Laboratory Medicine, Radboud university medical center</institution>
          ,
          <addr-line>Nijmegen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We are witnessing an enormous growth in the amount of molecular profiling (omics) data enabeling the integration of multi-omics data. Nonetheless, this is challenging due to the lack of FAIR -omics data and metadata. The storage of human -omics data in secure silos, for privacy reasons, further complicates their reuse. Federated analysis of FAIR data is a privacy-preserving solution to make optimal use of these multi-omics data and transform them into actionable knowledge. The Netherlands X-omics Initiative is a National Roadmap Large-Scale Research Infrastructure aiming for eficient integration of data generated within X-omics and external datasets. To facilitate this, we developed the FDCube, which adopts and applies the FAIR principles and helps researchers to create FAIR data and metadata, facilitate reuse of their data, perform federated analysis, and make their data analysis workflows transparent.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;FAIR</kwd>
        <kwd>Multi-omics</kwd>
        <kwd>FAIR Data Cube</kwd>
        <kwd>Metadata</kwd>
        <kwd>Federated Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>It is now widely acknowledged that in order to truly advance our understanding of health, it is
required to combine -omics data from diferent sources. Nonetheless, this remains challenging
as data and their associated metadata are not always findable, accessible, interoperable, and
reusable (FAIR). Furthermore, as most -omics data are derived from a human source, these data
are mainly stored in secure and protected data silos. It remains a challenge to re-use these
highly secured data sets without the risk of infringing the privacy of the individuals from which
the data are derived. Hence, there is a need for tools that enable federated data analysis.</p>
      <p>Here we present the X-omics FAIR Data Cube (FDCube). The FDCube helps to make -omics
data comply with the FAIR principles and provides a federated data analysis mechanism to
bring algorithms to data stations, in order to facilitate data reusing and analysis, while ensuring
data privacy.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Result</title>
      <p>The architecture of FDCube is presented in Figure 1A. The FDCube infrastructure allows
dataset owners to register data on the FAIR Data Point (FDP) and it incorporates the FAIR Data
Station, a metadata capture platform that facilitates making data FAIR at the source. Using
the Investigation-Study-Assay (ISA) metadata schema, metadata is transformed into a FAIR
machine-actionable resource stored in an RDF triplestore.</p>
      <p>Researchers can exploit the FDCube to find datasets and initiate computation requests to
dataset owners. These federated analysis requests are executed on the respective datasets
through Vantage6, and the results are communicated back.</p>
      <p>We adopted the Trusted World of Corona (TWOC)1 project as a demonstration to show how
to utilize the FDCube for integrated multi-omics federated analysis. TWOC is developing an
information platform containing scientific data &amp; information and real world clinical
observations on Corona. Figure 1B illustrates the use of FDCube on the TWOC dataset by showing the
pipelines covering multiple functionalities.</p>
      <p>The FDCube is now listed as a catalog item in the SURF Research Cloud, in which an in-a-box
solution is provided to deploy the collection of software applications used by FDCube, including
the FAIR Data Point, GraphDB, and FAIR Data Station.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>