<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Achieving data FAIRification in a distributed analytics research platform for rare diseases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anna Bernasconi</string-name>
          <email>anna.bernasconi@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cinzia Cappiello</string-name>
          <email>cinzia.cappiello@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Ceri</string-name>
          <email>stefano.ceri@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pietro Pinoli</string-name>
          <email>pietro.pinoli@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electronics, Information and Bioengineering - Politecnico di Milano</institution>
          ,
          <addr-line>Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Data-driven medicine is fundamental to improving the accessibility and quality of the healthcare system. The availability of data is crucial for this purpose. In the context of a distributed analytics platform for analyzing healthcare data - employing the Personal Health Train paradigm - we propose to implement a solid data FAIRification infrastructure. This will allow us to achieve findability, accessibility, interoperability, and reusability of data, metadata, and results within a network of several medical centers participating in the BETTER Horizon Europe project, where the study of rare diseases (such as intellectual disability and inherited retinal dystrophies) will be targeted. Impacts will be visible to a large population of healthcare practitioners, prospectively influencing health policymakers.</p>
      </abstract>
      <kwd-group>
        <kwd>FAIR principles</kwd>
        <kwd>healthcare</kwd>
        <kwd>distributed analytics</kwd>
        <kwd>rare diseases</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Analytics paradigm called the Personal Health Train (PHT) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. PHT can be explained via
a railway system analogy that includes trains, stations, and train depots. The trains use the
network to visit diferent stations to transport several goods, which in this analogy correspond
to analytical tasks. By adapting this concept to BETTER, the analytical task is brought to the
data provider (i.e., a medical center), whereas the data instances remain in their original location
(called station).
      </p>
      <p>
        As a technical partner of the project, Politecnico di Milano (i.e., the authors) will particularly
focus on BETTER’s objective to guide medical centers in collecting patients’ data following
a common schema in order to promote interoperability and re-use of datasets in scope. This
includes legal/ethical data protection authorizations as well as data FAIRification (data
documentation, cataloging, and mapping to well-established ontologies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). We will design a unified
schema repository for medical centers’ (meta)data integration, keeping a high abstraction level
to encourage maximum interoperability (see [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]).
      </p>
      <p>Importantly, the project will aim at the integration of external sources such as European
Health Data Space (EHDS), the 1+Million Genomes initiative (1+MG), and the European Open
Science Cloud. Legal and ethical implications will be duly considered, and data access and
re-use procedures will be proposed. Data pseudonymization will be performed as a default
preprocessing step, mitigating the risk of personal data leaks. A real-world large-scale data
integration framework (based on well-established ontologies) will be demonstrated taking into
account heterogeneous datasets. The platform will be tested primarily on two rare disease use
cases, inherent to pediatric intellectual disability and inherited retinal dystrophies.</p>
      <p>In conclusion, the BETTER project relies on “bringing computation to data” via incremental
and federated learning, which avoids unnecessary data moving across medical centers while
exploiting much of the information encoded in such data. The project will enable EU medical
centers and beyond to make full use of the potential ofered by a safe and secure exchange, use,
and reuse of health data fostered by robust data FAIRification. In the context of intellectual
disability and inherited retinal dystrophies – with the potential of expanding the same paradigm
to other diseases – the generated analytical tools will help healthcare professionals become
more proficient in cutting-edge digital technologies, data-driven decision support, health risk
surveillance, and control activities, monitoring and management of healthcare quality levels.
Acknowledgements. The work is supported by BETTER, Grant agreement 101136262.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Welten</surname>
          </string-name>
          , et al.,
          <article-title>DAMS: A distributed analytics metadata schema</article-title>
          ,
          <source>Data Intelligence</source>
          <volume>3</volume>
          (
          <year>2021</year>
          )
          <fpage>528</fpage>
          -
          <lpage>547</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O.</given-names>
            <surname>Beyan</surname>
          </string-name>
          , et al.,
          <article-title>Distributed analytics on sensitive medical data: the personal health train</article-title>
          ,
          <source>Data Intelligence</source>
          <volume>2</volume>
          (
          <year>2020</year>
          )
          <fpage>96</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <article-title>Ontology-driven metadata enrichment for genomic datasets</article-title>
          ,
          <source>in: SWAT4HCLS</source>
          <year>2018</year>
          , volume
          <volume>2275</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <article-title>Conceptual modeling for genomics: building an integrated repository of open data</article-title>
          ,
          <source>in: ER 2017</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>325</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          , et al.,
          <article-title>A review on viral data sources and search systems for perspective mitigation of COVID-19, Briefings in Bioinformatics 22 (</article-title>
          <year>2021</year>
          )
          <fpage>664</fpage>
          -
          <lpage>675</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>