<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>DOLAP</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>HealthMesh: An Architectural Framework for Federated Healthcare Data Management</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>AniolBisquert</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AchrafHmimou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josep Ll. Berra</string-name>
          <email>josep.ll.berral@upc.ed</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AlbertoGutierrez-Tor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>reand</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>OscarRomero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Architectural framework</institution>
          ,
          <addr-line>Healthcare, Data Mesh, Data Governance, Data Management, Federated data</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Barcelona Supercomputing Center</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Languages and Analytical Processing of Big Data</institution>
          ,
          <addr-line>co-located with</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universitat Politècnica de Catalunya</institution>
          ,
          <addr-line>UPC-BarcelonaTech</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>26</volume>
      <abstract>
        <p>Recently, significant milestones have been achieved in the field of healthcare data analysis. However, alongside these accomplishments, substantial data-related challenges have emerged in the domain of big data management. Modern healthcare projects are no more dealing with a single data repository but many heterogeneous ones and must overcome data variety, privacy and governance issues. Yet, current solutions face a privacy-decentralization trade-of. To address this dual challenge, we introduce HealthMesh, a novel layered architectural framework based on the Data Mesh principles, providing a domaindecentralised paradigm. In addition, the framework incorporates a Semantic Data Model which establishes robust governance, enables interoperability and guarantees policy compliance for all the data assets. To demonstrate the capabilities of the proposed approach, we provide an illustrative example inspired by the use case of the INCISIVE project for breast cancer analytics. Overall, this work makes a significant contribution on collecting key challenges, identifying actors and providing a set of components and guidelines for establishing a holistic framework for the complex field of healthcare data management.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
1. Introduction
velopment of the data management infrastructure. This
imbalance between the progress in data analytics and
In recent years, the healthcare landscape underwednattaa management is due to several multifaceted
chalremarkable transformation driven by the digitalisaletniognes, which collectively compromise the eficient
develof a wealth of health-related data and the adventoopfmBeingt of (big) data-driven solutions in health4c, a5r].e [
Data analytics (e.g., machine learning -ML- technologieHs)e.althcare data management, like any other data
manThese developments have opened the path for novel daatgae-ment system, requires means to ingest, store, process
driven techniques where the incorporation of ML taonoldsanalyze data. However, the specificities of this
dohas enabled a shift from subjective interpretations mtaoin have been traditionally ignored in the general field
a more objective and accurate approach in diagnosotficdsata management, but they naturally raise new
chaland treatment1][. A representative example of this apl-enges when tackling projects such as INCISIVE, which
proach is the INCISIVE projec2t],[ a major European we summarize as follows:
initiative3][ that aims to create an interoperable fedF-ederated data management for healthcare is a must
erated pan-European data repository with secure dsiantcae these projects require minimizing centralized
apthe diagnosis, prediction, and monitoring of c1a.ncer (even if distributed) management system sindactea
govsharing and distributed analytical capabilities relaptreodatcohes. Indeed, it is not acceptable to build a single</p>
      <p>However, the rapid advancement of such initiativeesrinnance would then be centralized. Instead, data assets
the healthcare domain often outpaces the concurrent(id) ea-re often distributed across various providers
(typi</p>
      <p>0009-0001-8246-3767 (A. Bisquert)0;009-0004-0448-3406
(A. Hmimou); 0000-0003-3037-3580 (J. Ll. Berral);
0000-0002-5548-3359 (A. Gutierrez-Torre0)0;00-0001-6350-8328
(O. Romero)
CEUR
htp:/ceur-ws.org
ISN1613-073</p>
      <p>CEUR</p>
      <p>Workshop ProceedingsC(EUR-WS.org)
1https://incisive-project.eu/
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License
cally in diferent medical centres or even distributed in
the same medical centre), making it dificult to access and
share critical information, and (ii) come with a strong
sense ofdata ownership from health institutio6n]s [
federation. For this reason, distribution alone is not a
solution and, instead, a federated data governance protocol
should be defined to provide clear guidelines to enable
healthcare organizations to harness data efect4i]v.ely [</p>
      <p>Data privacy and security compliance present huge
obstacles for researchers in this fiel7d].r[eviews privacy
preservation methods used in healthcare, including
encryption and anonymization, pointing their limitations.</p>
      <p>In addition, it is typically ignored that privacy antdhaset-all together provide answers to the challenges above
curity principles also demand that data computatiidoennstified: The Data Product Layer, the Federated
Compumust be executed locally and data cannot be moved frtoamtional Governance Layer and the Data Platform Layer.
where it resides. In this conteFxetd,erated Learning The Data Product Layer defines all the data products (i.e.,
(FL) is a promising solution for this probl8e]m. H[ow- the data assets) federated into the system. The Federated
ever, within the healthcare domain, a lack of consenCsoumsputational Governance Layer contains the artefacts
on global privacy policies for enhancing data sharingthoamsanage and govern all the data products. Finally, the
a detrimental impact on research stu9d]i.es [ Data Platform Layer acts as a gateway to utilize all the</p>
      <p>Data variety is a major challenge in healthcare dadtaata management processes and workflows provided by
management since this domain encompasses data rteh-e system, such as registering a new data source or
perlated to diagnosis, testing, monitoring, treatmentf,oarnmding an analytical study over selected data products.
health data stored in heterogeneous storage systeAmts, the core of the Federated Computational
Goveroften in varying standards and formats1[0]. Variety nance Layer lies Saemantic Data Model, which
capmay refer to (i) standards and format-related issues stinucrees the relationships between data products and the
healthcare data is typically produced following stanpdraersdcsr.iptive guidelines established byctohnesortium
[11] reviews predominant standards including openE H(iR.e,., the governing body of the resulting federated
sysISO13606, HL7, DICOM, etc. that are serialized follotwem-) together with semantic metadata. These guidelines
ing specific formats (e.g., JSON, CSV). Also, it may referhave a computational nature, serving to validate the
into (ii) hardware-specific issues, since a fair portion otefgrity and compliance of data products with the
conhealth data is generated by medical equipment (e.g.,sCoTrtium agreements. The model portrays the system’s
scan, X-ray, ventilators, etc.). The use of diverse mcoodm- plexity and facilitates the automation, integration,
els of equipment can introduce bias to measuremendtiss,covery and governance of data products while
guarstemming from variations in manufacturing origins, tahneteeing compliance with the policies defined. Further,
utilization of specific methods for image production, anadnalytical pipelines that need to adhere to specific
polidiferences in scans, such as varying backgrounds (e.gc.,ies can also be represented in the semantic model to
black vs. white) or alterations in image contrastg.ovInern analytical studies as well.
addition, addressing variety in healthcare also requiCreosntributions. HealthMesh is a novel architectural
precise domain interpretation which implies that dfraatmaework for federated healthcare data management
must be interoperable at tsehmeantic level [11]. Oth- with the following contributions: (i) it introduces a
doerwise, it may compromise the quality of care providmedain decentralized paradigm, grounded on the data mesh
to patients and waste resour1c2e]s. [ concept, granting autonomy and data ownership of
fed</p>
      <p>Despite the relevance of these challenges, current setraattee-d data products within healthcare institutions. At
of-the-art architectural solutions sufer frpormivacy- its core, (ii) it incorporates a Semantic Data Management
decentralization trade-of . Solutions collecting datmaodel, which governs the federated data products and
centrally fall short of privacy and security needs whegrueaarsantees their compliance with privacy and security
current distributed solutions do not provide meanpsofloircies set by the consortium. And (iii), this layer
fagoverning a federation and, therefore, lacking a federcailtietdates the discoverability of relevant data products,
fagovernance model and interoperability framework. cilitates their integration (overcoming data variety) and</p>
      <p>To cover the above-mentioned challenges, we presetnrtiggers federated learning by means of robust
goverHealthMesh, an innovative architectural framewornkafnocre mechanisms.
federated healthcare data management. HealthMesh is
grounded on the principlesDoafta Mesh [13],
advocating for the decentralization of heterogeneous2d.atRaelated Work
assets into autonomous and independent units referred</p>
      <p>Current solutions for healthcare data management fall
to as d“ata products” that can execute code locally and</p>
      <p>short of covering the gaps discussed in the motivation.
share results with the federation. Simultaneously, this</p>
      <p>Specifically, they either provide a centralized approach
approach fosters data ownership and accessibility due</p>
      <p>not meeting the privacy and security requirements or
to a domain-decentralized organization: i.e., data
prod</p>
      <p>they support distribution but not the creation and
manucts are categorized into domains and associated with</p>
      <p>agement of a federation. Further, the challenges
introconsensus-driven policies (constraints of use) and
available analytical services. HealthMesh also incorpodrautceeds by data variety are not properly addressed.</p>
      <p>Many current healthcare solutions are based on Data
a data governance layer responsible for managing these</p>
      <p>Warehousing architectures that prevent the unleashing
data products towards the establishment of a dynamic
yet robust federated big data ecosystem. of the potential of health da1t4a].h[ighlights their
limi</p>
      <p>The HealthMesh framework comprises three layetrastions in scalability, interoperability and privacy. As an
evolution, Data Lakes rely on Cloud infrastruc1t5u].respr[oject. Further, the data variety aspect, specifically in
However, these solutions sufer from several limitatiohnesa,lthcare, is not considered. However, there is yet no
especially privacy concern1s6][ but also the fact that thaevyailable federated data platform covering all the
probdo not create a federation but a distributed data malneamgsep-reviously discussed. For this reason, we propose
ment system. Current solutions trying to address privHaecaylthMesh, a novel architectural framework
addressconcerns (e.g.,1[7], which discusses the Blockchain benin-g the privacy-decentralization trade-of efectively and
efits and limitations in healthcare) fall into scalaboipleitraytionalizing it for the healthcare domain, while
proand interoperability problems. Approaches discussveiding means to tackle data variety in this domain.
above are either centralized, falling short with privacy
concerns (Data warehouse, Data Lake), or decentralized
falling short with governance and interoperability3.. The HealthMesh Framework
In response to this dilemma, innovative
architec</p>
      <p>HealthMesh is composed of a set of defined requirements
tures supported by semantic-based solutions were raised,
which tackle the lack of data governance in other aarncdhai-n architecture design which includes descriptions
of the components, roles and workflows. We pay special
tectures. Specifically, data governance may be definedattention to the Federated Computational Governance
as to what decisions must be made to ensure efective data Layer, which sits at the core of HealthMesh.
management and data usage and who makes the decision
(locus of accountability for data assets) [18]. In this
category, we focus on two: Data Fabric and Data Mesh. 3.1. Requirements</p>
      <p>Data Fabric1[9] is defined as a collection of architecH-ealthMesh must cover the whole data life cycle
followtural principles as specific modules. Based on a knowiln-g the challenges previously defined.
edge graph (data catalog), the architecture enables
work</p>
      <p>Functional Requirements: (i) Data registration:
Ining with data at the logical level instead of at the physical</p>
      <p>corporate new data assets into the big data system. (ii)
level through data virtualization, providing robusDtadtaatdaiscovery: Search and filter capabilities of the
ingovernance and interoperability. However, defining and</p>
      <p>gested data using metadata
parametersD.a(itiai)analymanaging data by a central organization, as discussisse:dAbility to perform diferent types of analytical studies
by the authors, make it fall into privacy and securituysiinsg- data assets of interest.
sues, following the same pattern observed in centralizeNdon-Functional Requirements: (i)
Domainapproaches. Indeed, this solution, like other semandteicce-ntralization: Data assets should be
domainbased solutions, does not allow the creation of a data</p>
      <p>decentralized meaning that they should be organized
federation. and aligned with the federation policies and analytical
Data Mesh 1[3] is a decentralized architecture built</p>
      <p>requirements. (iiC)ompliance: A contract is established
upon four fundamental principles. FirstDleyc,e“ntral- between a consortium and the owners of a data asset.
ized domain data ownership” advocates for ensuring thaAtny federated data assets should be compliant with
those closest to the data take control. SecDoantdalays, “ the contract and therefore respond to the expected
a product” emphasizes the integration of data, metadata,</p>
      <p>behaviour agreed. (iiPir)ivacy and Security: Data must
and code as a logical unit for sharing. Thirdly, therceomna-in where it resides. Only processed results, in the
cept of a s“elf-serve data platform”empowers data owners form of aggregates, can be retrieved, using Federated
to manage the entire life cycle of their data prodLuecatrsn.ing techniques. Individual data pieces should
Lastly, F“ederated Computational Governance” establishes never be compromised. (iv)Interoperability: The
a model that strikes a balance between domain auitnofnra-structure must facilitate the integration and usage
omy, global conformance, interoperability, and securoiftynew heterogeneous data assets, regardless of the
within the mesh. Data Mesh advocates for the decensttraanl-dard, format or hardware-specific issues. We refer
ization of data assets, emphasizing data ownership and</p>
      <p>to this as semantic interoperability among data assets.
team autonomy, ultimately enhancing data quality and
unlocking the full potential of analytical in2s0i]g.hts [</p>
      <p>The analysis of the existing literature reveals a g3a.p2i.n Running example: Breast Cancer
the current architectural solutions, particularly in the aAbn-alytics within the INCISIVE project
sence of a robust decentralized framework able to provide</p>
      <p>In this paper we will use the INCISIVE project, briefly
federated governance and privacy measures. The
theo</p>
      <p>introduced in the introduction, as a running example.
retical concept of Data Mesh is a promising alternative</p>
      <p>One of the most crucial application areas of INCISIVE
to properly manage all the factors previously mentioned.</p>
      <p>is that oBfreast Cancer Analytics. This use case is
However, this paradigm sits at a high level of abstraction</p>
      <p>based on [21], which introduces a comprehensive
malacking concrete descriptions and definitions, which does</p>
      <p>chine learning solution for mammography classification
not allow to operationalize their principles in a given
The goal of this layer is to manage and govern data
products. This layer is maintained byFeaderated Team
that provides the guidelines for all data products to be
discovered, integrated and consumed. This is a
multidisciplinary team consisting of domain experts. Platform,
legal and analytical experts create the guidelines
(constraints to guarantee when federating a data product)
and features (i.e., specific analytical services) for data
products in a consensus-driven way by means of the
global definitions, policies and analytical services
components. Healthcare institutions negotiate and establish a
contract with the federated team when registering their
data assets.</p>
    </sec>
    <sec id="sec-2">
      <title>Global Definitions. Global Definitions provide</title>
      <p>means to enable governance and interoperability, and
include the set oDfomains  , Common Data Models
usingBIRADS score, which is a quality control syste m and the referencOentology    . Domains 
that refers to the mammography assessment categoraierse. a set of healthcare disciplines given their analytical
requirements, providers, etc. The federated team is
re</p>
      <p>Example. Figure1 describesdata asset 1, owned by sponsible for the definition and evolution of domains.
Hospital A, a XNAT server2[2] with mammography im- Consequently, every data product must be associated
ages inDICOM format alongside its metadata (PatientwIDit,h at least one specific domain. Every domain has
owner,BIRADS, etc) annotated in the same file headeras.Common Data Model that functions as a data
Analogouslyd,ata asset 2 (owned byHospital B), is a file standard essential to enable interoperability. It sets a
system withTIFF mammography images also with sim- structure and content for the data a s sets. are
ilar metadata but stored separately in an ad-hoc Evoxcealbularies (i.e., the day-by-day terminology used by
ifle. In this example, in both data sets, patient identifieernsd users), typically in the form of ontology, that
enhave been anonymized. Also, within the same data typaebsl,e precise interpretation of data and, therefore, remove
there may be diferences which should be treated, e.g. tahmebiguity when interpreting the data meaning.
diference of contrast in images due to diferent scanner
machines (diferent brands, models...). Example. Data assets 1 and 2 are assigned to the</p>
      <p>In the following, we will show how to manage a”nBdreast Cancer Analytics” domain 1. Moreover, the
facilitate the integration of these heterogeneousfeddaertaated team agrees to uDsIeCOM as Common Data
assets to enable researchers to perform a federated sMtuoddyel (  ), a widely used standard for imaging
to obtain a single BIRADS score classification model dbyata, andSNOMED CT2 vocabulary  (   ),
using HealthMesh. one of the largest and most widely used collections of</p>
      <p>OWL vocabularies that enable sharing medical records,
3.3. Architecture Design clinical trials, and other healthcar1e1d].ata [
HealthMesh (see Figur2e) includes three layers: the FedC-omputational Catalogues. The computational
erated Computational Governance Layer, the Data Pcraotadl-ogues  store the Policy Check ertshat
impleuct Layer and the Data Platform Layer. In our approach,</p>
      <p>ments the agreed Polic ie.sAll the procedures stored
data assets are registered and represented as dataipnrtohde- computational catalogues are defined over the
ucts. Data products are decentralized self-contained</p>
      <p>previously defined to allow interoperability across
hetentities encompassing comprehensive elements,
includ</p>
      <p>erogeneou s .
ing data, metadata, and accompanying code responsible
for their maintenance. Every data product must have a
designateddata owner responsible for its accessibility Policies and Policy Compliance Checkers.
Poliand maintenance. In this section we introduce the comcpieos-  ∈  are defined by the federated team and
emnents of each layer and explain their functionalitiebsobduytt,he diferent guidelines that data products must be
due to space constraints, we focus on the most relevcaonmtpliant with. Regulatory experts within the domains
ones that show the feasibility of the overall approach.</p>
      <p>2https://www.snomed.org/
come together and agree that all the related data maugsrtebeed to adhere)(to fulfil the analytical requirements.
compliant with relevant laws, regulations, and industry
standards related to the handling, processing, and storEaxgaemple. In our running example, BIRADS
mamof data. Similarly, domain-specific policies are defined tomogram classification 1 is defined by the analytical team
validate data integrity within its context. within the domain1. Legal representatives in the
fed</p>
      <p>Policy compliance checkers  ∈  are computa- erated team define that data consumed b1yshould be
tional resources implementing Policies per domain.cIonmpliant with1, which states that personal data must
HealthMesh we implement them as test functions tobebecollected, processed, and stored in compliance with
executed on data products. T h usi,s shipped and ex- privacy regulations such as GDPR, CCPA, HIPAA, etc.
ecuted on each data product and, if a data productInfatilhsis context, policy checke1ris the computational
to meet a specific agreed policy for a given domain, thafutnction that addresse1sensuring compliance with a
data product is not available for exploitation. specific typology of data (e.g., there should not be any
personal name or identifiable data). Similarly, po l2icy</p>
      <p>Analytical Services. Analytical experts in the fewd-hich has been specifically defined for this service, states
erated team establish and develop a series of analythicaatl all mammogram image data should be annotated
services ( ), designed to operate on the data produwctitsh BIRADS score using DICOM headers.
within the system, generating aggregated results and
comprehensive reports. An analytical se rviscerelated Semantic Data Model. Data governance is an
essento a specific domain and is tailored to the specific typolt-ial requirement for the proposed architecture. This
comogy of data and the set of policies that the data propdouncetnst orchestrates all the components previously
introduced and defines the metadata needed to descraibtetributes), owner, version, etc. This is unique to each
and govern the data assets. T hus,in each domain data asse t  . specifies the typology of the data asset
must be mapped to its specific data standard ( i.e., ) to properly categorize the data pro d uctc.ontains
and the vocabulary (i. e.,  ). Similarly, the data prodi-nformation about the data product format (e.g., text,
uct should adhere to its policies ag re)eadn(d could annotated images, etc.). The sa me can be used in
be eligible for the analytical serv ice)sd(efined for various domains . contains all information to grant
that domain. All this metadata is described utilizainugthaorization and access to data from a technological
knowledge graph. point of view. includes the data access layer
creden</p>
      <p>Further, the semantic data moΔde(Fligure3) estab- tials and data repository (e.g., access URL) metadata.
lishes the relation between the Data Product metadctastaas an agreement between data providers and the
provided by their owners and the guidelines defined bfyederated team. It maps the data profile schema to the
the federated team to enhance the integration, gFoevdeerr-ated Computational Governance layer to facilitate
nance, and discoverability of data products. The gdoaatla integration. T he definition:
of this model is to provide a consistent semantic model
across the entire framewoΔrki.s aknowledge graph,  = ⟨  = { 1,  2, ...,   },
leveraging its capacity to ofer a holistic and intercon-   = {→ 1, → 2, ..., →  }, (1)
nected perspective of data. Knowledge graphs are a good
choice because they are flexible, heterogenous, intuitive,   = {→ 1, → 2, ..., →  } ⟩
formal and scalabl1e9[]. Further, several previous workcsontains the Data Product Sch e ma)(and the
semanhave discussed the relevance of knowledge graphstitcoattribute mappin gs ( ) to and   ,
respectackle governance in Big Data scenarios (2e3.g]).., [From tively. It also contains the set of policy che ckteors
a logical perspective, A data product)c(an be repre- guarantee its compliance with the domain policaineds
sented as a set o&lt;f  ,   ,  ,  &gt; containing a compatibility with analytical servic.eFsrom a data
inProfile (  ), Dataset Type Template (  ), Tech- tegration perspective, themaps the local data source
nology Aspects (  ) and aData Contract ( ). All schema (i.e.,  ) to the integration schema (i.e., the
this metadata is defined at the time of the data prodauncdt   ). This is a direct application of the knowledge
registration proce s s. includes all the metadatgaraph data federation approach presente2d3]i,nw[hich
related to the data asset, including its schema (i.e., leinstabolfes querying the data sources (i.e., the data products)
via the integration schema. Without a valid data conBt.raDcatt,a Product Layer
a data product cannot be part of the federation.</p>
      <p>The semantic data modΔelis the key component toData products ( ) are self-contained entities
encompassing data, metadata and code. Therefore, physical data
guarantee that heterogeneous medical data assets can be</p>
      <p>assets are stored and maintained by participating
instituefectively integrated, categorized, accessed, and
main</p>
      <p>tions/providers. This approach promotes data ownership
tained through the utilization of the resources previously
defined and, from a semantic point of view, acts as anand autonomy and is strongly favoured by hospitals and
orchestrator. Furthermore, leveraging ontology ladnat-a owners6[].
guages such as OWL or DL-Lite famil2y4[], the semantic Data Product owners are responsible for the life cycle
of the data product and its maintenance. Data owners
data model can benefit from reasoning to validate the
resultingΔ and infer additional informati2o5n,2[4]. are the ones closest to the data and they can understand
how it should be interpreted within each domain.</p>
      <p>Sidecar An adjunct component in the form of a
sideAlgorithm 1 Data product registration</p>
      <p>car ( ) facilitates seamless integration with the broader
Require: ,  ,   mesh ecosystem. The sidecar is installed inside the
in  ←Δ.recTemplate(  ) stitution/provider infrastructure but it is maintained by
,  ←Δ.generateMappings(  ,  ) the platform representatives of the federated team.
 ′ ←Δ.getPolicies(, ,  ) can retrieve the data of a data product through the data
for  in  ′ do access layer specified in  . Each contains a Data
 ←Δ.getPolicyChecker( ,   ) Contract that is retrieved froΔm.</p>
      <p>′.AddPolicychecker( ) Algorithm2 illustrates the process of consuming a
end for for a specific  . Each time a data product is consumed,
 = &lt; ′, ,  &gt;  validates i t s to verify that data adheres to the
Δ.addDPMetadata(  ,   ,   ,  ) mappings and policies specified. If data products are not
interoperable or compliant, comprehensive reports are</p>
      <p>Following Algorith1m,  and  are provided given to the data product owner specifying the errors
by the data product owners and the domain assignoebdtained during validation. This way, the integrity and
by the federated team. Wi t h , the semantic datacompliance of the data product are always validated in
modelΔ determines the most suita ble . Based on that,run-time guaranteeing that it conforms with its most
and using as input the global definiti ons and profile recent contract. If none of the reports has facialendb,e
  it semi-automatically generates the mapp ingesxecuted through the Sidecarto process the validated
and to and   , respectively. The policies t o . The sidecar returns results in the form of
aggregabe followe d ′ are obtained using the dom ainand map- tions. Therefore, individual data is never compromised.
pings following the approach i2n6][. MoreoverΔ, infers This approach creates a robust security measure while
the respectiv e′ based on ′ and  . To complete the still allowing for analytical tasks to be performed in the
process, all metadata that consti tutiessintegrated context of Federated Analytics.
intoΔ.</p>
    </sec>
    <sec id="sec-3">
      <title>Algorithm 2 Data product consumption</title>
    </sec>
    <sec id="sec-4">
      <title>Example. In our example, both data assets are reRge-quire: , , ,</title>
      <p>istered using Algorith1minto domai n 1. Therefore, as  ←SC.validateRS(, ., . )
input, the data product owner must provide the profil e, ←SC.validateC( , . ,  )
which for simplicity, let us consider only contains thife and are validthen
attribute ”Subject” (which stands as a patient identifier). aggResult ← .executeAS( )
First, HealthMesh would assig n as ”annotated im- else
ages”. Then, with the help of the data owner, who must return Failed reports
supervise the process, the system generates the mappingsend if
to (in this example, we defined  as 
and      as    ). Thus,  2 mappings: Data products configuration strives to adhere to the
→ 2 (   2,   _  ) ∈   2 and→ 2 FAIR principles of data managemen2t7][. It is
character(   2, SCTID:116154003  ) ∈   2. In ad- ized by a concerted emphasis on fostering data ownership
dition, policy checkers1 and  2 are determined toand the enhancement of data quality within the domain
apply for 1 (by checking the policies related to thatodfoh-ealthcare data.
main viaΔ) and added to their respec t ive.</p>
      <p>Example. Within our ongoing case study,1 and
 2 are candidates to be consumed for analytical service
 1. Following2,  1 would scrutinize both 1 and 2 until the global model converges. Furthermore, a similar
through their for any identifiable data. Given thaptrocedure can be used to perform federated exploratory
both data assets are anonymi zed1,is expected to returndata analysis to better understand the underlying data.
successful results, confirming compliance with privacy
standards.</p>
      <p>However, the report obtained thr ough on 4. Conclusions and Future Work
 2 would inform a format issue indicating tha2tis
not available in DICOM format. The report wouldWbeepresented HealthMesh, an architectural framework
sent to HospitalB to state that data should transfdoersmigended for the healthcare domain. Building upon data
to DICOM to adhere to . mesh principles, we present a design encompassing
multiple layers, components and workflows that we
illus</p>
      <p>Considering th at 2 owner applies the necessary
processes over the data to be compliant wit h i,tbsoth trated employing a real ongoing example. HealthMesh
adopts a federated approach, ensuring that data remains
 1 and 2 would be technically and semantically inwteirt-hin healthcare institutions to uphold security and
prioperable in terms of DICOM standard and SNOMED-CT</p>
      <p>vacy. The framework strategically employs a
Semanvocabulary. Moreover, the data would be anonymized</p>
      <p>tic Data Model in conjunction with computational
reand annotated with BIRADS in DICOM format.
Therefore, 1 could be operated in both Data Products. sources to achieve data interoperability and governance.</p>
      <p>HealthMesh is a novel architectural framework in the
ifeld that has been built upon the requirements identified
C. Data Platform Layer collaboratively with experts from INICISVE. Our work
The data platform layer functions aisnatenrface en- has certain limitations that we plan to address in the
compassing various tools/services to enable data prodnueactr future. For example, there is an absence of in-depth
workflows such as (i) data product registration, (ii) dtiesc-hnical considerations due to space constraints and an
covery and (iii) execution of federated analytical taesxkpse.rimental evaluation with real data in real
scenar</p>
      <p>Data Consumers andData Product owners use the ios. Currently, HealthMesh is a relevant step in the right
platform to perform analytical studies and managedtirheection collecting concepts of relevance, their
relationData Products, respectively. ships, and the identification of key actors, which is a key</p>
      <p>Data product registration is a process to incorporatecontribution in the complex and limited field of federated
new data assets into the system. The process is semdai-ta management for healthcare.
automated with the supervision of a data product ownTerh.e development of HealthMesh opens the door for</p>
      <p>Data discovery requires a query () provided by data future work. For example, to study how blockchain can
consumers containing keywords and/or filters in terbmesintegrated into the framework, the potential of Graph
of  . The function leveragΔesto efectively identify Neural Networks leveraging the Semantic Data Model,
the most appropriate data products. etc. Last, but not least, we also plan to explore the
fea</p>
      <p>This architectural framework is specifically designseidbility of generalizing this solution to other domains
to enable and enhance secuarnealytical tasks in the requiring a data federation (e.g., Data Spaces).
realm of Federated Analytics, including Federated
Learning. Upon selection of desired data products by dAatcaknowledgments
consumers, a can be performed over the
interoperable versions acquired through Algori2tthomgenerate The work reported here was supported by the EU-H2020
results. programme under GA.952179 INCISIVE, Horizon</p>
      <p>Europe Programme under GA.101095717 (SECURED)</p>
      <p>Example. In our ongoing use casDe,ata asset 1 and and GA.101135513 (CYCLOPS), the Spanish Research
Data asset 2 are registered bHyospital A andHospital B State Agency (MICINN/AEI, ERDF/FEDER) under
data assets owners as 1 and 2, respectively. DataGA.MCIN/AEI/10.13039/501100011033/FEDER (UE
consumers can use the Data Discovery interface to poDstAaLEST), the Spanish Ministerio de Ciencia e
Inquery 1 containing the keywords ”Breast Cancer” tonliosvtación under project PID2020-117191RB-I00 /
all data products related to do m1 asiunch as 1 and AEI/10.13039/501100011033 (DOGO4ML), and the
 2 throughΔ. In this context, data consumers can Gsee-neralitat de Catalunya (AGAUR) 2021-SGR-00478
lect 1 in the Analytical Service Interface to be execu(CtRedOMAI).
over the previously discovered data products.
Consumption of both1 and2 using 1 would provide a local
classification model. The local models can be aggregated
into a global model for BIRADS breast cancer
classification. Notice that this process could be done iteratively
data architectures, Procedia Computer Science 196 doi:10.1038/sdata.2016.18.
(2022) 263–271. doi:10.1016/J.PROCS.2021.12.</p>
      <p>013.
[21] I. Tzortzis, S. Sykiotis, I. Rallis, N. Doulamis, An
integrated framework for classifying mammograms
according to birads scale and breast tissue density., in:
Proceedings of the 16th International Conference
on PErvasive Technologies Related to Assistive
Environments, PETRA ’23, Association for Computing
Machinery, New York, NY, USA, 2023, p. 728–731.</p>
      <p>doi:10.1145/3594806.3596577.
[22] D. S. Marcus, T. R. Olsen, M. Ramaratnam, R. L.</p>
      <p>Buckner, The extensible neuroimaging archive
toolkit: an informatics platform for managing,
exploring, and sharing neuroimaging data,
Neuroinformatics 5 (2007) 11–34. do1i:0.1385/ni:5:1:11,
pMID: 17426351.
[23] S. Nadal, A. Abelló, O. Romero, S. Vansummeren,</p>
      <p>P. Vassiliadis, Graph-driven federated data
management, IEEE Trans. Knowl. Data Eng. 35 (2023)
509–520. doi:10.1109/TKDE.2021.3077044.
[24] D. Calvanese, G. De Giacomo, D. Lembo, M.
Lenzerini, A. Poggi, M. Rodriguez-Muro, R. Rosati,
Ontologies and Databases: The DL-Lite Approach,
Springer Berlin Heidelberg, Berlin, Heidelberg,
2009, pp. 255–356. doi:10.1007/978- 3- 642- 037
54-2_7.
[25] X. Chen, S. Jia, Y. Xiang, A review: Knowledge
reasoning over knowledge graph, Expert Systems
with Applications 141 (2020) 112948. dhoti:tps:
//doi.org/10.1016/j.eswa.2019.112948.
[26] J. Flores, K. Rabbani, S. Nadal, C. Gómez, O. Romero,</p>
      <p>E. Jamin, S. Dasiopoulou, Incremental schema
integration for data wrangling via knowledge graphs,
Semantic Web – Interoperability, Usability,
Applicability accepted, tbp (2023). URhLt: tps://www.se
mantic-web-journal.net/content/incremental-sch
ema-integration-data-wrangling-knowledge-gra
phs-0.
[27] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg,</p>
      <p>G. Appleton, M. Axton, A. Baak, N. Blomberg, J. W.</p>
      <p>Boiten, L. B. da Silva Santos, P. E. Bourne, J.
Bouwman, A. J. Brookes, T. Clark, M. Crosas, I. Dillo,
O. Dumon, S. Edmunds, C. T. Evelo, R. Finkers,
A. Gonzalez-Beltran, A. J. Gray, P. Groth, C. Goble,
J. S. Grethe, J. Heringa, P. A. t Hoen, R. Hooft,
T. Kuhn, R. Kok, J. Kok, S. J. Lusher, M. E. Martone,
A. Mons, A. L. Packer, B. Persson, P. Rocca-Serra,
M. Roos, R. van Schaik, S. A. Sansone, E. Schultes,
T. Sengstag, T. Slater, G. Strawn, M. A. Swertz,
M. Thompson, J. V. D. Lei, E. V. Mulligen, J.
Velterop, A. Waagmeester, P. Wittenburg, K.
Wolstencroft, J. Zhao, B. Mons, The FAIR guiding
principles for scientific data management and
stewardship, Scientific Data 2016 3:1 3 (2016) 1–9.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>