<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Minimum Metadataset for Data Lakes Supporting Healthcare Research</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Piantella</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierluigi Reali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Priyansh Kumar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Letizia Tanca</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano - Department of Electronics</institution>
          ,
          <addr-line>Information, and Bioengineering Via G. Ponzio 34/5, 20133 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>32</volume>
      <fpage>23</fpage>
      <lpage>26</lpage>
      <abstract>
        <p>While data lakes have emerged as a solution for storing vast amounts of heterogeneous and often unstructured data, responding to the growing need for flexible data storage, integration, and analytics in diferent domains, the digital transformation of healthcare processes has led to an exponential increase in various types of health records, necessitating eficient data management solutions and making this domain an ideal arena for experimenting data lake eficacy. In data lakes, efective metadata extraction and management are crucial for describing raw data, establishing connections, and ensuring interoperability among datasets ingested into the lake. To address this, we propose a minimum set of metadata tailored for clinical research, which includes relevant information common to significant branches of healthcare. Our metadataset not only streamlines data ingestion processes but also enhances the accessibility and usability of healthcare datasets for research purposes. By standardizing the collected metadata within the clinical research domain, we also facilitate data integration, analysis, and exploration, facilitating comprehensive data description and management within the data lake environment.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;medatata</kwd>
        <kwd>healthcare</kwd>
        <kwd>data lakes</kwd>
        <kwd>interoperability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Responding to the pressing demand for flexible and easily-accessible data analytics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], an
emerging trend involves data lakes as repositories for vast amounts of data and documents in
the big-data context [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Notably, data lakes operate without a predefined schema, enabling the
ingestion of raw data in various formats (including relational data, images, text, data streams,
and logs) without the need for prior preprocessing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This adaptability empowers users and
organizations to seamlessly store and access their data, facilitating data analytics, data-driven
applications, and machine learning tasks.
      </p>
      <p>
        In the field of medicine, the transition to digital healthcare processes and services has led to
an exponential increase in medical data. Within hospitals, daily operations generate a multitude
of (often unstructured) digital documents, including medical images, nursing notes, discharge
letters, and laboratory results. Moreover, advancements in medical devices, applications, and
monitoring technologies have digitized patient data, resulting in the collection, analysis, and
storage of vast amounts of heterogeneous information. In fact, it seems that, by 2025, the annual
growth rate of healthcare data will surpass that of generic data, reaching 36% compared to
circa 27% [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These challenges make the realm of medicine the ideal one for experimenting the
efectiveness of the use of Metadata.
      </p>
      <p>
        With the increasing availability of Electronic Health Records facilitating real-world-evidence
clinical trials [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a significant application of healthcare data management is medical research.
In this context, the ability to collect and analyze data from heterogeneous sources is crucial [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
and, given the diverse formats of healthcare data and its sheer volume, a data lake is a very
interesting solution. Since the datasets ingested by a data lake are extremely heterogeneous,
accessing and manipulating the stored raw data can be very expensive in terms of computational
and time complexity, therefore efective metadata extraction and management, establishing
connections among the ingested datasets [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], are essential for describing raw data. In fact,
metadata provide valuable information regarding the data without the need to directly analyze
the datasets.
      </p>
      <p>To achieve this, we propose a minimum set of metadata (i.e., a minimum metadataset) for
the context of clinical research, which encloses the relevant information common to the main
branches of healthcare. Data feeders can then specify additional metadata that further describe
the datasets.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology and Related Work</title>
      <p>We consider metadata models and tools specifically tailored for the healthcare context.</p>
      <p>Our primary objective is to construct a minimum metadata model that not only ofers essential
information relevant to associated healthcare data but also creates a distinctive framework
that facilitates the sharing of clinical data coming from diverse formats and sources. This
improves interoperability, which, in turn, supports seamless data exchange and collaboration
across diferent healthcare organizations. Our proposed minimum metadata model serves as a
foundation that can be further enhanced and specialized for each specific scope of use.</p>
      <p>We now briefly describe the existing clinical metadata models and management tools we
analyzed, which contributed to the design of our minimum metadataset.</p>
      <sec id="sec-2-1">
        <title>2.1. Genosurf</title>
        <p>
          Genosurf [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is a metadata integration and search system designed to eficiently analyze
genomics datasets from various sources in biological and clinical research settings. It leverages
a Genomics Conceptual Model (GCM) and implements a multi-ontology semantic search system.
The metadata repository includes millions of metadata entries from multiple datasets, focusing
on significant genomics data. The system ofers a web-based interface that allows users to
perform targeted searches based on specific metadata attributes and values. In this way, users
can inspect descriptions of matching datasets, explore the related metadata, and obtain the link
to the original datasets. Moreover, Genosurf facilitates free-text searches and ofers query
preparation functionalities for further data processing.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. PDXFinder</title>
        <p>
          Patient-Derived tumor Xenograft (PDX) models are essential tools to study the efects of
chemotherapy on tumors. PDXFinder [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] provides centralized access to an extensive collection
of PDX models. It supports advanced search functionalities that enable users to filter and refine
their searches based on specific criteria such as cancer type, molecular characteristics, and
treatment history. As a result, researchers can access detailed information about each PDX
model, including clinical annotations, molecular profiling data, histopathological features, and
associated research publications.
2.3. HL7 - FHIR
The Health Level Seven International (HL7) Fast Healthcare Interoperability Resources
(FHIR) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is widely acknowledged as a fundamental metadata model for achieving healthcare
data interoperability, presenting a standardized approach to data representation and exchange.
FHIR provides extensive information, encompassing various aspects of healthcare data, such
as patient demographics, clinical observations, medications, and procedures. These resources
are designed to be easily accessible using RESTful APIs[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], further enhancing its appeal and
ease of implementation. The standardized nature of FHIR and its support for RESTful APIs
can enable data exchange and sharing between diverse healthcare systems, regardless of their
underlying technology and platforms.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4. Datacite</title>
        <p>
          Datacite [
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ] is an internationally recognized organization that provides persistent
identiifers, known as DOIs (Digital Object Identifiers), for research data. Although not specifically a
metadata model, Datacite significantly contributes to data discoverability, access, and reference.
By assigning DOIs to research datasets, Datacite ensures their long-term accessibility and
establishes a standardized approach for referencing and linking data. The metadata ofered
by Datacite includes essential details about the dataset, such as its title, authors, publisher,
publication date, version, and any related resources. This metadata is usually presented in a
standardized format, leveraging a metadata schema, defining specific data elements and their
required or recommended attributes.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.5. EOSC FAIR principles</title>
        <p>The European Open Science Cloud (EOSC) initiative1 aims to create a seamless and open
research environment by providing access to research data, services, and infrastructures across
Europe. EOSC recently published its guidelines and recommendations to promote the Findability,
Accessibility, Interoperability, and Reusability (FAIR) of research data and services [14]. The
EOSC FAIR principles emphasize the importance of making research data and related resources
easily discoverable, accessible, and interoperable. By adhering, researchers and data providers
1https://digital-strategy.ec.europa.eu/en/policies/open-science-cloud. These principles align with the broader FAIR
data movement, which seeks to maximize the value and impact of research data by ensuring its usability and
long-term preservation.
adopt standardized metadata (categorized as mandatory, recommended, and optional), data
formats, and interoperability standards. This enables eficient data discovery, access, integration,
and reuse, promoting collaboration, knowledge sharing, and interdisciplinary research within
the EOSC ecosystem.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.6. Standardized terminologies and coding systems</title>
        <p>Standardized terminologies and coding systems are vital for achieving healthcare data
interoperability [15], providing a common vocabulary and coding structure, to ensure that clinical
concepts and metadata are represented in a consistent and standardized manner across diferent
healthcare systems and datasets. For example, SNOMED-CT [16] is a comprehensive clinical
vocabulary widely used in healthcare. It allows for the precise and uniform encoding of clinical
observations, diagnoses, procedures, and other medical concepts. Similarly, LOINC [17] is a
standardized coding system specifically designed for clinical laboratory observations and results.
It provides a unified representation of laboratory tests, measurements, and observations.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.7. Clinical Document Architecture</title>
        <p>The Clinical Document Architecture (CDA) [18], developed by HL7, serves as a minimum
metadata model for exchanging clinical documents. CDA defines the structure and semantics
of clinical records, enabling the standardized sharing of healthcare information. It enables
interoperability across diferent healthcare organizations by providing a common framework for
representing patient clinical summaries, discharge letters, progress notes, and other healthcare
documents.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Minimum metadataset</title>
      <p>We report in Table 1 our proposed minimum metadataset for healthcare. Following several
research works [19, 20, 21], we decided to employ for our model three main categories: (i)
administrative metadata, (ii) data provenance metadata, and (iii) descriptive metadata.</p>
      <sec id="sec-3-1">
        <title>Category</title>
      </sec>
      <sec id="sec-3-2">
        <title>Administrative</title>
      </sec>
      <sec id="sec-3-3">
        <title>Data provenance</title>
      </sec>
      <sec id="sec-3-4">
        <title>Descriptive</title>
      </sec>
      <sec id="sec-3-5">
        <title>Attributes</title>
        <p>Administrative metadata refer to the administrative aspects of data management,
facilitating efective data governance, management, and administration. They include the following
metadata, regarding ownership, access authorizations, and policies:
• GUID (Globally Unique Identifier) : a unique identifier assigned to each dataset for
identification and referencing purposes.
• Creator: the entity responsible for creating or generating the dataset.
• Owner: the entity that owns the dataset and holds responsibility for its management.
• Rights: the permissions or restrictions associated with accessing and using the dataset.
• Terms of access: the terms and conditions that govern the access and usage of the dataset.
Data provenance metadata serve as a detailed record of the data lifecycle. They ofer valuable
insights on reliability and quality by capturing information about collection methods, processing
steps, and modifications:
• Publication year: the year in which the dataset was oficially published or made available.
• Upload date: the date when the dataset was uploaded into the repository.
• Acquisition method: a description of the acquisition process employed to collect the data.
• Acquisition tools: SW and HW used to collect the data, with the related version details.
• Download URL: it refers to the specific web address that enables users to download the
dataset to their local systems.
• Checksum: a hash value that acts as a verification mechanism for data integrity.
• Encryption algorithm: if, for privacy reasons, the dataset is encrypted, this reports the
algorithm used to protect sensitive information.
• File version: an identifier or label that denotes the version or the revision of the dataset.</p>
        <p>It allows healthcare professionals, researchers, and stakeholders to track and manage
diferent instances of the dataset, ensuring proper documentation and version control.
• Update/modification date : it stores the date when the dataset was last updated or modified.</p>
        <p>This provides valuable information about the currency and freshness of the data, allowing
the users to ascertain the relevance and applicability of the dataset for their specific needs.
• Update frequency: it indicates the regularity or frequency at which the dataset is updated.
Descriptive metadata refer to the content and characteristics of a dataset, freeing the users
from the need to examine the resource itself in detail. This category is essential for classifying
and organizing datasets, enabling eficient search and retrieval, and facilitating decision-making
about which resources better fit the needs of the users:
• File description: a brief description or summary of the dataset, providing an overview of
its purpose, scope, and data content.
• File format: the specific file format in which the dataset is stored (e.g., CSV, XML, or</p>
        <p>DICOM).
• Min age: the minimum age of the patients represented in the dataset.
• Max age: the maximum age of the patients represented in the dataset.
• Ethnicity: the ethnic background of the patients represented in the dataset.
• Patient sex: the sex of the patients included in the dataset.
• Blood group: the blood type of the patients included in the dataset.
• Primary site: the primary anatomical site or organ associated with the data collected.
• Collection site: the location or institution where the data was collected or originated.
• Disease names: names of the diseases or medical conditions primarily represented in the
dataset.
• Disease types: the classification or type of diseases or medical conditions.</p>
        <p>• Disease variants: specific variants or subtypes of diseases or medical conditions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Adequacy of the proposed model</title>
      <p>In Table 2 we compare our minimum metadataset with the models described in Section 2,
leveraging some of the main metadata commonly employed for both general-purpose [19, 20,
21, 22, 23] and healthcare-related [24, 25] domains. Moreover, we evaluated the quality of our
model by studying how it addresses some of the most common challenges encountered in –
but not limited to – clinical data science and integration, which are acknowledged as critical
barriers also in the National Institute of Health (NIH) strategic plan for data science research [26].</p>
      <p>Our model Genosurf PDXFinder CDA/FHIR Datacite EOSC</p>
      <p>
        [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">18, 10</xref>
        ] [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ] [14]
Domain
Terms&amp;Conditions
File details
Content description
File format and structure
Provenance
Publication date
Reference to vocabularies
Access rights
Integrity information
Encryption information
Patient: blood group
Patient: age
Patient: gender
Patient: ethnicity
Observation: disease
Observation: collection site
(Hineagletnhecraarle) Genomic
✓ ✗
✓ ✗
✓ ✗
✓ ✓
✓✓ ∼ ✗ *
✗ ✗
✓ ✗
✓✓ ∼ ✗ +
✓ ✗
✓ ✓
✓ ✓
✓ ✓
✓ ✓
✓ ✓
⋆ License type only.
* Techniques only.
† Software used only.
+ Claimed to be stored, although not displayed.
‡ Predefined ranges only.
      </p>
      <p>Cancer
∼ ⋆
✗
✗
✓
✓
✗
✗
✗
✗
✗
✗
∼ ‡
✗
✗
✓
✓</p>
      <p>Health General General
documents purpose purpose
✗ ✓ ✓
✗ ✓ ✓
✗ ✓ ✓
✓ ✓ ✓
✗ ✗ ∼ †
✓ ✓ ✓
✓ ✗ ✗
✗ ✗ ✓
✗ ✗ ✓
✗ ✗ ✗
✓ ✗ ✗
✓ ✗ ✗
✓ ✗ ✗
✗ ✗ ✗
✓ ✗ ✗
✓ ✗ ✗</p>
      <sec id="sec-4-1">
        <title>4.1. Lack of standard structure and policies</title>
        <p>The lack of consistent data standards and formats across diferent healthcare systems poses
a significant challenge in achieving efective interoperability. Healthcare organizations often
employ diverse coding schemes, data structures, and terminologies, leading to inconsistencies
and incompatibilities when exchanging health information. This inconsistency may result in
errors and misinterpretations and possibly make data integration more complex.</p>
        <p>
          Proposed Solution The minimum metadataset we suggest in Table 1 provides a
standardized model for organizing and describing essential information about healthcare data for
research purposes. By adopting this model, healthcare organizations can establish a common
structure for data representation, promoting consistency and compatibility in data exchange.
To ensure high interoperability, this solution is in line with state-of-the-art standards and
frameworks, such as HL7 FHIR [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and DICOM (Digital Imaging and Communications in
Medicine)[27], as well as the others mentioned in Section 2. Moreover, leveraging the proposed
minimum metadata model, healthcare systems could map their local data elements and
terminologies to the standardized model, facilitating accurate interpretation and integration of
health information. Finally, the standardized metadata elements can also help in designing a
data catalog for a data lake that can easily accommodate diferent types of healthcare data.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Privacy concerns</title>
        <p>Healthcare data is inherently sensitive and requires robust protection to maintain confidentiality
and ensure the secure and reliable exchange of health information. The already stringent privacy
regulations implemented in the US (HIPAA [28]) and Europe (GDPR [29]) must comply with
other privacy regulations specific to each country or region.</p>
        <p>Proposed Solution Our model prioritizes data protection and privacy by excluding
sensitive information such as patient names, dates of birth, and unique identifiers that could
potentially disclose patient or clinician identities. The model minimizes the risk of privacy
breaches by carefully selecting and including only non-identifying attributes. This approach
aligns with privacy and security policies, safeguarding the confidentiality of healthcare data
and promoting a secure environment for data exchange and interoperability.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Incomplete and inaccurate data</title>
        <p>The quality and usefulness of healthcare datasets can be compromised by inconsistent data
capture and incomplete or inaccurate data entry practices. These issues significantly impact the
integrity and reliability of the exchanged data. Inconsistent data capture refers to variations in
how data is collected and recorded across diferent healthcare systems or organizations. This
can be due to discrepancies in terminology, coding systems, and acquisition processes, making
it challenging to compare and integrate information accurately. Incomplete or inaccurate data
entry practices further compound these challenges by introducing errors or missing information
into the exchanged data.</p>
        <p>Proposed Solution The model incorporates essential attributes describing data provenance,
ensuring both the elicitation of details regarding the acquisition process and the tracking and
management of diferent versions of datasets over time. These attributes allow healthcare
professionals and researchers to clearly understand the acquisition methods and identify the
most up-to-date version of a dataset, reducing the risk of utilizing outdated or incomplete data.
By clearly indicating the dataset version, our model promotes data integrity and ensures that
users work with the most accurate and complete information.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Data bias</title>
        <p>Data bias is a significant concern in healthcare research [ 30, 31], as it can lead to unequal
treatment, inaccurate research findings, and disparities in patient outcomes. Bias can arise
from several aspects, e.g., the demographics of the population sampled, the methods used to
collect and analyze data, and other intrinsic biases. For example, if a dataset primarily includes
information from individuals of a certain age or ethnicity, the findings and conclusions drawn
from that data may not be applicable or representative of the broader population. Similarly,
biases can occur when selecting variables to be measured, leading to incomplete or skewed
representations of health conditions. Addressing data bias is crucial to ensure fair and reliable
research insights.</p>
        <p>Proposed Solution Addressing data bias is a complex task that requires a multifaceted
approach. While our proposal focuses on a minimum metadata model, this does not solve
the problem completely. Therefore, scientists, researchers, and medical professionals must
employ various methodologies to tackle this issue comprehensively [32, 33]. We recognize
the significance of including attributes such as ethnicity, sex, and collection site, which can
help researchers and professionals analyze the demographic and geographical scope of the
datasets, thus assessing potential biases and accounting for them in their analyses. By combining
the strengths of the minimum metadata model, which addresses the identification of possible
data bias through attribute inclusion, with other approaches [34, 35, 36], researchers can work
towards mitigating and minimizing data bias, ultimately enhancing the quality and fairness of
their research outcomes.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.5. Data discovery</title>
        <p>The rapid generation of healthcare data brings the dificulty of finding relevant datasets for
specific research or clinical purposes, in terms of required variables, population demographics,
or specific clinical parameters. This issue is further compounded by the lack of standardized
data formats, inconsistent data labeling, and varying data storage practices across diferent
healthcare systems and organizations.</p>
        <p>Proposed Solution Researchers can leverage the proposed metadata model to establish a
standardized framework for organizing and describing essential information about healthcare
datasets. This includes attributes such as data structure, variables, demographics, diseases, and
anatomical sites. Data consumers can then utilize this standardized metadata to eficiently
search and filter through the vast amount of available datasets.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and future works</title>
      <p>Metadata can be used to support the storage, retrieval and analysis of complex datasets without
the need to directly accessing raw data.In this paper we demonstratedthis possibility using
the example of clinical metadata, which shows the essential information that data feeders
should attach to each dataset before ingesting it into a data lake. We have shopwn as well
that the use of metadata enhances data findability across multiple datasets, helping researchers
acquire suitable data for their studies. Future extensions of this work will include bounding the
values of the metadata fields to specific vocabularies to reduce representation ambiguities. In
the clinical domain, an attractive solution could be exploiting the Unified Medical Language
System (UMLS) [37], a controlled compendium of medical vocabularies including, among others,
SNOMED-CT [16] and LOINC [17]. The multi-language support of UMLS could certainly
facilitate the adoption and usage of our minimum metadataset by clinicians and researchers.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was carried out within the MICS (Made in Italy - Circular and Sustainable) Extended
Partnership and received funding from Next-Generation EU (Italian PNRR - M4 C2, Invest 1.3
D.D. 1551.11-10-2022, PE00000004). CUP MICS D43C22003120001.
[14] O. Corcho, M. Eriksson, K. Kurowski, M. Ojsteršek, C. Choirat, M. Van de Sanden, F.
Coppens, EOSC interoperability framework, Report from the EOSC Executive Board Working
Groups FAIR and Architecture, 2021.
[15] O. Bodenreider, R. Cornet, D. J. Vreeman, Recent developments in clinical terminologies</p>
      <p>SNOMED-CT, LOINC, and RxNorm, Yearbook of medical informatics 27 (2018) 129–139.
[16] K. Donnelly, SNOMED-CT: The advanced terminology and coding system for ehealth,</p>
      <p>Studies in health technology and informatics 121 (2006) 279.
[17] C. J. McDonald, S. M. Huf, J. G. Suico, G. Hill, D. Leavelle, R. Aller, A. Forrey, K. Mercer,
G. DeMoor, J. Hook, W. Williams, J. Case, P. Maloney, LOINC, a universal standard for
identifying laboratory observations: a 5-year update, Clinical chemistry 49 (2003) 624–633.
[18] R. H. Dolin, L. Alschuler, S. Boyer, C. Beebe, F. M. Behlen, P. V. Biron, A. Shabo, HL7
clinical document architecture, release 2, Journal of the American Medical Informatics
Association 13 (2006) 30–39.
[19] C. Lagoze, C. A. Lynch, R. Daniel Jr, The Warwick Framework: A Container Architecture
for Aggregating Sets ofMetadata, Technical Report, Cornell University, 1996.
[20] A. J. Gilliland, Setting the stage, Introduction to metadata 2 (2008) 7.
[21] U.S. National Archives, Metadata in electronic records
management, https://records-express.blogs.archives.gov/2016/11/21/
metadata-in-electronic-records-management/, 2016. Online; accessed April-2024.
[22] R. Gabriel, T. Hoppe, A. Pastwa, Classification of metadata categories in data warehousing
- A generic approach, in: Sustainable IT Collaboration Around the Globe. 16th
Americas Conference on Information Systems, AMCIS 2010, Lima, Peru, August 12-15, 2010,
Association for Information Systems, 2010, p. 133.
[23] J. Greenberg, A quantitative categorical analysis of metadata elements in image-applicable
metadata schemas, Journal of the American Society for Information Science and
Technology 52 (2001) 917–924.
[24] J. Pierson, L. Seitz, H. Duque, J. Montagnat, Metadata for eficient, secure and extensible
access to data in a medical grid, in: Proc. 15th International Workshop on Database and
Expert Systems Applications, 2004., IEEE Computer Society, 2004, pp. 562–566.
[25] R. Badawy, F. Hameed, L. Bataille, M. A. Little, K. Claes, S. Saria, J. M. Cedarbaum,
D. Stephenson, J. Neville, W. Maetzler, A. J. Espay, B. R. Bloem, T. Simuni, D. R.
Karlin, Metadata concepts for advancing the use of digital health technologies in clinical
research, Digital biomarkers 3 (2020) 116–132.
[26] U.S. National Institutes of Health, NIH strategic plan for data science, https://datascience.</p>
      <p>nih.gov/nih-strategic-plan-data-science, 2018. Online; accessed April-2024.
[27] M. Mustra, K. Delac, M. Grgic, Overview of the DICOM standard, in: 2008 50th International</p>
      <p>Symposium ELMAR, volume 1, IEEE, 2008, pp. 39–44.
[28] I. G. Cohen, M. M. Mello, HIPAA and protecting health information in the 21st century,</p>
      <p>Jama 320 (2018) 231–232.
[29] C. J. Hoofnagle, B. Van Der Sloot, F. Z. Borgesius, The European Union general data
protection regulation: what it is and what it means, Information &amp; Communications
Technology Law 28 (2019) 65–98.
[30] I. G. Cohen, R. Amarasingham, A. Shah, B. Xie, B. Lo, The legal and ethical concerns
that arise from using complex predictive analytics in health care, Health afairs 33 (2014)
1139–1147.
[31] A. Rajkomar, M. Hardt, M. D. Howell, G. Corrado, M. H. Chin, Ensuring fairness in machine
learning to advance health equity, Annals of internal medicine 169 (2018) 866–872.
[32] N. Norori, Q. Hu, F. M. Aellen, F. D. Faraci, A. Tzovara, Addressing bias in big data and AI
for health care: A call for open science, Patterns 2 (2021).
[33] C. Criscuolo, T. Dolci, M. Salnitri, Towards assessing data bias in clinical trials, in: VLDB
Workshop on Data Management and Analytics for Medicine and Healthcare, Springer,
2022, pp. 57–74.
[34] J. R. Marcelin, D. S. Siraj, R. Victor, S. Kotadia, Y. A. Maldonado, The impact of unconscious
bias in healthcare: how to recognize and mitigate it, The Journal of infectious diseases 220
(2019) S62–S73.
[35] C. FitzGerald, S. Hurst, Implicit bias in healthcare professionals: a systematic review, BMC
medical ethics 18 (2017) 1–18.
[36] J. Odgaard-Jensen, G. E. Vist, A. Timmer, R. Kunz, E. A. Akl, H. Schünemann, M. Briel, A. J.</p>
      <p>Nordmann, S. Pregno, A. D. Oxman, Randomisation to protect against selection bias in
healthcare trials, Cochrane database of systematic reviews (2011).
[37] O. Bodenreider, The unified medical language system (UMLS): integrating biomedical
terminology, Nucleic acids research 32 (2004) D267–D270.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <article-title>Managing data lakes in big data era: What's a data lake and why has it became popular in data management ecosystem</article-title>
          ,
          <source>in: 2015 IEEE International Conference on Cyber Technology in Automation, Control, and Intelligent Systems (CYBER)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>820</fpage>
          -
          <lpage>824</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Piantella</surname>
          </string-name>
          ,
          <article-title>A research on data lakes and their integration challenges</article-title>
          ,
          <source>in: Proceedings of the 30th Italian Symposium on Advanced Database Systems</source>
          , SEBD, volume
          <volume>3194</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>616</fpage>
          -
          <lpage>621</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Koutras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jarke</surname>
          </string-name>
          ,
          <article-title>Data lakes: A survey of functions and systems</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>D. R.-J. G.-J. Rydning</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Reinsel</surname>
            ,
            <given-names>J. Gantz,</given-names>
          </string-name>
          <article-title>The digitization of the world from edge to core</article-title>
          ,
          <source>Framingham: International Data Corporation</source>
          <volume>16</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>U. FDA</surname>
          </string-name>
          ,
          <article-title>Framework for FDA's real-world evidence program, Silver Spring, MD: US Department of Health and Human Services Food</article-title>
          and Drug
          <string-name>
            <surname>Administration</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kondylakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Koumakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tsiknakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Marias</surname>
          </string-name>
          ,
          <article-title>Implementing a data management infrastructure for big healthcare data</article-title>
          ,
          <source>in: 2018 IEEE EMBS International Conference on Biomedical &amp; Health Informatics (BHI)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>361</fpage>
          -
          <lpage>364</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ravat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Metadata management for data lakes</article-title>
          ,
          <source>in: New Trends in Databases and Information Systems: ADBIS 2019 Short Papers</source>
          ,
          <string-name>
            <surname>Workshops</surname>
            <given-names>BBIGAP</given-names>
          </string-name>
          , QAUCA, SemBDM, SIMPDA, M2P, MADEISD, and Doctoral Consortium, Bled, Slovenia, September 8-
          <issue>11</issue>
          ,
          <year>2019</year>
          , Proceedings 23, Springer,
          <year>2019</year>
          , pp.
          <fpage>37</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Canakoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernasconi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Colombo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Masseroli</surname>
          </string-name>
          , S. Ceri,
          <article-title>GenoSurf: metadata driven semantic search system for integrated genomic datasets</article-title>
          ,
          <year>Database 2019</year>
          (
          <year>2019</year>
          )
          <fpage>132</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Conte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Mason</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Halmagyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Neuhauser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mosaku</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yordanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chatzipli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Begley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Krupke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Parkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Meehan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Bult</surname>
          </string-name>
          , PDX Finder:
          <article-title>A portal for patient-derived tumor xenograft model discovery</article-title>
          ,
          <source>Nucleic acids research</source>
          <volume>47</volume>
          (
          <year>2019</year>
          )
          <fpage>D1073</fpage>
          -
          <lpage>D1079</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. H.</given-names>
            <surname>Dolin</surname>
          </string-name>
          , L. Alschuler,
          <article-title>Approaching semantic interoperability in health level seven</article-title>
          ,
          <source>Journal of the American Medical Informatics Association</source>
          <volume>18</volume>
          (
          <year>2011</year>
          )
          <fpage>99</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ehsan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A. M.</given-names>
            <surname>Abuhaliqa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Catal</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Mishra, RESTful API testing methodologies: Rationale, challenges</article-title>
          , and solution directions,
          <source>Applied Sciences</source>
          <volume>12</volume>
          (
          <year>2022</year>
          )
          <fpage>4369</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brase</surname>
          </string-name>
          ,
          <article-title>DataCite: a global registration agency for research data, in: 2009 fourth international conference on cooperation and promotion of information resources in science and technology</article-title>
          , IEEE,
          <year>2009</year>
          , pp.
          <fpage>257</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Scott</surname>
          </string-name>
          , R. Worden,
          <article-title>Semantic mapping to simplify deployment of HL7 v3 clinical document architecture</article-title>
          ,
          <source>Journal of biomedical informatics 45</source>
          (
          <year>2012</year>
          )
          <fpage>697</fpage>
          -
          <lpage>702</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>