<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>May</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Derek Corrigan</string-name>
          <email>derekcorrigan@rcsi.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qurratal Ain Fatimah</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Syeda Mah-e-Fatima</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ali Hasnain</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FutureNeuro SFI Research Centre, Royal College of Surgeons in Ireland</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Pharmacy and Biomedical Sciences, Royal College of Surgeons in Ireland</institution>
          ,
          <addr-line>Dublin</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University Hospital Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>29</volume>
      <issue>2022</issue>
      <abstract>
        <p>This paper discusses the importance as well as the practical issues that can be expected while working towards the FAIRification of Electronic Health Records (EHRs). Electronic health records are medical records of a patient that document their illness history as well as other relevant personal information. Patients interact with health care professionals in a multitude of capacities and thus, there is a stream of medical data available at any given time. So far, electronic health records have not been utilized to their full potential, which in part is attributable to the many dimensions the health care profession is split into, and the lack of interconnectivity between disciplines. Having patient data either in one source or alternatively interconnected fulfilling the principles of FAIR data, i.e. findable, accessible, interoperable and reusable, will be of immense benefit to both the healthcare system as well as for the patient. Streamlined FAIR compliance electronic health records will mean all the pertinent data should be linked or connected for better findability and accessibility to make it reusable and interoperable. In this position paper we look at the principles of FAIR data, identify priorities in relation to making data reusable, as well as recognizing points of action and complications in delivering FAIRified EHRs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>An Electronic Health Record (EHR) is a longitudinal collection of electronic health information
about a patient. The ubiquity of EHR systems has transformed the health care landscape over
several decades. Yet, even as improved patient care and cost savings have begun to emerge,
significant usability impediments have been documented. These include the duplication and
fragmentation of EHRs that hampers interoperability eforts and impact on patients’ safety.
Facilitating better access to and sharing of structured and unstructured health data is crucial to
ensure greater accessibility, availability, and afordability of healthcare. It will also stimulate
innovation in health and care for better treatment and outcomes and foster innovative solutions
that make use of digital technologies.</p>
      <p>SeWebMeDA-2022: 5th International Workshop on Semantic Web solutions for large-scale biomedical data analytics,
https://www.rcsi.com/people/profile/alihasnain (A. Hasnain)
CEUR
Workshop
Proceedings</p>
      <p>© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>The collection, access, use and re-use of health data in healthcare poses specific challenges in
ifnding the right balance between measures that facilitate data sharing while preserving the
interests and rights of individuals, including their personal data protection.</p>
      <p>
        It is important to highlight that unstructured data represent the 80% of data contained in
EHRs [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Unstructured content is the text of anamnestic notes, physical examination sheets,
medical and nurse diaries, surgical forms, specialist test reports, discharge letters written during
the patient’s stay, and any comment extending the standard patient-reported outcome measures
collected during the follow-up phase. This heterogeneous textual content contains relevant,
detailed and nuanced information about the illness trajectory and care processes undertaken by
and upon the patients. This makes the challenge to automatically extract accurate information
from narrative notes worthwhile.
      </p>
      <p>Digital health technology and data pose a critical opportunity and challenge for researching,
practicing, and experiencing healthcare, promoting public health policies, and changing the
way in which medicine is understood. However, digital health data is currently fragmented and
dispersed in diferent and non-homogeneous repositories.</p>
      <p>
        Standardization and common structures for facilitating interoperability and re-use of these
data for clinical practice and for research and innovation is a challenge, especially when it
comes to exploiting unstructured information in a secure way and making this information
interoperable [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Personal data has also emerged in the economic literature as the world economy’s new asset
class. These data ofer a cost-efective and technically feasible pathway to personalized medicine;
create wealth and value, attract private and public investment and provide a return to the whole
society by means of better and more eficient prevention and treatment management protocols.
However, data assetization and exploitation (including its analysis and synthesis) needs further
evidence, developments and research into capitalisation, techno-scientific applications and
socio-cultural and socio-technical identification of resources and possibilities, including legal
and structural facilitators and barriers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Exchange of medical information has been associated with benefits for patients, health
delivery systems and the economy. In fact, real-time health information systems, integrating all
relevant information on a patient(s) and their healthcare process, can substantially improve
coordinated care, patient safety, quality, and eficiency.</p>
      <p>For this reason, it is critical to improve EHRs interoperability to achieve their full potential,
that is, the ability of health information systems to work together within and across
organizational boundaries to advance the efective delivery of healthcare. While electronic health records
(EHRs) are of immense value in healthcare and medicine their potential has not been fully
exploited due to fragmentation of the healthcare sector. This is due to the lack of interconnectivity
and interoperability between diferent EHR systems.</p>
      <p>
        To this purpose, FAIR data principles [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] provide clear guidelines to increase the findability,
interoperability, accessibility and re-usability of data. The sensitive nature of data and information
available in EHRs brings challenges to FAIRify it. An important point made in the EU FAIR data
report [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (that was published before the COVID pandemic) was the value of FAIR data - making
data more widely available and reusable to support response to public health emergencies in a
timely manner will be a valuable asset. This allows easier access and sharing of data across
individual nations and beyond the boundaries across the European Union. The need to unlock
the data and the value of national EHR systems across the European Union as a resource for
the promotion of population health poses more pressure in the current circumstances of a
pandemic.
      </p>
      <p>
        An argument made is that the research community has not always suficiently acknowledged
the inherent value related to the production of analytical datasets to support research initiatives.
One of the recommendations from the EU commission report (step 3, recommendations 12
and 13) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is that development of FAIR data ecosystems allows for acknowledgement of the
explicit value of research data that can be rewarded through the development of incentives and
new metrics that measure data reuse. This can then incentivize and reward those who make it
actively do it well through “good data stewardship”.
      </p>
      <p>In this position paper, we advocate the importance of FAIRification led interconnectivity of
EHRs by presenting two motivational scenarios that arise on a daily basis in clinical settings
e.g. accidents and emergency units that can be supported using well-connected interoperable
EHRs at least at the national level.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Electronic Health Records</title>
      <p>Electronic Health Records (EHRs) are primarily known as a digital collection of medical
information about an individual patient that includes but is not limited to the information about
a patient’s health history, diagnoses, prescription medicines, diagnostic and monitoring tests,
known allergies and immunizations. It can also include healthcare encounters and referrals to
specialists and ambulatory/outpatient encounters in hospitals.</p>
      <p>Exchange of medical information has been shown to be associated with higher quality and
safer care for patients. It also has clear benefits for healthcare organizations- providers of care
in community, outpatient and hospital settings as well as health insurers and the health sector
more broadly. EHRs help healthcare providers better manage care for patients and provide
better health care by:
• Providing accurate, up-to-date, and complete information about patients at the point of
care at specific time
• Enabling quick access to patient records for more coordinated and eficient care
• Exploiting the potential of sharing electronic information with patients and other
clinicians in a secure manner
• Helping healthcare providers more efectively diagnose patients, reduce medical or human
errors while providing better and safer care
• Improving patient and health care provider interaction and communication, as well as
health care convenience
• Enabling safer, eficient and more appropriate prescribing of drugs</p>
      <p>The sensitive nature of healthcare data creates fragmented, siloed and mostly private
repositories of data by design. Although online portals have increased patients’ ability to view their
records, the structuring of this access is limited by design. A minority of systems allow direct
patient access to data and healthcare results. The degree to which data should be shared directly
with patients is still being debated within the healthcare community due to ever pressing
requirements of safety and security of such sensitive data in the digital space.</p>
      <p>On the other hand, holding and oversight of health data has largely been in the hands of
multinational technology organizations. This data collection, acquisition, storage and holding
of health data has evolved around the concept and motivation of the personal Electronic Health
Record data that supports collection of data relating to individuals’ exercise and lifestyle as
well as physiological measurements such as pulse rate, rhythm and variability. Initiatives have
emerged to develop these person-based EHRs, the most notable examples being Microsoft’s
HealthVault1 and Google Health Google Health,2. These technology innovations failed to
generate traction because of poor user engagement and concerns about security of personal
data. Most critically there were significant limitations in terms of integration and interoperability
with EHRs. The overlap between health data relating to diet, exercise and lifestyle activity with
more traditional medical data including morbidity conditions, medication codes, and healthcare
utilization (diagnostic testing, monitoring of disease, and contacts with health professionals
in primary and secondary care) has never been fully resolved. This may well be explained
by the confidential nature of medical data and the fact that prescription drugs require input
and understanding from health professionals (doctors and pharmacists) who are concerned
with prescribing, monitoring and dispensing of medicines needing specialist knowledge and
regulatory oversight.</p>
    </sec>
    <sec id="sec-4">
      <title>3. FAIR Data Overview</title>
      <p>
        In 2018 the European Commission expert working group on FAIR data produced a detailed
report about how to promote the concept of FAIR data to support more open science in the
European Union [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This detailed report sets out some of the motivations for promoting FAIR
data principles:
“It has long been recognised that it is not suficient simply to post data and
      </p>
      <p>
        other research-related materials onto the web and hope that the
motivation and skill of the potential user would be suficient to enable reuse.” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
      </p>
      <p>
        FAIR is a set of principles that collectively describe requirements to develop a FAIR ecosystem
supporting digital data reuse [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ]. The basic goal of FAIR is to promote wider reuse and
sustainability of digital datasets produced from research. This promotes better linkage and
connections between diferent research initiatives both within and across research domains and
disciplines. This in turn can accelerate discovery and increase the replicability of science. The
FAIR acronym stands for:
• Findable – datasets are uniquely identified and indexed in publicly available resources
that enable searchable discovery e.g., internet search engines
1Microsoft Cloud for Healthcare, https://www.microsoft.com/en-ie/industry/health/microsoft-cloud-for-healthcare?
rtc=1. (last accessed 14-03-2022)
2Google Health, https://health.google/. (last accessed 14-03-2022)
• Accessible – datasets are accessible based on defined authentication requirements in
publicly available repositories accessed using open protocols e.g., http
• Interoperable – datasets are described with metadata using open standards that
determine dataset content based on accepted controlled terminologies and vocabularies to
promote wider reuse
• Reusable – datasets are richly and accurately described by metadata that uses standards
most appropriate to the domain of knowledge and includes the provenance, licensing and
dataset content and appropriate use.
3.1. FAIR Digital Object
The basic building block to represent a data resource is the FAIR digital object. The object
consists of a link to the dataset itself along with a set of metadata that fully describes the
data according to the FAIR data principles. The metadata and the dataset are typically stored
separately to allow for authentication and access control (a description of the data may be
accessible but the actual data access may not if deemed sensitive).
      </p>
      <p>The digital object ensures that the FAIR resource is uniquely identified, semantically described
using open technical standards, and appropriately shared according to defined access controls
based on licensing information.</p>
      <p>
        An important consideration as highlighted in the EU report is the need for development of
distributed FAIR ecosystems by storing collections of FAIR digital objects across distributed
FAIR repositories that belong to searchable and uniquely defined organizations defined in
FAIR registries [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The goal is the publication of datasets in a FAIR ecosystem, with defined
licensing access and provenance that allow for long term sustainability and access to datasets
that eliminates the need to make multiple copies of that dataset across multiple sites. It is
envisioned that the datasets, once FAIRified are accessed from source:
      </p>
      <p>
        “Data Federations ofer a means to establish agreements between repositories or registries to
carry out certain tasks collaboratively and therefore will be essential to this distributed system.
Data will increasingly remain at various locations for reasons such as the expense of copying
data or because of legal or ethical restrictions. Distributed queries, managed by brokering
software, will be used to virtually integrate data. The need for such distributed analysis across
multiple data sets is one of the major drivers and use cases for FAIR data: it requires metadata
to find the data resources, protocols to access them, agreed specifications such that the data can
interoperate and rich provenance information so that the data can be reused with confidence”
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>An important distinction to be made is that FAIR data does not necessarily imply ‘Open data’.
As part of the metadata associated with FAIR datasets, it is envisioned that licensing will also
be described which may allow full access or restricted access to only the metadata descriptions
rather than the actual raw datasets themselves where it is considered that the data is sensitive or
should be restricted. This also implies that a FAIR infrastructure must allow for authentication
mechanisms that define how data can be shared across diferent FAIR data repositories, groups
and disciplines while respecting the EU legislation across borders as well (such as General Data
Protection Regulations). Whilst open data is considered desirable the EU envisions that data
should be ‘as open as possible, as closed as necessary’.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Motivating Scenarios</title>
      <p>Integration of diferent patient repositories as well as electronic health records add value to the
case that such integrations could make healthcare safer for healthcare professionals as well as
patients. Consideration of FAIR data principles can be a foundational step for the digitization
and integration of EHRs that are more findable, accessible, interoperable and reusable for better
accessing clinical scenarios. The following practical real world scenarios show the motivation
behind the FAIRification led integration of electronic health records.</p>
      <p>Motivating Scenario 1 Susy is a patient who is involved in a road trafic accident and
presents to a hospital to be seen by the doctor. Due to the nature of injuries sustained, the
hospital orders, among other tests, a CT scan of the brain. The patient has the CT scan but
is waiting many hours to be seen by the doctor. She eventually leaves without receiving the
results of the tests and presents to another hospital. Since no records are available and the
patient is unable to tell the doctor at the second hospital what tests she got in the first place,
she has another CT scan at the second hospital. The patient is eventually discharged but has
now had two CT scans of the head. The radiation exposure is significant. One in every 1800 CT
scans leads to one excess cancer. This patient has now had exposure at two facilities due to data
inaccessibility between the two facilities. If data is findable and accessible as proposed by the
FAIRify principles, better clinical decisions such as limiting the amount of harmful radiation
a patient is exposed to, can be made, and repetitive testing can be avoided. This will lead to
better patient outcomes as well as reducing cost of patient care.</p>
      <p>In a nutshell the accessibility of patient data could have been achieved through FAIRification
led digitization of EHRs</p>
      <p>Motivating Scenario 2 In a second scenario, a patient has been discharged from hospital
following a prolonged in-patient stay. Upon discharge, a discharge summary has been completed
by the hospital many days following discharge. The patient had multiple investigations in
the hospital including blood tests, electrocardiogram (ECG) and radiological investigations,
but none of the results are available to the GP who is the primary physician of the patient in
the community. Some of the doses of medications have been adjusted, and needs referrals to
multiple outpatient specialties for continuity of care, but as the GP is unaware due to a delay in
the discharge summary and recommendations getting to them, there is a delay ultimately in
the patient getting optimal care. If data is findable and accessible as proposed by the FAIRify
principles to the corresponding GP (obviously in a secure and closed environment), better
clinical decisions as well as timely diagnosis and recommended medication modifications would
have been possible by the GP. This will also lead to better patient outcomes as well as reducing
cost of patient care.</p>
      <p>Motivating scenario 3 FAIRifying EHR will play a significant role in integrating research
(interoperable data) along multiple platforms. If an interesting clinical scenario such as a rare
cancer or genetic condition is observed in facility A, for instance, with EHR being accessible
across a network of healthcare, it will be easy and feasible for another institution to not only
gain access to the information due to accessibility of data, but also compare the findings and
consolidate their own research with research from other sites. EHR can also then be used to
recruit for clinical studies, following trends across institutions.</p>
      <p>Accessibility of patient data available in hospital in-patient and its interoperability/
connectivity (in a protected, secured and closed environment) with the EHRs available at GP practices
could be supported through FAIRification led digitization of EHRs.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Describing the FAIR Data Ecosystem for EHRs</title>
      <p>
        The implementation of FAIR data principles are described in the EC report “Turning FAIR into
reality” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which identifies several priorities in relation to the promotion of the re-use and
sharing of health data across Europe that will be the points of action for delivering FAIR data
implementations. Key points are:
• Standardized and interoperable data selection
• Long-term data stewardship
• Accessibility (by both person and machine)
• Legal interoperability
• Timeliness of sharing
      </p>
      <p>
        Based on these priorities, any integrated healthcare data, research outputs, or results can be
subject to a FAIRification process [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>This process of FAIRifying the healthcare data, (more precisely Electronic Health Records)
and the implementation of technical components needed to support the process of FAIRification
are presented as follow but were originally described in the recommendations of the GO FAIR
initiative (Figure 1). The seven distinct points to consider as follows:
1. Creation of a FAIR provider registry to uniquely identify digital objects made available
by the data provider (e.g., healthcare datasets, publications)
2. Creation of FAIR digital objects (FDOs) from EHR system extracts
3. Creation of local FAIR data repositories (FDRs) for digital objects hosted by each data
provider
4. Deployment of FDOs to FDRs.
5. Definition of access policies and licenses to control data sharing and wider access to</p>
      <p>FDRs &amp; FDOs
6. Assessment of FDRs based on FAIR maturity models, certification standards and FAIR
metrics.
7. Development and deployment of supporting tools and services to allow search, browsing,
and controlled access to FDOs between providers and third parties.</p>
      <p>
        It is clear from this description of the FAIRification process that the technical implementation
of the FAIR concept in its entirety and at scale requires several separate software components
that collectively work together to deliver FAIRificiation of data in practice. These important
sets of individual but integrated software components are required to implement FAIR data
objects and FAIR data repositories. We propose three most important set of tools namely 1) FAIR
Data Object Annotation Tool, 2) FAIR Data Object Manager and 3) Local FAIR Data Repository
Manager as described below:
• FDO Annotation Tool: A software that will allow health data providers to describe
their FDOs using metadata tags in a standardized way using accepted ontologies such
as use of appropriate clinical terminologies to provide standardized semantic
interoperability and coding for FAIR data (most notable examples but not limited to these are
ICD [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], LOINC [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], SNOMED CT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], etc.). Other proposed standard ontologies include
Human Phenotype Ontology [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Online Mammalian Inheritance in Man (OMIM) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
ORPHANET [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], providing for cross disciplinary phenotypical and genetic interpretation
and understanding of how to use, interpret and link to the FAIR data for research purposes.
It is also proposed that the Metadata annotation should be based on the structure of the
TRIPOD guidelines [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for predictive model development and validation to fully describe
the research datasets according to recognized best practice.
• FDO Manager: A software component for creation, packaging, storage, and curation
of FDOs, along with their unique identifiers, dataset description, provenance metadata,
licensing and Data Management Plans (DMPs) to local FAIR Data Repositories. The HL7
FHIR [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for FAIR implementation standard is an example of standards to support the
implementation of FAIR repositories.
• Local FDR Manager: A software tool for managing local FDRs at each health data
provider. This will allow for managing security, governance, and availability of local
FDRs to participants and third parties. This will allow for searching and creation of
FAIR repositories and making FAIR digital objects available and searchable as collections
of publically accessible REST API endpoints. The HL7 FHIR for FAIR implementation
standards will be used and extended as needed to implement FAIR repositories.
      </p>
      <p>The pictorial depiction of aforementioned tools, software components and the way these
individual components are connected can be seen in Figure 1 below.</p>
      <p>It is evident that the delivery of a FAIR data ecosystem that supports sharing of EHR data
is therefore not simply a case of adopting health interoperability standards. It must also be
supported by a broader ecosystem that uniquely identifies the EHR resources available, makes
them searchable, discoverable, and most importantly controls access to only those third parties
who are authorized to use those resources.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Discussion (Challenges for FAIRifying EHRs)</title>
      <p>In the multiple clinical scenarios we have mentioned in this paper, it can be deduced that
EHRs, if accessible and interoperable across healthcare modalities, can benefit immensely any
given healthcare system in terms of cost reduction by avoiding multiple testing and multiple
presentations, as well as by enhancing patient safety. Data collected from example studies of
rare genetic disorders or clinical presentations can be consolidated from more than one location
and be built upon for clinical studies with more ease.</p>
      <p>
        The original focus for developing FAIR data standards and a broader FAIR ecosystem was
cantered around the sharing and transparency of research-related datasets[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. By their nature,
research datasets contain finite datasets that are needed to answer specific research questions
guided by the research questions that the study is investigating. Research datasets are also
largely expected to be static in nature, unless they are subsequently revised or corrected due to
inaccuracies found in the original research. These datasets are captured at a point in time during
the conduct of the research and may be reported to support the publication of the research
study results. The concept of retaining the integrity of data in research is fundamental and
consistent with FAIR data principles where data is published and has a unique resource identifier
associated with that static dataset forever, supporting the publication of results which cannot
subsequently be changed.
      </p>
      <p>On the contrary, the nature of EHR data is fundamentally diferent. EHR data typically have
very wide coverage of collected data. This data may potentially be captured to support decision
making across a much wider spectrum of clinical care including phenotypic data, diagnostics,
therapeutics, imaging, and genetics. In any discussion about FAIRifying EHR data we therefore
need to be clear about whether it is feasible to FAIRify an entire complex EHR in its entirety, or
whether several defined subsets of that EHR data are being constructed and FAIRified separately.
The data complexity of a typical EHR system suggests that a more manageable approach would
be to extract several data subsets from an entire EHR that can be FAIRified to support the defined
needs of research or real-world-evidence creation. In the context of FAIR data, we potentially
therefore have multiple FAIR datasets created from a single data source that would appear to
have some sort of relationship or linkage between them.</p>
      <p>EHR data collected in frontline clinical care is also highly dynamic with changes to the
underlying data occurring frequently during a single day. The creation of dataset extracts
as part of a FAIRification process is therefore only reflecting a point-in-time snapshot of a
defined subset of the overall EHR data. If multiple FAIR datasets are required to reflect the data
complexity of the EHR as mentioned previously then it only makes sense to create them all at
the same point in time to maintain integrity between the diferent related sets of data.</p>
      <p>
        A final distinction needs to be made between research datasets and EHR systems. The data
captured as part of the conduct of research studies is typically subject to ethical approval and
patient consent to gather such data. The ethical constraints under which EHR data that is
captured for the primary purposes of routine clinical care may not be so explicitly defined.
Under General Data Protection Regulation (GPDR) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] guidance, this raises questions about
secondary processing of such data and the legal basis for onward data sharing to support
the conduct of research or real-world-evidence generation. In the absence of explicit patient
consent for onwards data sharing for explicitly informed research purposes, this suggests that
anonymization of EHR data is a prerequisite for FAIRification of EHR data which incurs an
additional data processing overhead.
      </p>
      <p>
        These distinctions in complexity, dynamic nature, and ethical approval relating to EHR data
suggest that the FAIR data concept was not originally defined to support the concept of sharing
of data from frontline clinical care systems in its current form. The concept of ‘Findability’ of
data becomes problematic in this context where each data resource should have a uniquely
defined and unchanging identifier that defines it as a uniquely identifiable and shared data
resource. It is not clear if multiple EHR extracts should have a single identifier and whether
that identifier should dynamically change with each snapshot of data that is extracted from the
EHR system [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        One potential solution, according to Stein et al, [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], that is being used is to associate a unique
identifier of an associated research publication based on the underlying FAIR data source and
to quote that as the findable resource 3. This is a workaround rather than a solution to this
problem. A proper solution would suggest a diferent type of resource identifier can be used
for these more dynamic and complex datasources that has a core unique identifier that can
be adjusted with some sort of time stamp with a common root that indicates where diferent
datasets are related from the same core underlying datasource.
      </p>
    </sec>
    <sec id="sec-8">
      <title>7. Future Directions</title>
      <p>
        The FAIR data concept is still evolving and is now being more universally adopted in the context
of sharing data generated from other areas beyond static research datasets. This has been noted
as being problematic while indicating a gap in the context of applying FAIR data principles to
more dynamic and larger (perhaps linked) datasets that collectively expose EHR data for the
purposes of real world evidence more generally [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. However, this relative limitation of the FAIR
data concept and the eagerness of organizations to demonstrate ‘open data’ credentials has led
to difering interpretations of the degree to which there is adherence to FAIR principles, which
seems to vary across diferent data repositories. This suggests that a mutual understanding of
minimum requirements is needed, as this has become an ever-pressing need [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
34DN Data Portal https://data.4dnucleome.org/help/user-guide/faq. (last accessed 10-03-2022)
      </p>
      <p>
        Authors also believe that this will emerge and become clearer as expert groups from the EU
have been established to measure adherence to and examine alternative metrics some of which
may be appropriate for measuring compliance with FAIR data principles [
        <xref ref-type="bibr" rid="ref19 ref20 ref7">7, 19, 20</xref>
        ].
      </p>
      <p>
        The expert group stated: “A major additional challenge in the data domain is the adoption of
a new set of metrics to assess FAIRness, i.e., compliance with the FAIR principles. We propose
the following as a basic minimum standard: discovery metadata, persistent identifiers and
access to the data or metadata”. It will be important to standardize FAIR metrics globally and to
coordinate initiatives and some are under way to develop FAIR maturity models or assessment
tools [
        <xref ref-type="bibr" rid="ref20 ref21 ref7">7, 21, 20</xref>
        ].
      </p>
    </sec>
    <sec id="sec-9">
      <title>8. Conclusion</title>
      <p>FAIR data is a set of principles that provides the guidelines and recommendation for data being
ifndable, accessible, interoperable, and reusable. Applying FAIR data principles to electronic
health records is an evolving process. The complex and dynamic nature of medical data, as
well as the ethical complexities that do not critically apply to records in any other field, has
demonstrated that more dimensions and associated metrics will be needed to be applied to
electronic records if they will be interoperable and accessible. Nevertheless, taking the timely
decision the EU has already set up committees to address the feasibility, possibilities and
complications of the same which is a judicious step in the right direction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.-J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <article-title>Managing unstructured big data in healthcare system</article-title>
          ,
          <source>Healthcare informatics research 25</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hersh</surname>
          </string-name>
          , H. Liu,
          <article-title>On mapping textual queries to a common data model</article-title>
          ,
          <source>in: 2017 IEEE International Conference on Healthcare Informatics (ICHI)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vezyridis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Timmons</surname>
          </string-name>
          ,
          <article-title>E-infrastructures and the divergent assetization of public health data: Expectations, uncertainties, and asymmetries</article-title>
          ,
          <source>Social Studies of Science</source>
          <volume>51</volume>
          (
          <year>2021</year>
          )
          <fpage>606</fpage>
          -
          <lpage>627</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Stall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yarmey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cutcher-Gershenfeld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hanson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lehnert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nosek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Parsons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wyborn</surname>
          </string-name>
          ,
          <article-title>Make scientific data fair</article-title>
          ,
          <source>Nature</source>
          <volume>570</volume>
          (
          <year>2019</year>
          )
          <fpage>27</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Collins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Genova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Harrower</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hodson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Laaksonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mietchen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Petrauskaitė</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wittenburg</surname>
          </string-name>
          ,
          <article-title>Turning fair into reality: Final report and action plan from the european commission expert group on fair data (</article-title>
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Wilkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Aalbersberg</surname>
          </string-name>
          , G. Appleton,
          <string-name>
            <given-names>M.</given-names>
            <surname>Axton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Blomberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-W.</given-names>
            <surname>Boiten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B. da Silva</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Bourne</surname>
          </string-name>
          , et al.,
          <article-title>The fair guiding principles for scientific data management and stewardship</article-title>
          ,
          <source>Scientific data 3</source>
          (
          <year>2016</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hasnain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <article-title>Assessing fair data principles against the 5-star open data principles, in: The Semantic Web: ESWC 2018 Satellite Events: ESWC 2018 Satellite Events</article-title>
          , Heraklion, Crete, Greece, June 3-7,
          <year>2018</year>
          ,
          <source>Revised Selected Papers 15</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>469</fpage>
          -
          <lpage>477</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C. J.</given-names>
            <surname>McDonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Huf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Suico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Leavelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Forrey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mercer</surname>
          </string-name>
          , G. DeMoor, J.
          <string-name>
            <surname>Hook</surname>
          </string-name>
          , et al.,
          <article-title>Loinc, a universal standard for identifying laboratory observations: a 5-year update</article-title>
          ,
          <source>Clinical chemistry 49</source>
          (
          <year>2003</year>
          )
          <fpage>624</fpage>
          -
          <lpage>633</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Donnelly</surname>
          </string-name>
          , et al.,
          <article-title>Snomed-ct: The advanced terminology and coding system for ehealth</article-title>
          ,
          <source>Studies in health technology and informatics 121</source>
          (
          <year>2006</year>
          )
          <fpage>279</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Köhler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Vasilevsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Engelstad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Foster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McMurry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aymé</surname>
          </string-name>
          , G. Baynam,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Bello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. F.</given-names>
            <surname>Boerkoel</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Boycott</surname>
          </string-name>
          , et al.,
          <source>The human phenotype ontology in 2017, Nucleic acids research</source>
          <volume>45</volume>
          (
          <year>2017</year>
          )
          <fpage>D865</fpage>
          -
          <lpage>D876</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Scott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Amberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brylawski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. A.</given-names>
            <surname>McKusick</surname>
          </string-name>
          ,
          <article-title>Omim: Online mendelian inheritance in man, Bioinformatics: Databases and systems (</article-title>
          <year>1999</year>
          )
          <fpage>77</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Weinreich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mangon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sikkens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Teeuw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornel</surname>
          </string-name>
          ,
          <article-title>Orphanet: a european database for rare diseases</article-title>
          ,
          <source>Nederlands tijdschrift voor geneeskunde 152</source>
          (
          <year>2008</year>
          )
          <fpage>518</fpage>
          -
          <lpage>519</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Collins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Reitsma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Moons</surname>
          </string-name>
          ,
          <article-title>Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod) the tripod statement</article-title>
          ,
          <source>Circulation</source>
          <volume>131</volume>
          (
          <year>2015</year>
          )
          <fpage>211</fpage>
          -
          <lpage>219</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sartipi</surname>
          </string-name>
          ,
          <article-title>Hl7 fhir: An agile and restful approach to healthcare information exchange</article-title>
          ,
          <source>in: Proceedings of the 26th IEEE international symposium on computer-based medical systems</source>
          , IEEE,
          <year>2013</year>
          , pp.
          <fpage>326</fpage>
          -
          <lpage>331</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Regulation</surname>
          </string-name>
          ,
          <article-title>General data protection regulation</article-title>
          ,
          <source>Intouch</source>
          <volume>25</volume>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Löbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Matthies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stäubert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Meineke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <article-title>Problems in fairifying medical datasets</article-title>
          .,
          <source>in: MIE</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>392</fpage>
          -
          <lpage>396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ceol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Aloy</surname>
          </string-name>
          ,
          <article-title>3did: identification and classification of domain-based interactions of known three-dimensional structure</article-title>
          ,
          <source>Nucleic acids research</source>
          <volume>39</volume>
          (
          <year>2010</year>
          )
          <fpage>D718</fpage>
          -
          <lpage>D723</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dunning</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. De Smaele</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Böhmer</surname>
          </string-name>
          ,
          <article-title>Are the fair data principles fair?</article-title>
          ,
          <source>International Journal of digital curation 12</source>
          (
          <year>1970</year>
          )
          <fpage>177</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wilsdon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bar-Ilan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Frodeman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Lex</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wouters</surname>
          </string-name>
          ,
          <article-title>Next-generation metrics: Responsible metrics and evaluation for open science (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jacobsen</surname>
          </string-name>
          , R. de Miranda Azevedo,
          <string-name>
            <given-names>N.</given-names>
            <surname>Juty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Batista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Coles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cornet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Courtot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Crosas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. T.</given-names>
            <surname>Evelo</surname>
          </string-name>
          , et al.,
          <source>Fair principles: interpretations and implementation considerations</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>N.</given-names>
            <surname>Krans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ammar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nymark</surname>
          </string-name>
          , E. Willighagen,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quik</surname>
          </string-name>
          ,
          <article-title>Fair assessment tools: evaluating use and performance</article-title>
          ,
          <source>NanoImpact</source>
          <volume>27</volume>
          (
          <year>2022</year>
          )
          <fpage>100402</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>