<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Syeda Amna Sohail[</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Normative and Empirical Evaluation of Privacy Utility Trade-o in Healthcare</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Syeda Amna Sohail</string-name>
          <email>s.a.sohail@utwente.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Data Management and Biometrics (DMB), Faculty of Electrical Engineering</institution>
          ,
          <addr-line>Mathematics and Computer Science (EEMCS)</addr-line>
          ,
          <institution>University of Twente</institution>
          ,
          <addr-line>Enschede</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>0000</year>
      </pub-date>
      <volume>0001</volume>
      <fpage>11</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>Post-GDPR, the public/private (healthcare) enterprises, while performing (sensitive) Big Data Analytics (BDA), encounter the dilemma of abiding by the privacy regulations on one hand and extracting maximum value from (healthcare) metadata on the other. Concerning this, one of the major issues is the Privacy Utility trade-o (PUT). The PUT a ects each phase including (healthcare) metadata collection, formulation, storage, and resharing amongst (healthcare) enterprises. So far in healthcare, PUT concerning issues are identi ed and resolved in a remote, disintegrated manner. It's high time to resolve the issue by taking a holistic approach. This Ph.D. research work strives to achieve the same with normative (should be) and empirical (as-is) evaluation of PUT in Dutch care metadata share landscape. For clarity, the problem area is segregated into four fundamental dimensions. For each dimension, empirical evaluation is performed using Process Mining (discovery/conformance checking) techniques on real-world healthcare eventlog(s). Based on data analytics, the conceptual modeling frameworks are formulated using e3 value modeling or/and REA ontologies. For normative evaluation, two alternative approaches; the `Content Analysis', to formulate the conceptual modeling framework(s) and `BPMN text extraction', for documents `Rule Mining' for drawing the respective business model(s), are used. Later, the (in- eld) IT expert(s) further evaluates the proposed conceptual model(s). The aim is to evaluate the technical (IS-based privacy-preserving tools and techniques) and respective organizational (access governance, data ownership) measures of Dutch healthcare providers. The research work will (ultimately) contribute standardized conceptual modeling framework(s) with technical and respective organizational measures to e ciently cope with the PUT in handling sensitive (healthcare) metadata.</p>
      </abstract>
      <kwd-group>
        <kwd>privacy utility tradeo • conceptual modeling • process mining • healthcare</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Information Technology (IT) is quintessential in how contemporary research and
industrial undertakings proceed and execute [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. IT (computers, software apps,
and telecommunication) together with the business process modeling (and
evaluation) created Information Technology Engineering (ITE) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. ITE is an
amalgamation of business processes, techniques, and systems that improve business
proceedings in better achieving business goals. Because of ITE, public/private
enterprises formulate, store and share valuable metadata (i.e. data/information about
other (big) data) for Big Data Analytics (BDA). Big Data Analytics (BDA) is
(meta) data evaluation for valuable information gain. The BDA of metadata
facilitates the extraction of insights and actionable decisions in achieving business
goals in a cost and time-e cient manner [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In healthcare, the BDA improves
the care business models, foresees the long/short term treatment outcomes, and
does patient and disease centric strati cation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Post Covid-19, the BDA is
essential in aiding infectious control measures and respective policies. However,
the use of BDA predominantly relies upon metadata sharing and posits some
serious ethical concerns including privacy preservation of patient's personally
identi able information [
        <xref ref-type="bibr" rid="ref5 ref6 ref8">5, 6, 8</xref>
        ] within and across healthcare sub-domains.
      </p>
      <p>
        Privacy comprises the autonomous decision-making and direct/indirect
control over personal information [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The privacy concerns do not imply the lack
of trust in BDA, rather it demands responsible and fair metadata sharing [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ].
Unfortunately, the pace of privacy-preserving tools' (and techniques) formulation
(and implementation) lag far behind in comparison to the use of BDA across
domains especially in healthcare [
        <xref ref-type="bibr" rid="ref13 ref8">8, 13</xref>
        ]. Caregivers, for patients' e ective/e cient
clinical care, are bound to share the un-anonymized/pseudonymized patients'
metadata with other internally and externally located counterparts such as labs
and pharmacies [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Simultaneously, the care providers are obliged to ful ll
the privacy legislature/regulations [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] to avoid paying hundreds of thousands
of Euros as compensation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. For example, Haga hospital and Menzis
insurance company had to compensate for privacy lapses with hefty amounts [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
Privacy-Utility-Tradeo (PUT) is the performance impairment of data
analytics in ascertaining data privacy [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The issue demands quick,
comprehensive, and e cient technical/organizational solutions, to identify and evaluate
the data utility-prone privacy-preserving measures for the care provider's
Information Systems (ISs). Such measures will facilitate the healthcare providers to
avoid paying hefty compensations and in turn infamous reputation.
      </p>
      <p>Section 1 (introduction) comprises introduction of the problem domain and
two subsections namely the related work and research objectives and research
questions where we highlight the current state of the art and this research work's
objectives and questions, Section 2 comprises the standardized data analytics
approach and research methodology for all four dimensions. In Section 3, we
present the current results employing the two currently published papers
and an under-review paper in the rst year of Ph.D. research work. Section
4 highlights the threats to the validity of this research work, Section 5 gives a
detailed description of the dimension-wise contribution and the uniqueness of this
research work. Section 6 includes a conclusion, acknowledgments and is followed
by references.
1.1</p>
      <sec id="sec-1-1">
        <title>Related Work</title>
        <p>
          In the contemporary world of the information economy, the BDA is essential
for (private/public) enterprises to shape their business goals and information
system engineering with privacy by design measures [
          <xref ref-type="bibr" rid="ref25 ref6">6, 25</xref>
          ]. To ensure the
responsible Business Information System Engineering (BISE) the main concern,
amongst others, is the Privacy Utility Tradeo (PUT) [
          <xref ref-type="bibr" rid="ref22 ref6">6, 22</xref>
          ]. In healthcare, the
PUT raises graver repercussions because of the involvement of (highly)
sensitive personally identi able data on one hand and the e ciency and e ectiveness
of healthcare performance on the other [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. In this regard, research studies (in
healthcare) remotely focus on privacy and IoT [
          <xref ref-type="bibr" rid="ref12 ref16">12,16</xref>
          ], privacy in AI and machine
learning [
          <xref ref-type="bibr" rid="ref13 ref19">13, 19</xref>
          ], privacy in Process Mining (focussing on third party process
analytics) [
          <xref ref-type="bibr" rid="ref17 ref21 ref27">17, 21, 27</xref>
          ] by not paying much attention to the overall context of the
issue. Similarly, the data pipeline (from data collection to valuable insights) and
provenance records (pipeline description) is often ignored while sharing of the
data is emphasized concerning PUT [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>
          It is important here to realize that PUT concerning issues of (healthcare)
BDA is both technical and organization-based. Thus, the issue requires an
evaluation of (care provider's) integrated techniques, processes, and systems (i.e.
ITE) in (care) metadata share landscape. Privacy by design only provides for
the technical solutions of real-world problems [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] but the people performing the
tasks behind their computer screens are the quintessential source of issue
resolving. People who are interwoven in the organization's architectural setup and are
accountable to the organization. Besides, to identify, 'what (i.e. techniques) is
happening in the BISE', it is integral to identify and evaluate that how (i.e.
business processes) it is happening in that fashion? Process Mining with the event
logs gives us a glimpse of the same using the datasets extracted directly from an
organization's IS. Moreover, to check the business system's and processes'
compliance on both technical and organizational grounds, the normative (should be)
evaluation is done. Normative evaluation is done using two alternative
methodologies namely, 'Content Analysis' and organization's (o cial) documents' rule
mining using 'BPMN text extraction'. For simpli cation, the ndings are
represented with conceptual modeling frameworks using REA and e3 value modeling
ontologies. The ontologies are further evaluated by (in- eld: working in the same
domain) IT experts.
1.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Research Objectives and Research Questions</title>
        <p>
          The method of the research work aims to answer dichotomous knowledge
questions (Design Science Methodology [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]), that include Analytical research work,
using Process Mining tools/algorithms on real-world healthcare event log(s) (i.e.
data set(s) that are extracted directly from healthcare provider's Information
Systems) for the empirical (as-is state of a airs) evaluation of PUT in Dutch care
        </p>
        <p>Access
governance
evaluation</p>
        <p>3. Data
Ownership
evaluation</p>
        <p>Data Utility
indicators and
quantification
evaluation
1. Privacy
Preservation (inter and
intra organizational) in
care metadata sharing)</p>
        <p>FAIR/FACT
based data?
evaluation
2. Quantification
of Data Value
Privacy Utility Tradeoff
(PUT) in Healthcare</p>
        <p>Data Quality
indicators and
quantification
evaluation</p>
        <p>Data Ecosystem
Integration etc
measures i.e. from
logistics domain
4. Co-operate to Compete:
other domains' suitable
privacy preserving,
utilityprone tools and techniques</p>
        <p>Privacy Indicators
and quantification methods
evaluation</p>
        <p>Technical Solutions
privacy preserving, utility prone tools and</p>
        <p>techniques which are pro
legislative/regulatory prerequisites
metadata share landscape. And Exploratory research work, using either
'Content Analysis' or the 'BPMN text extraction' for normative (should-be state of
a airs) evaluation. The goal is the normative and empirical evaluation of PUT
concerning tools/techniques in the care provider's technical and organizational
metadata sharing set-up.</p>
        <p>For clarity the research goal is subdivided into Four dimensions (see Fig.1)
based on four sub-research objectives : Objective 1 : Evaluate privacy
measures in care metadata share landscape within Dutch care providers' inter/intra
organizational setup. Objective 2 : Identify the data utility and data quality
indicators in healthcare and assess whether they are pro/against privacy indicators
(PUT) on one hand and FAIR and FACT-based data indicators on the other.
Objective 3 : Evaluate both normatively and empirically the data ownership and
access governance on both technical and organizational grounds in the care
metadata share landscape. Objective 4 : identify privacy-preserving, higher data utility
prone measures from other domains (i.e. logistics, etc) that are e ectively
applicable to healthcare and are per regulatory and legislative requirements of the
EU.</p>
        <p>The respective four dimensions (each with an assigned color see
Fig.1) and their Research Questions (RQs) are as follow: Dimension 1 :
Privacy evaluation in care metadata share landscape at the backdrop of care
providers inter and intra-organizational setup. RQ1 : What are the
privacypreserving indicators in the care metadata share the landscape, how are they
implemented and assessed in Dutch inter and intra-organizational setup?
Dimension 2 : Quanti cation of the data value in healthcare and its relevance to
privacy (PUT) and FAIR and FACT-based data. RQ2.1 : What are the
respective indicators for data utility and data quality while sharing healthcare
metadata in inter/intra organizational setup. RQ2.2 : How respective indicators
support/ discourage the privacy indicators on one hand and FAIR (uninterrupted
data) and FACT (responsible data) based data indicators on the other?
Dimension 3 : Privacy evaluation of data ownership in healthcare metadata within and
amongst care providers. RQ3 : What is the normative (should be) and empirical
(as is) state of a airs of data ownership and access governance in care
metadata share at inter/intra organizational levels? Dimension 4 : Identify
privacypreserving, higher data utility prone measures from other domains i.e. logistics,
etc, that are e ectively applicable to healthcare and are per regulatory and
legislative requirements of the EU. RQ4.1 : What are useful privacy-preserving
tools/techniques in safeguarding an organization's integrity in addition to the
simultaneous sharing of valued metadata with other counterparts for the collective
bene t? RQ4.2 : How are those privacy-preserving, data utility-prone measures
applicable to healthcare, and are they e ective? Give normative and empirical
evaluation.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Data Analytics Approach and Research Methodology</title>
      <p>
        For each dimension (see Fig. 1), the following standardized data analytics
approach (and methodology) is applied (see Fig. 2 for an overview). The approach
combines the empirical and normative evaluation and aims to validate care
providers' integrated techniques, processes, and systems concerning PUT in the
Dutch (care) metadata share landscape. Empirical evaluation is done using
Process Mining (PM) discovery and conformance checking techniques (to validate
the empirical evaluation) on healthcare event logs i.e. datasets that are extracted
directly from Hospital Information System (HIS) [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. The objective is to
evaluate data utility in comparison to privacy preservation of sensitive data. Based
on the empirical analytics, conceptual modeling frameworks are drawn (using
REA or e3 value modeling ontologies) and are evaluated by in- eld (i.e. from
within the organization) IT expert(s). The aforementioned conceptual
modeling frameworks are selected for clarity and convenient understanding of the key
actors, their interactions, and mutual value gain for technical (and
organizational) proceedings. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. For normative evaluation, two alternative approaches
are followed as per the respective research objective to ascertain the provenance
record (brie y explained in the next paragraph) making. One approach follows
the `Content Analysis' (using literature review, o cial websites, and (online)
content) methodology to formulate conceptual modeling framework(s). We also
plan to include the `BPMN text extraction' methodology for document rule
mining [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The deduced BPMN model will be evaluated by the IT expert(s). So far,
publicly available event logs (comprising metadata share within and across a
local hospital) are used for research ndings, we are in process of a prospective
collaboration with an EU project for some interesting datasets and respective
project collaboration concerning inter-organization exchange of health data with
a special focus on privacy preservation.
      </p>
      <p>
        The PUT concerning issues either depend upon or directly in uence the
prescribed rst three dimensions. Which in turn incorporate the data pipeline
and data provenance [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Data pipeline allows automatic data gathering from
diverse sources and its integration into a data warehouse. In healthcare,
espeAND
(
udohS roNm
-esb itaev
ta e
te va
fo lau
fa it
ia on
f)r
s
(sA Em
-s ip
is ir
      </p>
      <p>c
ttae lea
fao lavu
iffa itao
)rs n</p>
      <p>Datfraoamcqmuuisltiitpiolen/sinotuergcreastion DaIntafoerxmtraatciotinonSyfrsotmemH(oHsIpSit)al PDraotcaeasnsaMlyitnicinsgu(sPinMg) De(dPuecteridNbeutss)ineevsaslumatoiodnels Furthienr-fdiealdtaITeveaxlupaetritosn by
Multiple Data sources</p>
      <p>HIS</p>
      <p>Event log</p>
      <p>Data Analytics
(Process Mining)</p>
      <p>Petri Nets
Evaluation</p>
      <p>Expert Opinion
(in-field)
Literature Review, Websites</p>
      <p>Documents</p>
      <p>Publicly/privately
available real-world
(care) datasets
Data Analytics
(Content Analysis)</p>
      <p>
        Conceptual Modeling
(REA,e3 value modeling)
Data Analytics
(Rule Mining)
(BPBMuNsiTneexstsEMxtordacetlion)
DpDraoatravteaidvaaiececqrwsoqu',nuiosoliisifnfXtfifiiteiocOiconiicaRanolflrfndwortoomeemcnbultsimctietaeerrasne,ttusre sCptbaroaynncdtdDtepaieconrrseidctvigiuAapznmcreni/yndapeclbrnypiiXyvtpsrsaaOiapsccRloRsyttuiolcifcbloeelyyorpMcpfporaroiirnintvleiiccacnaisycpgrtyaaeftolonbsdrdylafcoaodtacraredraepsitzeridgievandat/cay RCEBBoAuPfnr,saMceimneN3epXevstTwOuasealoRuxmlrtekmoE-mudoxsedtoirlnedauglecisntliiignongng
Further evaluation by
in-field IT experts
cially in HIS, the data is assimilated from diverse sources that include, amongst
others, administrative, medical, and physiological (including sensor-based) data
resources. Provenance is the documentation of the \data entities, systems and
processes" to avoid data manipulation and system misuse [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The potential
usefulness of provenance records depends directly upon the veracity (genuine
description) of the data pipeline. Provenance records allow: easier assessment
in complying with the regulatory prerequisites, identi cation, and recovery of
bottlenecks concerning systems and processes, and data security and privacy
preservation [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The aforementioned methods and their combination were
intentionally selected for not only analytical but also normative evaluation of PUT
to validate the integrity of the systems, tools, and processes in the contemporary
Dutch healthcare landscape. The prescribed fourth dimension, however, largely
relies upon the availability of the time.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Current Results</title>
      <p>
        By following the above-mentioned approach and methodology, two papers are
published and one paper is under review [
        <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
        ]. Normative evaluation is done
using the Content Analysis (explained above) methodology [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. A conceptual
modeling framework is designed on the fundamentals of the Padlock Chain Model
of e3 value modeling. The model is evaluated by the (in- eld) IT expert (An IT
head in a local hospital who is also a liated with a local diagnostic lab). The
model identi ed that the privacy is implemented with 'privacy by design,
'privacy by policy, and patients 'informed consent for care metadata sharing. The
indicators for privacy by design vary as per each healthcare provider's business
goals, and lack consistency across healthcare providers [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Privacy by
policy (Information Security Management System) indicators are regulated by ISO
(and NEN) and local regulatory authorities but these are only qualitative
evaluations. The lack of; consistent/above board privacy by design indicators and
quantitative evaluation of ISMS indicators, leaves ample room for technical and
organizational ambiguity for the care providers and in turn, gives vent to the
privacy lapses [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>
        Another research work identi es that the un-anonymized/pseudonymized
care metadata sharing within/across a local hospital poses serious (patients')
privacy concerns [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The empirical evaluation was done using Process Mining
on event logs from a local HIS. The evaluation identi ed that for the sake of e
cient/e ective performance-oriented care, the patients' un-anonymized metadata
is shared within horizontally oriented intra-organizational (i.e. between multiple
departments within the hospital, etc) caregivers [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The in- eld IT expert
evaluated that the ndings are generalizable to the vertically located (i.e. outpatient
caregivers such as lab, general practitioner, pharmacy, hospital) caregivers as
well. Based on the aforementioned empirical evaluation and (normative) content
analysis, the conceptual modeling framework using REA's Insurance Model [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
is extended [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The framework stresses recent advancements where
'Materialized Privacy Claims' are launched either by the patient or by any other potent
authority, such as the Dutch Data Protection O cer (DPO) and costs hundreds
of thousands of euros to the care providers [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>Another under review research work evaluates the PUT on care event log
from HIS. The PUT is evaluated using the ProM tool for identifying
noiseadding plugins which are data utility e cient as well. The plugins are evaluated
on three di erent datasets and two di erent versions of ProM and gave similar
results. The research work will assist the ProM tool's end-users to make use of
those plugins for privacy-preserving, utility-prone data analytics using Process
Mining. So far, the rst dimension and marginally the second dimension are
covered. The rest of the dimensions will be covered in the remaining years to
come.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Threats to Validity</title>
      <p>To increase the scope and to ascertain the functionality of the data
analytical approach, we are not con ning ourselves to one disease-speci c event log(s),
rather an analysis of more diversi ed datasets (with a focus on inter/intra
organizational data exchange) will allow us to conduct more realistic empirical
evaluation (in avoiding the selection bias). But on the other hand, this approach
can lead us to certain unforeseen pitfalls while aggregating the data information
as the results will be diversi ed yet non-conforming to one another. To avoid this
loophole, in the future, we aim to gather at least two or more event logs from
a similar sub-domain. Additionally, to gather the diversi ed datasets, we are
in discussion with the personnel who are directly involved with an EU project
and are interesting in a prospective collaboration with us concerning inter/intra
organizational care data exchange with a special focus on privacy preservation.</p>
      <p>1. Clarity regarding privacy definition and
its qualitative assessment measures. Policymakers need to
standardize the Privacy by Design measures and entrepreneurs
have to apply standardized technical measures to effectively
safeguard privacy to avoid hefty compensations.</p>
      <p>3. Understanding the target areas
regarding data ownership and access governance both
technically and organization-wise. Giving</p>
      <p>recommendations for both
4. Evalua on
of privacy-preserving, data u lity prone tools,
and techniques from other domains and their
evalua on in healthcare</p>
      <p>CONTRIBUTION
(DIMENSION-WISE)</p>
      <p>2. Understanding that the FAIR data
supports utility prone metadata sharing, whereas FACT
measures facilitate privacy-preserving ethical data sharing.</p>
      <p>Resolving PUT issues in care metadata sharing will extend
the implications to FAIR and</p>
      <p>FACT-based data.</p>
      <p>rapp iDm
ev aco sen
la h io
So far the discursive technical solutions hamper the holistic understanding of
the PUT in the care data share landscape. This research work aims to provide
the footing for the same. Even if all four dimensions are not completed by the
end of this Ph.D. work, at least it will provide the infrastructure with the rst
three dimensions which will (ideally) serve the basic purpose/goal.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Contribution</title>
      <p>
        So far, various divergent, remotely conducted analytical research works in
healthcare such as for Internet of Medical Things (IoMT) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Arti cial Intelligence
(AI) in healthcare [
        <xref ref-type="bibr" rid="ref13 ref8">8, 13</xref>
        ], generic IT-related challenges [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], Process Mining [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
exist. Similarly, normative evaluation with the standardized privacy-aware
conceptual modeling frameworks for the scalable, privacy-preserving systems for
IoT exist [
        <xref ref-type="bibr" rid="ref15 ref7 ref9">7, 9, 15</xref>
        ]. The connection between the empirical and normative
evaluations of PUT in care metadata sharing is still lacking. The proposed research
work proposes a unique holistic approach in resolving PUT concerning issues
in healthcare. Fig. 3 (with a recurring color scheme for each dimension as that
of Fig. 1) gives a detailed description of the dimension-wise contribution of the
proposed research work (see Fig. 3).
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>ITE is an amalgamation of business processes, techniques, and systems that
improve business proceedings in better achieving business goals. Currently,
enterprises encounter privacy concerning issues in performing performance-oriented
data analytics. Privacy-Utility-Tradeo (PUT) is the performance impairment
of Big Data Analytics (BDA) in ascertaining data privacy. Normative (should
be) and empirical (as-is) evaluation of PUT is essential to better nd suitable
techniques and processes for (care providers) Information Systems against PUT
in healthcare.</p>
      <p>The empirical evaluation is performed on real-world care event logs. The
ndings are drawn using conceptual modeling frameworks using REA and e3 value</p>
      <p>Normative and Empirical Evaluation of Privacy Utility Trade-off in Healthcare 19
modeling ontologies. For normative evaluation, two alternatives approaches are
followed. One content analysis technique provides a basis for the 'conceptual
modeling' frameworks and the other 'BPMN text extraction' allows documents
rule mining and formulation of BPM. (In- eld) IT experts further evaluated the
conceptual models.</p>
      <p>By following the afore-mentioned data analytics approach and methodology,
two papers are published and one paper is under review in the rst year of
this Ph.D. research work. The rst paper identi ed the loopholes in privacy
by design and privacy by policy measures. The second paper identi ed the
unanonymized/pseudonymized care metadata share amongst horizontally and
vertically located caregivers in the Dutch metadata share landscape. Another
underreview paper locates the privacy-preserving data utility-prone noise-adding
plugins in publicly available PM tool i.e. ProM.</p>
      <p>The future work comprises the quanti cation of the data value in healthcare
and its relevance to privacy (PUT) and FAIR and FACT-based data, privacy
evaluation of data ownership/stewardship in healthcare metadata within and
amongst care providers, identi cation of privacy-preserving, higher data utility
prone measures from other domains (i.e. logistics, etc) and their evaluation in
healthcare. PUT concerning issues of (healthcare) BDA is both technical and
organization-based. Thus, the issue requires evaluation of (care provider's)
integrated techniques, processes, and systems to nd respective solutions in (care)
metadata share landscape.</p>
      <p>Acknowledgments. I am grateful to my mentor, dr. Maurice van Keulen, and
my rst supervisor, dr. Faiza Allah Bukhsh for their exceptionally discerning
reviews, their undaunting support, and guidance, throughout.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Bpmntextextraction, https://sudonull.com/post/299-
          <string-name>
            <surname>Business-</surname>
          </string-name>
          processes
          <article-title>-E xtract-BPMN-model-from-</article-title>
          <string-name>
            <surname>document-</surname>
          </string-name>
          Part-
          <volume>1</volume>
          ,
          <issue>20</issue>
          <year>March</year>
          ,
          <year>2020</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Dutch-dpa, https://autoriteitpersoonsgegevens.nl/en/about-dutch
          <article-title>-dpa/b oard-dutch-</article-title>
          <string-name>
            <surname>dpa</surname>
          </string-name>
          ,
          <issue>11</issue>
          <year>Nov</year>
          ,
          <year>2020</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Gdpr, https://gdpr-info.eu/, 11 Nov,
          <year>2020</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. go.fair, https://www.go-fair.org/fair-principles/, 20 March,
          <year>2020</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. thedigitalsociety, https://www.thedigitalsociety.info/themes/responsibledata-science/, 20 March,
          <year>2020</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bichler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heinzl</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Responsible data science (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Arruda</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          ,
          <article-title>Bulca~o-</article-title>
          <string-name>
            <surname>Neto</surname>
            ,
            <given-names>R.F.</given-names>
          </string-name>
          :
          <article-title>Toward a lightweight ontology for privacy protection in iot</article-title>
          .
          <source>In: Proceedings of the 34th ACM/SIGAPP symposium on applied computing</source>
          . pp.
          <volume>880</volume>
          {
          <issue>888</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bohr</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Memarzadeh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Arti cial intelligence in healthcare</article-title>
          . Elsevier Science &amp;
          <string-name>
            <surname>Technology</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Can</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yilmazer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Improving privacy in health care with an ontology-based provenance management system</article-title>
          .
          <source>Expert Systems</source>
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <year>e12427</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Davenport</surname>
            ,
            <given-names>T.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Short</surname>
            ,
            <given-names>J.E.:</given-names>
          </string-name>
          <article-title>The new industrial engineering: information technology and business process redesign (</article-title>
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Garattini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ra</surname>
            <given-names>e</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Aisyah</surname>
            ,
            <given-names>D.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sartain</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlakidis</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Big data analytics, infectious diseases and associated ethical impacts</article-title>
          .
          <source>Philosophy &amp; technology 32(1)</source>
          ,
          <volume>69</volume>
          {
          <fpage>85</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Guan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lv</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guizani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Achieving data utility-privacy tradeo in internet of medical things: A machine learning approach</article-title>
          .
          <source>Future Generation Computer Systems</source>
          <volume>98</volume>
          ,
          <fpage>60</fpage>
          {
          <fpage>68</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hlavka</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          : Security, privacy, and
          <article-title>information-sharing aspects of healthcare arti cial intelligence</article-title>
          .
          <source>In: Arti cial Intelligence in Healthcare</source>
          , pp.
          <volume>235</volume>
          {
          <fpage>270</fpage>
          .
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hruby</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Model-driven design using business patterns</article-title>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Iwaya</surname>
            ,
            <given-names>L.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giunchiglia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martucci</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hume</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fischer-</surname>
            Hubner, S.,
            <given-names>ChenuAbente</given-names>
          </string-name>
          , R.:
          <article-title>Ontology-based obfuscation and anonymisation for privacy</article-title>
          .
          <source>In: IFIP International Summer School on Privacy and Identity Management</source>
          . pp.
          <volume>343</volume>
          {
          <fpage>358</fpage>
          . Springer (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>K.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joukov</surname>
          </string-name>
          ,
          <source>N.: Information Science and Applications (ICISA)</source>
          <year>2016</year>
          , vol.
          <volume>376</volume>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mannhardt</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koschmider</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baracaldo</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidlich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
          </string-name>
          , J.:
          <article-title>Privacypreserving process mining</article-title>
          .
          <source>Business &amp; Information Systems Engineering</source>
          <volume>61</volume>
          (
          <issue>5</issue>
          ),
          <volume>595</volume>
          {
          <fpage>614</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Data provenance and security</article-title>
          .
          <source>IEEE Security &amp; Privacy</source>
          <volume>9</volume>
          (
          <issue>2</issue>
          ),
          <volume>83</volume>
          {
          <fpage>85</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>McSherry</surname>
            ,
            <given-names>F.D.</given-names>
          </string-name>
          :
          <article-title>Privacy integrated queries: an extensible platform for privacypreserving data analysis</article-title>
          .
          <source>In: Proceedings of the 2009 ACM SIGMOD International Conference on Management of data</source>
          . pp.
          <volume>19</volume>
          {
          <issue>30</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Mohammed</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fung</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohno-Machado</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Privacypreserving heterogeneous health data sharing</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <volume>462</volume>
          {
          <fpage>469</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Pika</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wynn</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budiono</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ter</surname>
            <given-names>Hofstede</given-names>
          </string-name>
          , A.H., van der Aalst,
          <string-name>
            <given-names>W.M.</given-names>
            ,
            <surname>Reijers</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.A.</surname>
          </string-name>
          :
          <article-title>Privacy-preserving process mining in healthcare</article-title>
          .
          <source>International journal of environmental research and public health 17(5)</source>
          ,
          <volume>1612</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pramanik</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lau</surname>
          </string-name>
          , R.Y.,
          <string-name>
            <surname>Hossain</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahoman</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Debnath</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashed</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uddin</surname>
            ,
            <given-names>M.Z.</given-names>
          </string-name>
          :
          <article-title>Privacy preserving big data analytics: A critical analysis of state-of-the-art</article-title>
          .
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge</source>
          Discovery p.
          <year>e1387</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sohail</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allah</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krabbe</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          :
          <article-title>Identifying materialized privacy claims of clinical-care metadata share using process-mining and rea ontology</article-title>
          .
          <source>In: 15th International Workshop on Value Modelling and Business Ontologies</source>
          ,
          <string-name>
            <surname>VMBO</surname>
          </string-name>
          <year>2021</year>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Sohail</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krabbe</surname>
            , J., de Alencar Silva,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bukhsh</surname>
            ,
            <given-names>F.A.</given-names>
          </string-name>
          :
          <article-title>Privacy value modeling: A gateway to ethical big data handling</article-title>
          .
          <source>In: 14th International Workshop on Value Modelling and Business Ontologies</source>
          ,
          <string-name>
            <surname>VMBO</surname>
          </string-name>
          <year>2020</year>
          . pp.
          <volume>5</volume>
          {
          <fpage>15</fpage>
          .
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Srinivasan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Guide to Big Data Applications</article-title>
          , vol.
          <volume>26</volume>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Van Der Aalst</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Data science in action</article-title>
          . In: Process mining, pp.
          <volume>3</volume>
          {
          <fpage>23</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>von Voigt</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fahrenkrog-Petersen</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janssen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koschmider</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschorsch</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mannhardt</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landsiedel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidlich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Quantifying the reidenti cation risk of event logs for process mining</article-title>
          .
          <source>In: International Conference on Advanced Information Systems Engineering</source>
          . pp.
          <volume>252</volume>
          {
          <fpage>267</fpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Wieringa</surname>
          </string-name>
          , R.J.:
          <article-title>Design science methodology for information systems</article-title>
          and software engineering. Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>