<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Analyzing the heterogeneity of rule-based EHR phenotyping algorithms in CALIBER and the UK Biobank</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Spiros Denaxas</string-name>
          <email>s.denaxas@ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Helen Parkinson</string-name>
          <email>parkinso@ebi.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalie Fitzpatrick</string-name>
          <email>n.fitzpatrick@ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cathie Sudlow</string-name>
          <email>Cathie.Sudlow@ed.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harry H</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>mingw</string-name>
          <email>h.hemingway@ucl.ac.uk</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Medical Informatics, Usher Institute of Population Health Science and Informatics, University of Edinburgh</institution>
          ,
          <addr-line>Edinburgh</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>European Bioinformatics Institute</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Health Data Research UK London/Cambridge/Scotland</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Institute of Health Informatics, University College London</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>UCL Hospitals Biomedical Research Center</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>6</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>Electronic Health Records (EHR) are data generated during routine interactions across healthcare settings and contain rich, longitudinal information on diagnoses, symptoms, medications, investigations and tests. A primary use-case for EHR is the creation of phenotyping algorithms used to identify disease status, onset and progression or extraction of information on risk factors or biomarkers. Phenotyping however is challenging since EHR are collected for different purposes, have variable data quality and often require significant harmonization. While considerable effort goes into the phenotyping process, no consistent methodology for representing algorithms exists in the UK. Creating a national repository of curated algorithms can potentially enable algorithm dissemination and reuse by the wider community. A critical first step is the creation of a robust minimum information standard for phenotyping algorithm components (metadata, implementation logic, validation evidence) which involves identifying and reviewing the complexity and heterogeneity of current UK EHR algorithms. In this study, we analyzed all available EHR phenotyping algorithms (n=70) from two large-scale contemporary EHR resources in the UK (CALIBER and UK Biobank). We documented EHR sources, controlled clinical terminologies, evidence of algorithm validation, representation and implementation logic patterns. Understanding the heterogeneity of UK EHR algorithms and identifying common implementation patterns will facilitate the design of a minimum information standard for representing and curating algorithms nationally and internationally.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>In the United Kingdom (UK), structured electronic health
records (EHR) spanning primary care, hospital care,
disease/procedure registries and death registries are used to
create longitudinal disease phenotypes for observational
research studies [Hemingway et al., 2018]. Through a
process called phenotyping, researchers create algorithms
which utilize multiple EHR sources to accurately extract
information on diseases (e.g. status, onset and progression),
lifestyle risk factors and biomarkers [Banda et al., 2018].
Phenotyping however is challenging due to the fact that
EHR are fragmented, curated using different controlled
clinical terminologies and collected for purposes other than
research (e.g. reimbursement, audit) [Morley et al., 2014].
Phenotyping requires a significant amount of resources and
mix of expertise, yet no common standard approach for
defining, validating and ultimately sharing EHR
phenotyping algorithms currently exists. In the UK,
structured primary care EHR have been used in &gt;1,800
peer-reviewed studies to date but only 5% of studies
published sufficiently reproducible phenotypes [Springate et
al., 2014]. Defining a standardized format to represent EHR
phenotypes will enable portability across data sources (and
healthcare systems) and facilitate the systematic sharing of
algorithms across the community [Mo et al. 2015].
Compared to the United States (US), the UK EHR research
landscape differs in two important ways: 1) researchers can
utilize multiple national EHR sources to create longitudinal
‘cradle to grave’ phenotypes [Kuan et al., 2019], and 2) UK
primary care EHR contain both healthy and unhealthy
individuals which allow researchers to capture information
on disease severity and progression over time. A recent
systematic review identified 66 different definitions used to
capture asthma status and exacerbations in research using
UK EHR [Al Sallakh et al., 2017] demonstrating significant
existing heterogeneity. While analyses have been
undertaken in the US to characterize the heterogeneity of
phenotyping algorithms [Conway et al., 2011], no such
analysis has been carried out in the UK.</p>
      <p>
        One of the aims of the newly-established national institute
for health data science, Health Data Research UK (HDR
UK, www.hdruk.ac.uk), is the creation of a national
Phenomics Resource: an open-access online resource where
EHR phenotypes can be deposited and curated. A critical
first step in this process is to establish a minimum
information standard for representing EHR phenotyping
algorithms. This involves exploring and documenting the
complexity, heterogeneity, design and implementation
patterns of contemporary phenotyping algorithms in the UK.
The concept of a minimum information standard has been
used successfully in other biomedical disciplines, e.g.
Minimum Information About a Microarray Experiment
(MIAME) defines standards for reporting microarray
experiments [
        <xref ref-type="bibr" rid="ref5">Brazma et al., 2001</xref>
        ]. Establishing a
standardized method for representing phenotypes in the UK
can potentially address these challenges and ensure
compatibility with other international initiatives such as
eMERGE and PCORNet [Fleurence et al. 2014; Gottesman
et al. 2013].
      </p>
      <sec id="sec-2-1">
        <title>Aims</title>
        <p>Despite the widespread use of UK EHR data sources for
research, contemporary research resources utilize different
approaches for algorithm creation, curation and validation.
The aims of this study were to: a) identify and characterize
the structural components, implementation logic and
heterogeneity of rule-based algorithms defining diseases,
lifestyle risk factors and biomarkers in structured national
EHR in the UK utilized by contemporary research
resources, and b) propose a minimum information standard
to represent UK EHR phenotyping algorithms.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Methods</title>
        <p>We identified, downloaded and reviewed published
phenotyping algorithms for diseases, biomarkers and
lifestyle risk factors from two large-scale contemporary UK
research resources: UK Biobank1 and CALIBER2.
The UK Biobank [Sudlow et al., 2015] is a prospective
cohort study of 500,000 (aged 40-69 at recruitment) adults
recruited in England, Scotland and Wales from 2006-2010.
For each participant, deep phenotypic and genotypic
information is available including biomarkers in blood and
urine, imaging (brain, heart, abdomen, bone, carotid artery),
lifestyle indicators, pathophysiological measurements and
genome-wide genotype data. Follow-up for health outcomes
is enabled by hospital EHR (Hospital Episode Statistics
(HES) in England, Patient Episode Data Warehouse in
Wales and Scottish Morbidity Registry in Scotland) and
linkages to primary care EHR are underway. CALIBER
[Denaxas et al., 2012; Denaxas et al., 2019] is a research
resource consisting of algorithms, tools and methods for
structured EHR linked across primary care (Clinical Practice
Research Datalink, CPRD), hospital care (HES) and a
mortality data (Office for National Statistics, ONS) in the
UK.</p>
        <p>In the UK, national EHR are recorded using controlled
clinical terminologies where terms are assigned at variable
timepoints i.e. in UK primary care the physician records
terms in real time during the consultation with the patient
whereas in hospital care terms are retrospectively entered
into databases by trained coders and data selected for billing
purposes. We identified and counted the number of
ontology terms each algorithm utilizes from five controlled
clinical terminologies which are widely used in the UK: a)
Read (primary care, subset of SNOMED-CT), b)
International Classification of Diseases 9th and 10th
Revision (ICD-9, ICD-10, secondary care diagnoses and
cause of mortality), c) OPCS Classification of Interventions
and Procedures (OPCS-4, hospital surgical procedures,
analogous to the Current Procedural Terminology ontology
used in the United States), and d) the Dictionary of
Medicines and Devices (DM+D) which is used to record
primary care prescriptions. Terms were automatically
extracted from documents and counted using regular
expressions in Python 3.63. We manually extracted and
counted terms across five randomly chosen algorithms to
verify the automatically-generated counts.</p>
        <p>
          EHR phenotype validation is a critical process guiding the
subsequent use of algorithms and we were interested in what
types, if any, of evidence were available to external
researchers. We classified the available material into six
non-overlapping categories which encapsulate all potential
1 http://biobank.ndph.ox.ac.uk/showcase/label.cgi?id=42
2 https://www.caliberresearch.org/portal/phenotypes
3 https://www.python.org/
approaches for obtaining validity evidence
          <xref ref-type="bibr" rid="ref16 ref19 ref9">(adapted from
[Denaxas et al, 2019] and recorded as used/not used)</xref>
          :
• Aetiological: Are the prospective associations with risk
factors consistent with previous published evidence
from both EHR and non-EHR studies?
• Prognostic: Are the risks of subsequent events
plausible and consistent with existing domain
knowledge?
• Case-note review: What is the positive predictive value
(PPV) and the negative predictive value (NPV) when
comparing the algorithm with clinician-led review of
case notes, self-reported information or a suitable “gold
standard” source?
• Cross-EHR-source concordance: To what extent is
the phenotype concordant across EHR sources?
• Genetic: Are the observed genetic associations
plausible and consistent in terms of magnitude and
direction of association with associations reported from
non-EHR studies?
• External populations: Has the algorithm been
evaluated in different countries or external sources?
For each algorithm, we documented the EHR sources the
phenotype is derived from (i.e. primary care, hospital care,
mortality register). We extracted information on the
representation components of phenotypes e.g. the presence
of tabular data and the use of a flowchart (or other graphical
presentation). We extracted and categorized information on
the different types of implementation logic, temporality and
algorithm implementation patterns (Table 1), partially based
on previous research in the US [Conway et al., 2011].
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Concept Definition Example</title>
        <p>Simple Simple PVD diagnosis during a
Boolean Boolean primary care
statements consultation OR
e.g. diagnosis of leg or aortic
“AND”, embolism or thrombosis
“OR” during a hospitalization
Complex Nested IF patient = diabetic: HT
Boolean statements threshold: SBP ≥140
with mmHg OR DBP ≥90
multiple mmHg ELSE: threshold
layers? SBP ≥150 mmHg OR</p>
        <p>DBP ≥90 mmHg
Negation Are No AF diagnosis term is
negation present, but the patient
statements record includes a
used? warfarin prescription in
the absence of prior
DVT or PE, or a digoxin
prescription but no HF
Temporal Temporal Iron deficiency anaemia
(simple) proximity record in primary care
future or OR hospital AND
past endoscopy in 30 days</p>
      </sec>
      <sec id="sec-2-4">
        <title>Temporal (complex)</title>
      </sec>
      <sec id="sec-2-5">
        <title>Biomarker</title>
      </sec>
      <sec id="sec-2-6">
        <title>Complex calculation</title>
        <p>
          We identified and reviewed 70 EHR phenotyping (Table 2)
algorithms available from the UK Biobank (n=19) and the
CALIBER resource (n=51). The majority of phenotyping
algorithms were created to ascertain disease status (n=54)
          <xref ref-type="bibr" rid="ref12 ref13 ref15 ref16 ref19 ref2 ref3 ref7 ref9">(e.g. heart failure [Gho et al. 2018; Uijl et al. 2019],
depression [Daskalopoulou et al. 2016])</xref>
          , ten algorithms
were created to extract information on biomarkers
          <xref ref-type="bibr" rid="ref11 ref12 ref15 ref18 ref2 ref3">(e.g.
heart rate [Archangelidi et al. 2018], blood pressure
[Rapsomaniki et al. 2014])</xref>
          and six algorithms were used to
identify lifestyle risk factors
          <xref ref-type="bibr" rid="ref1 ref10 ref17 ref20 ref4">(e.g. alcohol [Bell et al. 2017],
smoking [Pujades-Rodriguez et al. 2015])</xref>
          .
        </p>
        <p>All but one CALIBER phenotyping algorithm (n=50) used
information from primary care EHR with the exception of
socioeconomic status which was defined using the Index of
Multiple Deprivation (IMD) provided by the ONS.
Algorithms defining biomarker measurements (e.g. white
blood cells, heart rate) were based on primary care EHR
entirely while approximately half of the algorithms
ascertaining disease status (n=19 of 35) combined
information across all three EHR sources. All currently
available UK Biobank algorithms (n=19) combined
information recorded during the baseline assessment (data
not shown), diagnoses and/or surgical procedures recorded
during hospitalization and information based on the
underlying (or secondary) cause of death which is recorded
in the national mortality register. Primary care linkages in
UK Biobank are still underway and as a result none of the
currently available algorithms utilized information from
primary care EHR. However, primary care information for
just under half of the cohort (n=230,000) will be made
available for UK Biobank researchers in June 2019.
Algorithms incorporating primary care data for the
conditions already covered have been or are being
developed [Wilkinson et al 2019]. Along with a range of
additional algorithms expanding the range of health
outcomes available, they will be available from UK Biobank
later in 2019. Overall, based on current publicly available
information from CALIBER and UK Biobank, 75% (n=66)
of algorithms used data from secondary care EHR and 45%
(n=49) used information available in the death registry.
The most widely-used clinical terminology was Read with
4,729 (non-unique) terms used across all algorithms while
the second highest number of terms was derived from the
DM+D with 2,273 (non-unique) terms used to record
prescriptions in primary care EHR. Four algorithms (body
mass index, socioeconomic deprivation, sex, heart rate) did
not use any terms across any terminology systems and were
based on information which is derived from a structured
field of the EHR or externally linked such as in the case of
IMD. The atrial fibrillation algorithm used the highest
number of clinical terms (n=987) while across all algorithms
the pregnancy phenotype used the highest number of Read
codes (n=1,948). ICD-9 was the terminology least used: in
the UK Biobank it is used for recording diagnoses in older
Scottish hospital records and in CALIBER it is used to
record the cause of death prior to 1997. Algorithms defining
biomarkers contained the lowest number of terminology
terms as they relied on structured data fields combined with
a small number of diagnosis terms to denote the type of test
(e.g. Read code “42K..00 Eosinophil count”).</p>
        <p>With regards to algorithm implementation logic, 66 (93%)
of algorithms used Boolean statements, usually to identify
the presence of one or more diagnosis codes in a patient’s
EHR. Where Boolean statements were deployed, in nearly
half of the cases these were complex and involved either a
series of nested statements or joined information across
multiple sources, for example in the UK Biobank where
information is derived from self-reported, hospital and
mortality sources and events are further stratified as
‘prevalent’ (first reported prior to recruitment) or ‘incident’
(first reported after recruitment). A similar pattern of logic
was observed with regards to temporality where 66
algorithms utilized temporal rules and almost always this
included more complex statements and restrictions. Finally,
approximately half (n=43) of the algorithms used negation.
Only ten algorithms (16%) included more complex
calculations, usually to calculate the mean of multiple
measurements on the same day or to harmonize units for
laboratory measurements to a common format.</p>
        <p>Prognostic 86% (n=66) and cross-source concordance 54%
(n=43) validation approaches where the most widely-used
algorithm evaluation approaches. The least-widely used
validation approach was expert case note review, although
this type of validation has been completed for a few UK
Biobank algorithms, including dementia and its subtypes
[Wilkinson et al, 2019], and is underway for several others.
Most (93% [n=66]) of the algorithms used data stored in
tabular format since tables are predominantly used to store
lists of controlled clinical terminology terms. Only 25%
(n=15) of algorithms included a graphical representation of
the algorithm using a flowchart and all algorithms included
a textual description of the algorithm components.</p>
      </sec>
      <sec id="sec-2-7">
        <title>5. Discussion</title>
        <p>In this study we downloaded and reviewed 70 EHR
phenotyping algorithms from two large-scale, national
research resources in the UK. We reviewed algorithms in
terms of EHR data sources, controlled clinical terminologies
used, available evidence of algorithm validation, algorithm
representation formats and implementation logic patterns.
Similar to findings from US studies, we discovered that UK
EHR algorithms make extensive use of Boolean statements
and temporal logic. When these are used, they are often
complex i.e. combining multiple nested Boolean layers of
logic and defining temporal proximity rules within them.
This is expected given that algorithms utilize multiple
sources of information and include evidence from primary
care and hospital care (or self-reported information in the
case of the UK Biobank). Algorithms defining disease status
were the most frequent and complex algorithms reviewed
and utilized the greatest number of terms from controlled
clinical terminologies. Negation was another major
component of algorithms and is often used to exclude
concomitant diagnoses or procedures when trying to
ascertain diseases based on secondary information (e.g.
ascertaining AF cases based on a prescription of digoxin but
excluding patients which are diagnosed with HF).
The Read clinical terminology was the most popular
terminology used with the highest number of terms per
phenotype. These findings are expected as Read contains a
significant amount of duplication internally due to synonym
terms which can be potentially utilized. Additionally, the
clinical concepts contained within Read subsume the
concepts across all other terminologies i.e. Read contains
terms for diagnoses, symptoms, laboratory tests,
prescriptions and procedures. UK primary care clinical
coding is currently transitioning to SNOMED-CT which
should provide a more streamlined set of terms to be used.
In terms of validation, we observed a significant level of
heterogeneity with approaches seeking to evaluate and
replicate previously reported aetiological and prognostic
estimates from non-EHR studies being the most popular.
The presentation of the evidence however does not follow a
common standard and sometimes only included references
to published research rather than a more structured abstract
of the main findings of the analyses. In contrast with the
US, expert review of case records was the least frequently
used approach for evaluation due to the fact that large scale
corpuses of medical text do not exist in the UK owing to
information governance restrictions and the technical
challenges of integrating such data since they are held in a
wide range of formats by multiple different NHS
organisations. For similar reasons, none of the algorithms
reviewed utilize medical text and natural language
processing approaches to extract information from medical
notes which is prevalent in some clinical specialties such as
mental health [Wu et al. 2018].</p>
        <p>Significant heterogeneity was also observed in terms of
representation. UK Biobank algorithms were curated in
individual PDF files4 and included extended information on
the goal of the algorithm and useful background knowledge
and references. In contrast, CALIBER phenotypes were
stored in an online, openly-available Portal5, spanned
multiple pages and did not include much background
information. Flowcharts or similar graphical representations
were not widely-used and while they are not
machinereadable, they can potentially minimize errors during
translation of the algorithm to machine code.</p>
        <p>Our study has potential limitations. We reviewed algorithms
from only two UK sources. While other UK initiatives exist,
they tend to focus on curating lists of controlled clinical
terminology terms (referred to as codelists) rather than
selfcontained phenotypes i.e. terms, implementation, validation
evidence. We only focused on rule-based approaches and
did not cover machine learning approaches. While
rulebased methods are the most widely used in the UK,
datadriven high-throughput approaches including natural
language processing methods are emerging [Zhou et al.,
2016, Pikoula et al., 2019]. These approaches pose different
challenges and their requirements would need to be
documented and analysed in order to ensure their integration
[Hripcsak &amp; Albers 2013]. Finally, reproducible research
approaches [Denaxas et al., 2017, Goodman et al, 2016]
which are covered elsewhere would also need to be
carefully taken into consideration in order to ensure
algorithm portability.</p>
      </sec>
      <sec id="sec-2-8">
        <title>6. Steps towards a minimum information standard</title>
        <p>Based on our findings, we propose that an EHR
phenotyping algorithm representation combines metadata,
4 http://biobank.ndph.ox.ac.uk/showcase/label.cgi?id=42
5 https://www.caliberresearch.org/portal
implementation logic, validation evidence and use-cases.
We suggest the following components towards establishing
a minimum information standard with regards to rule-based
phenotyping algorithms for UK EHR:
Part 1 – Algorithm metadata: Succinct information about
the goal of the algorithm, the intended use-case, the data
sources and controlled clinical terminologies used,
applicable age groups and genders, list of authors and their
contact details and a set of SNOMED-CT terms to classify
the algorithm. A unique identifier, such as a Digital Object
Identifier (DOI), should be minted to enable usage tracking
in subsequent research.</p>
        <p>Part 2 – Implementation: Details on the implementation
logic of the algorithm with pseudocode to facilitate the
translation to machine code and documentation on decisions
made and reasoning. Where possible analytical scripts
should be attached using markdown or a similar approach.
The standard should support defining complex Boolean and
temporal logic across multiple EHR sources and clinical
terminologies. In the future, a computable phenotype format
should encapsulate this information as a stand-alone file.
Part 3 – Validation evidence: Description of the steps
taken to support phenotype validity across six categories
(aetiological, prognostic, genetic, expert review,
crosssource and external population). For each implementation,
the number of cases, controls, NPV and PPV values should
be reported and the format should support the embedding of
graphical files (e.g. forest plots).</p>
        <p>Part 4 – Use-cases: Links to published research utilizing
the phenotype algorithms, cross-referenced with DOI’s.</p>
      </sec>
      <sec id="sec-2-9">
        <title>7. Conclusion</title>
        <p>Our analyses identified a certain level of underlying
homogeneity in terms of how phenotyping algorithms are
defined and evaluated. We suggest four components
towards a minimum information standard that should be
used to represent phenotyping algorithms. These findings
provide a crucial first step towards curating and
disseminating phenotyping algorithms utilizing UK EHR.
Further work is required towards establishing a computable
format for phenotyping algorithms and ensuring
interoperability with other resources (e.g. PheKB).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <p>This work was supported by Health Data Research UK,
which receives its funding from HDR UK Ltd (LOND1)
funded by the UK Medical Research Council, Engineering
and Physical Sciences Research Council, Economic and
Social Research Council, Department of Health and Social
Care (England), Chief Scientist Office of the Scottish
Government Health and Social Care Directorates, Health
and Social Care Research and Development Division
(Welsh Government), Public Health Agency (Northern
Ireland), British Heart Foundation and the Wellcome Trust.</p>
      <p>The BigData@Heart Consortium is funded by the
Innovative Medicines Initiative-2 Joint Undertaking under
grant agreement No. 116074. This study was supported by
the Farr Institute of Health Informatics Research at UCL
Partners (MR/K006584/1). This paper represents
independent research part funded by the National Institute
for Health Research Biomedical Research Centre at UCLH.</p>
      <p>HH is a NIHR Senior Investigator. SD is an Alan Turing
Fellow.
6PrimarycareEHRavailableforparticipantsin2019;case-notereviewvalidationunderwayformultiplephenotypes.
Parkinsonism
PD
Stroke NS
VD
n/a
n/a
n/a
n/a
n/a
n/a
n/a
+
+
+
+
+
+
+
+
+
+
+
+
+
+
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
+
+
+
+
+
+
+
+
+
+
+
+
AAA Abdominal Aortic Aneurysm; AD Alzheimer's Disease; AF Atrial Fibrillation; AMI Acute Myocardial Infarction; AU Autoimmune Uveitis; BMI
Body Mass Index; BP Blood Pressure; BuP Bullous Pemphigoid; CHD Coronary Heart Disease; FTD Frontotemporal dementia; GCA Giant Cell Arteritis;
HCM Hypertrophic Cardiomyopathy; HDL High Density Lipoprotein cholesterol; HF Heart Failure; HIV Human Immunodeficiency Virus; HR Heart Rate;
HT Hypertension; ICH Intracerebral Haemorrhage; LDL Low Density Lipoprotein cholesterol; MS Multiple Sclerosis; NS Not Specified; PAD Peripheral
Arterial Disease; PBC Primary Biliary Cirrhosis; PMR Polymyalgia Rheumatica; RA Rheumatoid Arthritis; SA Stable Angina; SAH Subarachnoid
Haemorrhage; SCD Sudden Cardiac Death; TIA Transient Ischaemic Attack; UA Unstable Angina; UCD Unheralded Coronary Death; VD Vascular
Dementia; WBC White Blood Cell Count; COPD Chronic Obstructive Pulmonary Disease; ESRD End Stage Renal Disease; MND Motor Neuron Disease;
PD Parkinson's Disease and Parkinsonism; MSA Multiple System Atrophy; PSP Progressive Supranuclear Palsy; STEMI ST-Elevation AMI; NSTEMI
NonST Elevation AMI</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Al Sallakh et al.
          <year>2017</year>
          ]
          <string-name>
            <given-names>Al</given-names>
            <surname>Sallakh</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. A.</surname>
          </string-name>
          , et al.
          <article-title>Defining asthma and assessing asthma outcomes using electronic health record data: a systematic scoping review</article-title>
          .
          <source>Eur. Respiratory J.</source>
          ,
          <volume>49</volume>
          (
          <issue>6</issue>
          ),
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Archangelidi et al.,
          <year>2018</year>
          ] Archangelidi,
          <string-name>
            <surname>O.</surname>
          </string-name>
          , et al.
          <source>Clinically Recorded Heart Rate and Incidence of 12 Coronary</source>
          , Cardiac, Cerebrovascular and Peripheral Arterial Diseases in
          <volume>233</volume>
          ,970 Men and
          <article-title>Women: A Linked Electronic Health Record Study</article-title>
          .
          <source>Eur. J. of Preventive Cardiology</source>
          <volume>25</volume>
          (
          <issue>14</issue>
          ):
          <fpage>1485</fpage>
          -
          <lpage>95</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Banda et al.,
          <year>2018</year>
          ] Banda,
          <string-name>
            <surname>J. M.</surname>
          </string-name>
          , et al.
          <article-title>Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models</article-title>
          .
          <source>Annual Review of Biomedical Data Science</source>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Bell et al.,
          <year>2017</year>
          ] Bell,
          <string-name>
            <surname>S</surname>
          </string-name>
          , et al.
          <source>Association between Clinically Recorded Alcohol Consumption and Initial Presentation of 12 Cardiovascular Diseases: Population Based Cohort Study Using Linked Health Records. BMJ 356: j909</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Brazma</surname>
          </string-name>
          et al.,
          <year>2001</year>
          ] Brazma,
          <string-name>
            <surname>A.</surname>
          </string-name>
          et al.,
          <article-title>Minimum information about a microarray experiment (MIAME)- toward standards for microarray data</article-title>
          .
          <source>Nature Genetics</source>
          ,
          <volume>29</volume>
          (
          <issue>4</issue>
          ),
          <fpage>365</fpage>
          -
          <lpage>371</lpage>
          .
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Conway et al.,
          <year>2011</year>
          ] Conway,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , et al.
          <article-title>Analyzing the heterogeneity and complexity of Electronic Health Record oriented phenotyping algorithms</article-title>
          .
          <source>Proc. Am Med</source>
          Infor Assoc.,
          <fpage>274</fpage>
          -
          <lpage>283</lpage>
          ,
          <year>2011</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Daskalopoulou et al.,
          <year>2016</year>
          ] Daskalopoulou,
          <string-name>
            <surname>M.</surname>
          </string-name>
          et al.,
          <article-title>Depression as a Risk Factor for the Initial Presentation of Twelve Cardiac, Cerebrovascular, and Peripheral Arterial Diseases: Data Linkage Study of 1.9 Million Women and Men</article-title>
          .
          <source>PLOS ONE 11</source>
          (
          <issue>4</issue>
          ): e0153838,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Denaxas et al,
          <year>2012</year>
          ] Denaxas,
          <string-name>
            <surname>S.</surname>
          </string-name>
          et al.
          <article-title>Data resource profile: cardiovascular disease research using linked bespoke studies and electronic health records (CALIBER)</article-title>
          .
          <source>Int. J. Epidemiology</source>
          ,
          <volume>41</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1625</fpage>
          -
          <lpage>1638</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Denaxas et al.,
          <year>2019</year>
          ] Denaxas,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , et al.
          <article-title>UK phenomics platform for developing and validating EHR phenotypes: CALIBER</article-title>
          .
          <source>J Am Med Inf</source>
          <volume>10</volume>
          .1093/jamia/ocz105,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Denaxas et al. 2017]
          <string-name>
            <surname>Denaxas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          et al.,
          <article-title>Methods for enhancing the reproducibility of biomedical research findings using electronic health records</article-title>
          .
          <source>BioData Mining</source>
          ,
          <volume>10</volume>
          (
          <issue>31</issue>
          ),
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Fleurence et al.,
          <year>2014</year>
          ] Fleurence,
          <string-name>
            <surname>R.</surname>
          </string-name>
          , et al.
          <article-title>Launching PCORnet, a National Patient-Centered Clinical Research Network</article-title>
          . JAMIA
          <volume>21</volume>
          (
          <issue>4</issue>
          ):
          <fpage>578</fpage>
          -
          <lpage>82</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Gho et al. 2018]
          <string-name>
            <surname>Gho</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.
          <article-title>An Electronic Health Records Cohort Study on Heart Failure Following Myocardial Infarction in England: Incidence and Predictors</article-title>
          .
          <source>BMJ Open</source>
          <volume>8</volume>
          (
          <issue>3</issue>
          ):
          <fpage>e018331</fpage>
          .,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Goodman et al.,
          <year>2016</year>
          ] Goodman,
          <string-name>
            <surname>S.N.</surname>
          </string-name>
          , et al.
          <source>What does research reproducibility mean? Science Translational Medicine</source>
          ,
          <volume>8</volume>
          (
          <issue>341</issue>
          ), p.
          <fpage>341ps12</fpage>
          .,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Gottesman et al.,
          <year>2013</year>
          ] Gottesman,
          <string-name>
            <surname>O.</surname>
          </string-name>
          , et al. “
          <article-title>The Electronic Medical Records and Genomics (eMERGE) Network: Past, Present, and</article-title>
          <string-name>
            <surname>Future.</surname>
          </string-name>
          ”
          <source>Genetics in Medicine</source>
          <volume>15</volume>
          (
          <issue>10</issue>
          ):
          <fpage>761</fpage>
          -
          <lpage>71</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Hemingway et al.,
          <year>2018</year>
          ] Hemingway,
          <string-name>
            <surname>H.</surname>
          </string-name>
          , et al.
          <article-title>Big data from electronic health records for early and late translational cardiovascular research: challenges and potential</article-title>
          .
          <source>European Heart J.</source>
          ,
          <volume>39</volume>
          (
          <issue>16</issue>
          ),
          <fpage>1481</fpage>
          -
          <lpage>1495</lpage>
          ,
          <year>2018</year>
          [Hripcsak &amp; Albers, 2013] Hripcsak,
          <string-name>
            <given-names>G.</given-names>
            &amp;
            <surname>Albers</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.J.</surname>
          </string-name>
          <article-title>Next-generation phenotyping of electronic health records</article-title>
          .
          <source>JAMIA</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ),
          <fpage>117</fpage>
          -
          <lpage>121</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Kuan et al.,
          <year>2019</year>
          ] Kuan,
          <string-name>
            <surname>V.</surname>
          </string-name>
          et al.
          <article-title>A chronological map of 308 physical and mental health conditions from 4 million individuals in the English National Health Service</article-title>
          .
          <source>The Lancet Digital Health</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          ),
          <fpage>e63</fpage>
          -
          <lpage>e67</lpage>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Mo et al.,
          <year>2015</year>
          ] Mo,
          <string-name>
            <surname>H.</surname>
          </string-name>
          , et al.,
          <article-title>Desiderata for Computable Representations of Electronic Health Records-Driven Phenotype Algorithms</article-title>
          , JAMIA
          <volume>22</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1220</fpage>
          -
          <lpage>30</lpage>
          .,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Morley et al.,
          <year>2014</year>
          ] Morley,
          <string-name>
            <surname>K.</surname>
          </string-name>
          et al.,
          <article-title>Defining disease phenotypes using national linked electronic health records: a case study of atrial fibrillation</article-title>
          .
          <source>PLOS ONE</source>
          ,
          <volume>9</volume>
          (
          <issue>11</issue>
          ),
          <year>e110900</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Pikoula et al.,
          <year>2019</year>
          ] Pikoula,
          <string-name>
            <surname>M.</surname>
          </string-name>
          et al.,
          <article-title>Identifying clinically important COPD sub-types using data-driven approaches in primary care population based electronic health records</article-title>
          .
          <source>BMC Medical Informatics and Decision Making</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ), p.
          <fpage>86</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [
          <string-name>
            <surname>Pujades-Rodriguez</surname>
          </string-name>
          et al.,
          <year>2015</year>
          ]
          <article-title>Pujades-</article-title>
          <string-name>
            <surname>Rodriguez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.,
          <article-title>Heterogeneous Associations between Smoking and a Wide Range of Initial Presentations of Cardiovascular Disease in 1937360 People in England: Lifetime Risks and Implications for Risk Prediction</article-title>
          .
          <source>Int. J. of Epidemiology</source>
          <volume>44</volume>
          (
          <issue>1</issue>
          ):
          <fpage>129</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>