<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Initial Data-Driven Model for Estimating Impact of Antihypertensive Drug Amount on Blood Pressure Lowering?</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Almazov National Medical Research Centre</institution>
          ,
          <addr-line>St. Petersburg 197341</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ITMO University</institution>
          ,
          <addr-line>St. Petersburg 197101</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Due to the increasing popularity of clinical decision support systems, the problem of personalized drug dose identi cation becomes more relevant and substantial. In this paper, the authors introduce a data-driven model designed to operate in this case. Current work comprises general problem formulation, description of data and its preprocessing steps, model design overview, its rst stage model tuning and training, evaluation metrics used to estimate the quality, achieved values.</p>
      </abstract>
      <kwd-group>
        <kwd>Digital healthcare Personalized dose identi cation Classi cation algorithms Electronic health records Decision support systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Currently, decision support systems that consider patient characteristics are
gaining more popularity and impact on the process of treatment [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Individual treatment rule (ITR) that assigns an appropriate treatment to the speci c
patient based on his/her characteristics is one of decision support systems
important elements [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. An individual dosage rule (IDR) can be considered as a
part of ITR. It maximizes the expected treatment outcome for each patient by
de ning individual drug dosages.
      </p>
      <p>
        ITR was studied for various diseases, such as oncology [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or genome-guided
therapy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In this research, the ITR is constructed for patients with arterial
hypertension. A similar problem was considered in the case of antihypertensive
monotherapy, where the patients get treatment with a single drug class [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Copyright c 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
? The reported study was funded by RFBR according to the research project
#18-37-00441.</p>
      <p>
        This task of obtaining ITR was being solved in speci c areas with such
approaches as the Q-learning [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and O-learning (Outcome Weighted Learning) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
algorithms, and statistical random-e ects linear models [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>In this paper, the authors consider a supervised learning approach to obtain
ITR for personalized combined drug therapy.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem statement</title>
      <p>Given the vector Xj (ti) of patient pro le j in the moment of time ti before
treatment as:</p>
      <p>Xj (ti) = nx(jh)(ti)o
Obtain the set of antihypertensive therapy:</p>
      <p>Yj (Xj (ti)) = n(yj(k;1); yj(k;2))o;
(1)
(2)
where j = 1; m; i = 1; n; h = 1; p; k = 1; q and each drug yjk is a vector containing
a drug International Nonproprietary Name (INN) and optimal daily dosage. The
model, which authors propose in current paper, is designed to consist of three
data-driven submodels: the rst model receives a vector of patient features as an
input and predicts an optimal drugs count (nopt), the second model extends the
result specifying drug INNs nyj(k;1)onopt and the third model de nes the desired
daily dosages of each drug INN nyj(k;k2=) o1kn=op1t resulting with n(yj(k;1); yj(k;2))okn=op1t .
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data description</title>
      <p>The data used in this study were collected from 2010 to 2015, depersonalized
and provided by Almazov National Medical Research Centre.</p>
      <p>In the work, 16 features are grouped into a vector describing the patient
pro le. They are the following: age, sex, body mass index (BMI), systolic and
diastolic blood pressure before treatment, smoking status, impaired glucose
tolerance (IGT), left ventricular hypertrophy (LVH), chronic heart failure (CHF),
ischemic heart disease (IHD), dyslipidemia, diabetes, microalbuminuria,
cardiovascular diseases (CVD) among relatives, chronic kidney disease stage (CKD),
and combination of concomitant drugs.</p>
      <p>In the authors' previous study, eight patient clusters were been identi cation
based on the feature tuples. The patient groups accept the clinical
interpretation. As time passes, a patient pro le will be changing due to disease
development and ageing patient. It means that the same patient can belong to various
groups depending on the disease dynamics at di erent points in time. To
probabilistic model the arterial hypertension development in a patient, we use a
M arkov chain of transition from one cluster to another cluster. Therefore, the
dynamic model of the hypertensive patient transitions process from cluster i
(i = 1; 8) to cluster j (j = 1; 8) is presented as a graph, where the nodes are
clusters and the edges are transition probabilities Pij (Fig. 1). It's noted, the
probabilities Pij are determined in a way that the transition probabilities sum
is equal to one. Also, the patient condition will be able to remain the same then
the patient will transition to the same cluster. These transitions are presented
as a loop on between groups graph in Fig. 1. However, the certain clusters are
incompatible for relative transitions, in particular, due to gender characteristics
and/or the chronic concomitant diseases.</p>
      <p>0.22
Cluster 1</p>
      <p>
        Additionally, training dataset contains an outcome eld as a result of the
one-month treatment process. This value accepts various ways to be set. In this
particular research, it is based on clinical guidelines and points out if systolic and
diastolic blood pressure levels have reached the target values less than 140/90
mm Hg [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Since the task is to predict the optimal therapy for each patient based on
his/her features vector, such elds as drug names and daily dosages have to be
in the training data. So that, data was merged (on patient ID and outpatient
visit date) with records contained medical prescriptions for these patients. As
long as medical texts are written in natural language, they require additional
processing to distinguish desired information and ll these elds.</p>
    </sec>
    <sec id="sec-4">
      <title>Data preprocessing</title>
      <p>Due to the lack of training dataset, the processing task was resolved in this
research with a sequence of regular expressions and extracting rules
implementing so-called rule-based natural language processing (NLP). Below is the overall
pipeline that was applied to each outpatient visit:
1. Split the medical prescriptions eld into substrings with `nn' (newline)
delimiter. In the provided data most of such substrings contain, if they do,
only one prescription of the drug.
2. Find all entries of drugs in each substring using the dictionary prepared by
authors. This dictionary includes drug brand-names, their di erent writing
options that may show up in a natural language text, INNs, and
pharmacological classes.
3. The substring may have no medical prescriptions because it has general
guidance, referral to laboratory testing, etc. Such substring is not involved
in further processing.
4. If substring contains several drug brand-names, then distinguish them as an
alternative or as a combination. Patterns, which are used in this step, include
checking: their INNs { the same indicates the alternative, their location in
string boundaries, and the presence of `and' symbols, commas between them,
conjunctions.
5. Check the presence of words that mean cancellation, dosages and frequency
indicators. Substrings that don't have dosages and frequency can be involved
in the dataset with lling missing values using appropriate machine learning
algorithms in further preprocessing.
6. Extract dosages using regular expressions with measurement units,
frequencies { using regular expressions with parts of the day patterns.
7. Aggregate INNs of all extracted drugs in the INN eld, calculate their daily
dosages and write them in the Dosage eld using speci ed delimiter (in this
research `j').
5</p>
    </sec>
    <sec id="sec-5">
      <title>Classi er implementation</title>
      <p>This paper is aimed to describe the rst submodel in detail, which is proposed to
be an extension of a treatment outcome classi er. In the cycle, it concatenates the
vector of the patient features with every possible drug count g = 1; r and utilizes
the classi er to predict the probabilities of treatment ine ectiveness (negative
class, 0) and e ectiveness (positive class, 1) denoted as f(p(0); p(1))gg.</p>
      <p>The combination with the maximum outcome probability is assumed to
contain an optimal number of drugs.</p>
      <p>nopt = arg max (p(0); p(1))g
p(1)
(3)</p>
      <p>In the current Python implementation, several scikit-learn classi ers (library
version 0.21.3) were trained, tuned and evaluated, including C-Support Vector
Classi er (SVC), random forest classi er (RF), Multi-layer Perceptron (MLP),
classi er as well as LightGBM (library version 2.3.0).</p>
      <p>Hyper-parameters estimation using cross-validation led to the following
values (parameters not mentioned below are expected to have the default values
for the speci ed library version):
{ SVC: radial basis function (RBF) kernel, balanced class weights, enabled
probability estimation, kernel coe cient = 0:04, penalty parameter of the
error term C = 8:0.
{ Random forest classi er: entropy criterion, balanced class weights, with 150
trees in the forest of maximum depth 2 and 10 minimum samples to split.
{ MLP: stochastic gradient descent (sgd) solver, invscaling learning rate,
maximum number of iterations is 20, one hidden layer with 76 neurons,
regularization term is 0.2, using Nesterov's momentum, shu e samples in each
iteration.
{ LightGBM: random forest boosting type, balanced class weights, bagging
frequency is 1, bagging fraction is 0.9, learning rate is 0.01, number of
estimators is 110.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Assessment</title>
      <p>After the preprocessing step the dataset containing 4521 records were divided
into three datasets: training, validation, and test in a ratio of 0.63:0.27:0.1
respectively. It was decided to consider both Sensitivity (Recall) and Speci city
while estimating the treatment outcome classi ers parameters. This decision is
based on the requirement to e ciently identify both ine ective and e ective
treatment and use it in revealing the optimal drug amount.</p>
      <p>The rst two datasets were used in 5 splits with 7 repeats cross-validation,
results of which are presented in Table 1. Table 2 gives the results of classi ers
quality evaluation on the third (test) dataset. Although Sensitivity and
Specicity were considered as target metrics, the tables additionally include values of
such metrics as Accuracy, Precision, F1 score and ROC AUC (Receiver
Operating Characteristic Area Under ROC Curve).</p>
      <p>As can be seen from the tables that all classi ers avoided over tting. MLP
classi er performed worse than the others, which showed comparatively close
results. The random forest classi er is assumed to perform the most optimal
way.</p>
      <p>The most important patient features and their importances with more than
1% impact returned by random forest classi er are presented in Fig. 2. They
are calculated as the impurity decrease from each feature: the reduce in node
impurity weighted by the probability of reaching the node. It can be seen that
the drug count provides around 4.0% of the total decision.</p>
      <p>However, authors assume that using a bigger training dataset will improve
the results and the impact of the drug count feature, which is now considered
to be not su cient enough to reliably separate the e ective therapy from the
ine ective.</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion and Future works</title>
      <p>As a result of this work, the model that predicts the optimal antihypertensive
drug count was presented, implemented, trained, and evaluated. This model is
a part of the proposed general model predicting the optimal antihypertensive
drug dosages based on the patient features.</p>
      <p>Future works of this research include:
{ preparation and preprocessing of new data collected from 2016 to 2019 that
will be provided by Almazov National Medical Research Centre;
{ further training and parameters tuning of additional classi ers predicting the
optimal amount of prescription drugs for a patient with arterial hypertension;
{ development of the data-driven model predicting the most e ective
individual antihypertensive therapy including drug INNs and daily dosages.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements References</title>
      <p>The reported study was funded by RFBR according to the research project
#18-37-00441.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Somogyi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McMichael</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baranzini</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mousavi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greller</surname>
          </string-name>
          , L.D.:
          <article-title>10 Advanced data mining and predictive modelling at the core of personalised medicine</article-title>
          .
          <source>Studies in Multidisciplinarity 3</source>
          ,
          <issue>165</issue>
          {
          <fpage>192</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Darwich</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogungbenro</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinks</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Powell</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reny</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marsousi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daali</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fairman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lesko</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCune</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knibbe</surname>
            ,
            <given-names>C.A.J.</given-names>
          </string-name>
          , de Wildt,
          <string-name>
            <given-names>S.N.</given-names>
            ,
            <surname>Leeder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.S.</given-names>
            ,
            <surname>Neely</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Zuppa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.F.</given-names>
            ,
            <surname>Vicini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Aarons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.N.</given-names>
            ,
            <surname>Boiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Rostami-Hodjegan</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Why has model-informed precision dosing not yet become common clinical reality? lessons from the past and a roadmap for the future</article-title>
          .
          <source>Clin. Pharmacol. Ther</source>
          .
          <volume>101</volume>
          (
          <issue>5</issue>
          ),
          <volume>646</volume>
          {
          <fpage>656</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Barbolosi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciccolini</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lacarelle</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barlesi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andre</surname>
          </string-name>
          , N.:
          <article-title>Computational oncology-mathematical modelling of drug regimens for precision medicine</article-title>
          .
          <source>Nat Rev Clin Oncol</source>
          <volume>13</volume>
          (
          <issue>4</issue>
          ),
          <volume>242</volume>
          {
          <fpage>254</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bielinski</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pathak</surname>
          </string-name>
          , J.:
          <article-title>Preemptive genotyping for personalized medicine: Design of the right drug, right dose, right timedusing genomic data to individualize treatment protocol</article-title>
          .
          <source>Mayo Clinic Proceedings</source>
          <volume>89</volume>
          (
          <issue>1</issue>
          ),
          <volume>25</volume>
          {
          <fpage>33</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Semakova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zvartau</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bochenina</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konradi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Towards Identifying of E ective Personalized Antihypertensive Treatment Rules from Electronic Health Records Data Using Classi cation Methods: Initial Model</article-title>
          . In: Procedia Computer Science, pp.
          <volume>852</volume>
          {
          <fpage>858</fpage>
          .
          <string-name>
            <surname>Elsevier</surname>
            <given-names>B.V.</given-names>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Moodie</surname>
            ,
            <given-names>E.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kramer</surname>
            ,
            <given-names>M.S.:</given-names>
          </string-name>
          <article-title>Q-learning for estimating optimal dynamic treatment rules from observational data</article-title>
          .
          <source>Can J Stat</source>
          <volume>40</volume>
          (
          <issue>4</issue>
          ),
          <volume>629</volume>
          {
          <fpage>645</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kosorok</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          :
          <article-title>Personalized Dose Finding Using Outcome Weighted Learning</article-title>
          .
          <source>J Am Stat Assoc</source>
          <volume>111</volume>
          (
          <issue>516</issue>
          ),
          <volume>1509</volume>
          {
          <fpage>1521</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Diaz</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeh</surname>
          </string-name>
          , H.W., de Leon, J.:
          <article-title>Role of Statistical Random-E ects Linear Models in Personalized Medicine</article-title>
          .
          <source>Curr Pharmacogenomics Person Med</source>
          <volume>10</volume>
          (
          <issue>1</issue>
          ),
          <volume>22</volume>
          {
          <fpage>32</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>