<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Tree-based Regularization for Interpretable Readmission Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jialiang Jiang and Varun Chandola</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sharon Hewner</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science and Engineering, University at Buffalo</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Nursing, University at Buffalo</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Preventable hospital readmissions have been identified as
one of the primary targets for reducing costs and
improving healthcare delivery. However, most data driven studies for
understanding readmissions have produced non-interpretable
black boxes, which precludes them from being used
effectively within the decision support systems in the hospitals. A
novel strategy to improve the interpretability of a linear model
by incorporating domain knowledge is proposed here. The
central idea is to exploit the hierarchical relationships among
the features (medical diagnosis codes, in this case) using
a tree-structured sparsity-inducing regularization norm. The
proposed method transforms the hierarchical relations among
features into a graph and then applies graph-guided
regularization during the model learning. Additionally, an evaluation
metric is proposed to quantify the interpretability of a linear
model with respect to the domain hierarchy. Results on two
healthcare claims data sets are shown, where a model is learnt
to predict a patient’s risk of readmission, based on the
medical history and other relevant features. Results show that the
proposed method is able to learn a model which can predict
readmission risk with accuracies that are comparable to
existing methods, but produces a highly interpretable output,
which allows medical experts to draw clinically relevant
insights and identify key factors associated with hospital
readmissions. Some of these factors conform to existing beliefs,
e.g., impact of surgical complications and infections during
hospital stay. Other factors, such as the impact of mental
disorder and substance abuse on readmission, provide empirical
evidence for several pre-existing but unverified hypotheses.
The findings of this study will be instrumental in designing
the next generation decision support systems for preventing
readmissions.
Hospital readmissions are prevalent in the healthcare
system and contribute significantly to avoidable costs. In United
States, recent studies have shown that the 30-day
readmission rate among the Medicare beneficiaries1 is over 17%,
Copyright held by the author(s). In A. Martin, K. Hinkelmann, A.
Gerber, D. Lenat, F. van Harmelen, P. Clark (Eds.), Proceedings of
the AAAI 2019 Spring Symposium on Combining Machine
Learning with Knowledge Engineering (AAAI-MAKE 2019). Stanford
University, Palo Alto, California, USA, March 25-27, 2019.</p>
      <p>
        1A federally funded insurance program representing 47.2 %
($182.7 billion) of total aggregate inpatient hospital costs in the
with close to 75% of these being avoidable
        <xref ref-type="bibr" rid="ref22">(Mpa 2007)</xref>
        ,
with an estimated cost of $15 Billion in Medicare
spending. Similar alarming statistics are reported for other private
and public insurance systems in the US and other parts of
the world. In fact, management of care transitions to avoid
readmissions has become a priority for many acute care
facilities as readmission rates are increasingly being used as a
measure of quality
        <xref ref-type="bibr" rid="ref6">(Conway and Berwick 2011)</xref>
        .
      </p>
      <p>
        Given that the rate of avoidable readmission has now
become a key measure of the quality of care provided in
a hospital, there have been increasingly large number of
studies that use healthcare data for understanding
readmissions. Most existing studies have focused on building
models for predicting readmissions using a variety of available
data, including patient demographic and social
characteristics, hospital utilization, medications, procedures, existing
conditions, and lab tests
        <xref ref-type="bibr" rid="ref10 ref5">(Futoma, Morris, and Lucas 2015;
Choudhry et al. 2013; Donze et al. 2013)</xref>
        . Other
methods use less detailed information such as insurance claim
records
        <xref ref-type="bibr" rid="ref11 ref30">(Yu et al. 2013; He et al. 2014)</xref>
        . Many of these
methods use machine learning methods, mainly Logistic
Regression, to build classifiers and have reported consistent
performance on a variety of clinical data sets. In fact, most papers
about readmission prediction report AUC scores in the range
of 0.65-0.75.
      </p>
      <p>While the predictive models have shown promise, their
moderate performance means that they are still not at a
stage where hospitals could use them as “black-box”
decision support tools. Moreover, such models are not easily
interpretable to provide actionable insights to the decision
makers. At the same time, beyond the selection of the initial
set of features to learn from, these solutions do not explicitly
utilize the rich information available in the medical domain.</p>
      <p>
        In this paper, we explore incorporation of one such
domain information, into the model learning process.
Specifically, we utilize the hierarchical relationships among
different medical diagnosis codes, available as a taxonomical tree
(See Section 3 for details). The tree structure is utilized as
a regularization penalty, to enforce the model (logistic
regression) to learn a sparse solution such that the non-zero
weights are localized within a few sub-trees. The key idea
is that such a solution would be easier to interpret compared
to a solution in which the weights are “scattered” across. A
graphical illustration is provided in Figure 1. The proposed
method falls under the general class of structured
sparsity regularization based machine learning models (Mosci
et al. 2010), which consists of numerous schemes to
exploit different types of relationships among features,
including groups
        <xref ref-type="bibr" rid="ref31">(Yuan and Lin 2006)</xref>
        , sequential
        <xref ref-type="bibr" rid="ref27">(Tibshirani et
al. 2005)</xref>
        , and graphs (Chen et al. 2010). However,
regularization methods for scenarios where the features are related
over a tree are sparse, and the existing ones provide an
indirect way of capturing the tree structure
        <xref ref-type="bibr" rid="ref32">(Zhao, Rocha, and Yu
2009)</xref>
        , which, as observed later in the experiments, makes
them inadequate for the target problem of readmission
prediction.
      </p>
      <p>
        The proposed regularization scheme transforms the tree
structure into a weighted graph that uses the “tree-distance”
as the weight of the edge between the corresponding nodes
in a graph, and then employs a graph based penalty to force
the machine learning algorithms to favor solutions in which
the non-zero weights are strongly linked in the graph. The
regularizer is incorporated into a standard logistic
regression classifier, using truncated gradient descent
        <xref ref-type="bibr" rid="ref18">(Langford,
Li, and Zhang 2008)</xref>
        for the optimization step. This is used
to learn a readmission prediction model that uses
diagnosis codes from a patient’s medical history to predict his or
her readmission risk, as a binary label. Results on two data
sets, extracted from: 1). New York State Medicaid records
(MDW), and 2). MIMIC-III data set (a publicly available
data set), show that the proposed model not only performs
comparably, in terms of accuracy, to classical regularization
schemes such as LASSO and existing tree-based
regularizer
        <xref ref-type="bibr" rid="ref32">(Zhao, Rocha, and Yu 2009)</xref>
        , but learns a sparse model
that is significantly better than others in terms of the
interpretability. To this effect, we propose a quantitative metric
to assess the interpretability of a model in which the features
are arranged in a tree structure.
      </p>
      <p>
        By analyzing the model trained on the MDW data, we
infer several important insights to improve the
understanding of readmissions. Some of our findings conform to
existing beliefs, for example, the importance of bacterial
infections during hospital stay. Other findings provide empirical
evidence to support existing hypotheses amongst healthcare
practitioners, for example, the effect of the type of insurance
on readmissions
        <xref ref-type="bibr" rid="ref12">(Hewner et al. 2014)</xref>
        . Most interesting
findings from our study reveal surprising connections between
a patient’s non-disease background and the risk of
readmission. These include behavioral patterns (mental disorders,
substance abuse) and socio-economic background. For the
result of the analysis of MIMIC-III data, it also has similar
inference. For example, bacterial infections during hospital
stay, chronic circulatory and respiratory system diseases are
important factors to predict readmission. Moreover, profited
by semantic refining ability of tree-based regularization, We
can infer the significant disease classification to readmission
straightforward. For example, class of Diseases Of The
Circulatory System and Metabolic Disorders are highlighted for
understanding of readmissions.
      </p>
      <p>We believe that such findings can have a significant
impact on how healthcare providers develop effective
strate(a) No taxonomy-guided regularization</p>
      <p>(b) Taxonomy-guided regularization
gies to reduce readmissions. At present, the healthcare
efforts in this context have been twofold. First is the effort
to improve the quality of care within the hospital and the
second is to develop effective post-discharge strategies such
as telephone outreach, community-based interventions, etc.
The results from this study inform the domain experts on
both fronts.</p>
      <p>The rest of the paper is organized as follows. We review
existing literature on readmission prediction in Section 2.
We describe the data used for our experiments in Section 3
and formulate the machine learning problem in Section 4.
We discuss the classification methodology in Section 5. We
present the algorithm for measurement of model’s
interpretability in Section 6. The results are presented in
Section 7. We discuss the importance of interpretable outcome
by including a real world case study in Section 8.
Infectious &amp; Parasitic</p>
      <p>Intestinal Infections
Tuberculosis
Zoonotic Bacterial
Infections
: : :
Neoplasms
Malignant (Lip,
Oral Cavity, : : :)
Malignant
(Digestive)
: : :
: : :</p>
      <p>Injury &amp; Poisoning</p>
      <p>
        External cause
status
Activity
Railway Accidents
: : :
Coincident with the rising importance of readmissions in
reducing healthcare costs, there have been several papers
that use clinical and insurance claims information to build
predictive models for readmissions
        <xref ref-type="bibr" rid="ref17">(Kansagara et al. 2011)</xref>
        using different machine learning models including Deep
Neural Networks
        <xref ref-type="bibr" rid="ref14 ref19">(Jamei et al. 2017; Lin et al. 2018;
Xiao et al. 2018)</xref>
        , Logistic Regression
        <xref ref-type="bibr" rid="ref10 ref23 ref5">(Futoma, Morris,
and Lucas 2015; Choudhry et al. 2013; Donze et al. 2013;
Niu 2013)</xref>
        and Support Vector Machines
        <xref ref-type="bibr" rid="ref30">(Yu et al. 2013)</xref>
        .
However, most of these solutions have focused on improving
the accuracy of the predictive model, and not necessarily on
the interpretability of the model to improve the
understanding of the readmission problem. Papers that focus on
interpretability are limited to identifying the best features that
predict readmission
        <xref ref-type="bibr" rid="ref15">(Jiang et al. 2018)</xref>
        and have typically
focused on a small set of patients or hospitals
        <xref ref-type="bibr" rid="ref1 ref30">(Yu et al. 2013;
Amarasingham et al. 2010)</xref>
        . In this paper, we are focusing
on a more direct approach that is scalable to any problem
setting.
      </p>
      <p>
        Finally, the hierarchical relationship has never been
exploited for building predictive models for readmission.
Singh, et. al,
        <xref ref-type="bibr" rid="ref25">(Singh et al. 2014)</xref>
        have presented a similar
approach in the context of predicting disease progression,
however, the authors focus on using the disease hierarchy
to come up with new features that are fed into the
classifier. Additionally, there is no standard of measurement for
interpretability of prediction model especially for structured
based, while we propose a general methodology to address
the problem.
      </p>
      <p>3</p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>For the experiments, we explored two different data sets that
consist of healthcare insurance claims and electronic health
records (EHR).</p>
      <p>The fist dataset is obtained from the New York State
Medicaid Data Warehouse (MDW). Medicaid is a social health
2http://www.icd9data.com/2015/Volume1/
default.htm
care program for families and individuals with low income
and limited resources. We analyzed four years (2009–2012)
of claims data from the MDW. The claims correspond to
multiple types of health utilization including
hospitalizations, outpatient visits, etc. While the raw data consisted of
4,073,189 claims for 352,716 patients, we only included the
patients in the age range 18–65 with no obstetrics related
hospitalizations. The number of patients with at least one
hospitalization who satisfied these conditions were 11,774
and had 34,949 claims.</p>
      <p>For each patient we have information of patient’s
admission medical history extracted from four years of claims data
represented as a binary vector that indicates if the patient
was diagnosed with a certain disease in the last four years.</p>
      <p>
        The second dataset is Multi-parameter Intelligent
Monitoring in Intensive Care (MIMIC-III) public dataset
        <xref ref-type="bibr" rid="ref16">(Johnson et al. 2016)</xref>
        . This data is a large, freely-available
database comprising de-identified health-related data
associated with over forty thousand patients who stayed in
critical care units of the Beth Israel Deaconess Medical Center
between 2001 and 2012.
      </p>
      <p>While the database includes information such as
demographics, vital sign measurements made at the bedside (one
data point per hour), laboratory test results, procedures,
medications, caregiver notes, imaging reports, and mortality
(both in and out of hospital), we focus on admission records
to extract the medical codes as part of each patient’s history.</p>
      <p>According to the dataset we retrieved from the
MIMICIII dataset, there are 46516 patients in total with 3996 of
patients flagged as readmissions. The medical history for each
patient consists of 6783 diagnosis codes.</p>
      <sec id="sec-2-1">
        <title>Diagnosis Codes</title>
        <p>Disease information is encoded in insurance claims and
medical records using diagnosis codes. The International
Classification of Diseases (ICD) is an international
standard for classification of disease codes. The data used in
this paper followed the ICD-9-CM classification which is a
US adaptation of the ICD-9 classification. Conceptually, the
ICD-9-CM codes are structured as a tree (See Figure 2 for a
sample) with 19 broad disease categories at level 1. The
entire tree has 5 levels and has total of 14,567 diagnosis codes.
While the primary purpose of ICD taxonomy has been to
support the insurance billing process, it contains a wealth of
domain knowledge about the different diseases.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Readmission Risk Flag</title>
        <p>For each patient in the above described cohort, we assign a
binary flag for readmission risk. The readmission risk flag
is set to 1 if the patient had at least one pair of
consecutive hospitalizations within 30 days of each other in a single
calendar year, otherwise it is set to 0.</p>
        <p>4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Problem Statement</title>
      <p>
        Given a patient’s disease history, we are interested in
predicting the readmission risk (binary flag) for the patient. The
problem formulation is different from many existing
studies
        <xref ref-type="bibr" rid="ref10">(Futoma, Morris, and Lucas 2015)</xref>
        , where the focus is
on assigning a readmission risk to a single hospitalization
event. Our focus is on understanding the impact of
socioeconomic and behavioral factors on a readmission.
      </p>
      <p>We denote each patient i as a vector xi consisting of
11,881 elements for the MDW dataset and 6,873 elements
for the MIMIC-III admission dataset corresponding to the
number of disease codes showed in data respectively. Note
that while ICD-9-CM classification contains 14,567 codes,
only 11,881 and 6,783 codes are observed in each data set
used in this paper. We selected patients that age in between
18 and 65 and excluded pregnancy related diseases. All
elements in the vector are binary. The readmission risk flag
is denoted using yi 2 f0; 1g where 1 indicates readmission
risk.</p>
      <p>From machine learning perspective our task is to learn a
N
classifier from a training data set hxi; yiii=1 which can be
used to assign the readmission risk flag to a new patient
represented as x . Note that the input vector xi is highly sparse.
For example, in the NY Medicaid dataset, on average, there
are only 36 non-zeros out of total 11,884 possible codes.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <p>
        We use a logistic regression (LR) model
        <xref ref-type="bibr" rid="ref8">(Cox 1958)</xref>
        as the
classifier, which, is the most widely used model in the
context of readmission prediction
        <xref ref-type="bibr" rid="ref10">(Futoma, Morris, and Lucas
2015)</xref>
        . The LR model, for binary classification tasks,
computes the probability of the target y to be 1 (readmission
risk), given the input variables, x as:
p(y = 1jx) =
      </p>
      <p>1
1 + exp(
&gt;x)
(1)
Where is the LR model parameter (regression
coefficients). We assume that x includes a constant term
corresponding to the intercept.</p>
      <p>The model parameter are learnt from a training data set</p>
      <p>N
(hxi; yiii=1) by optimizing the following objective function:</p>
      <p>N
^ = arg min X log(1 + exp( yi &gt;xi)) +
( ) (2)
k Gij k1
where G denotes the tree constructed using the hierarchy of
the diagnosis codes. Gij denotes the jth node in the tree at
the ith level. Thus G01 denotes the root level, and so on.
( ) =</p>
      <p>1</p>
      <sec id="sec-4-1">
        <title>A Novel Sparse Tree-Structure Regularizer</title>
        <p>The TSGL regularizer, discussed above, treats the tree
structure as a special overlapped group, which ignores the hidden
relationship between nodes at different levels. To overcome
this deficiency, we propose a different way to exploit the tree
structure. The new regularization penalty is defined as:
m m
X X
Where the i and i are coefficients of features i and j,
respectively and Dij is the tree distance between features i
and j and will be introduced in next subsection. The first
penalty term ensures that the selected features are closer to
each other in the taxonomy tree while the second term, k k1,
ensures the overall sparsity of the solution.
where the first term refers to the training loss and the second
terms is a regularization penalty imposed on the solution;
being the regularization parameter.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Existing Regularization Schemes</title>
        <p>
          L1 Regularizer Different forms of regularization
penalties have been used in the past, including the widely used
l1 and l2 norms
          <xref ref-type="bibr" rid="ref28">(Tibshirani 1994)</xref>
          . While l2 norm ( ( ) =
k k2 = (Pj j2)1=2) is typically used to ensure stable
results, l1 norm ( ( ) = k k1 = Pj j j j) is used to promote
sparsity in the solution, i.e., most coefficients in are 0.
        </p>
        <p>
          However, l1 regularizer does not explicitly promote
structural sparsity. Given that the features used in predicting
readmission risk have a well-defined structure imposed by the
ICD-9 standards, we explore regularizers that leverage this
structure for model learning:
Sparse Group Regularizer This regularizer (also referred
to as Sparse Group LASSO or SGL) assumes that the input
features can be arranged into K groups (non-overlapping or
overlapping)
          <xref ref-type="bibr" rid="ref2">(Bach 2008)</xref>
          . The SGL regularizer is given by:
(3)
(4)
( ) =
k k1 + (1
        </p>
        <p>K
) X k Gk k2
k=1
where Gk are the coefficients corresponding to the group
Gk. The above penalty function favors solutions which
select only a few groups of features (group sparsity). For the
task of readmission prediction, we divide the features
corresponding to all numbers of diagnosis codes included into 19
non-overlapping groups, based on the top level groupings in
the ICD-9-CM classification (See Table 1).</p>
        <p>Tree Structured Group Regularizer This regularizer,
also referred to as Tree Structured Group LASSO (TSGL),
explicitly uses the hierarchical structure imposed on the
features. The TSGL regularizer is given by:
Distance Matrix for Tree The distance between any two
nodes in the tree, Dij is defined in terms of the length of the
path between the two nodes. If the node i is an “ancestor” of
node j, or vice-versa, the distance is defined as:</p>
        <p>Dij = (li
lj )2
where li denotes the level or the number of steps from the
root for node i. If the nodes i and j do not share any ancestral
relationship, then the distance is defined as:</p>
        <p>Dij = ((li
lc)2 + (lj
lc)2)3
where c is the node that is the nearest common ancestor for
nodes i and j. The cubic power is used to sharply increase
the weight with the number of levels to go up by to find
the common ancestor. Thus, Dij will be largest for two leaf
nodes whose common ancestor is the root node.</p>
        <p>The data matrix consisting of the distances between all
pairs of leaf nodes (features) from the MIMIC-III data set is
shown in Figure 3.</p>
        <p>Optimization To solve the optimization problem in (2)
using the regularization penalty defined in (5), we first convert
the first penalty term (tree structure sparsity) into a graph
constraint as:
where L is a m</p>
        <p>
          m matrix, such that:
m m
X X
Truncated Gradient Descent Note that the objective
function and the tree penalty term have a convex form such
that one can calculate the gradient of the objective function
with respect to the weight vector, , and use that within a
gradient descent algorithm. However, due to the presence
of the l1 term (k k1), a direct gradient descent
formulation is not possible. We employ Truncated Gradient
Descent
          <xref ref-type="bibr" rid="ref18">(Langford, Li, and Zhang 2008)</xref>
          which has been shown
to be effective in learning solutions under l1 regularization
penalties.
        </p>
        <p>The idea behind truncated gradient descent is to ignore
the l1 term when calculating the gradient, and round small
coefficients (that are not larger than a small threshold) to
zero after every k online steps, i.e., at every kth step:
(t) = T0( (t 1)
rJ )
Where rJ is the gradient without the l1 penalty term and
T0 defined by:</p>
        <p>T0( j ; ) =
0 if j j j &lt;
j
otherwise</p>
        <p>That is, we first apply the standard gradient descent rule,
and then round small coefficients to zero to enforce the l1
sparsity.
(11)
(12)</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>A Quantitative Measure for</title>
    </sec>
    <sec id="sec-6">
      <title>Interpretability</title>
      <p>In classical sparsity inducing models, sparsity is measured
using the number of non-zero coefficients or weights. While
this is reasonable for settings with “flat” structure, e.g.,
LASSO or Group LASSO, this does not reveal the true
interpretiveness of a solution, in the context of a tree structure.
For instance, Figure 1 illustrates how two solutions with
same number of non-zero coefficients can have different
interpretability.</p>
      <p>We propose a novel measure to assess the
interpretiveness of a solution. The proposed measure is calculated in a
bottom-up fashion, starting from the leaf nodes. For the ith
node in the tree, we define the purification noise, Pi as:
Pi = Ei + X</p>
      <p>Pj
j2Ci
where Ci is the set of non-leaf children of node i and Ei is
the Shannon Entropy of the current ith node by measuring
the information loss of all leaf child nodes, i.e., children that
are actual features:</p>
      <p>Ei =</p>
      <p>pi log2(pi)
where pi is the fraction of leaf children of node i with a
nonzero coefficient. Starting from the parents of the leaf nodes,
the purification noise is recursively computed, and finally
the purification noise for the root node, i.e., Proot is treated
as the overall purification noise for the entire solution.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Results</title>
      <p>
        In this section we present our findings by applying
logistic regression classifier for the task of readmission
prediction on the MDW data and MIMIC-III data described
earlier. We first compare the performance of the different
regularization strategies to the classification task using the area
under the ROC-curve (AUC) for each classifier as our
evaluation metric due to the imbalance of data. We also compare
the different strategies to report the purification noise
(interpretability score) value for each solution. For each strategy,
we run 10 experiments with random 80-20 splits for
training and test data, respectively. The optimal values for the
regularization parameters for each strategy are chosen using
cross-validation. We use the MATLAB package, SLEP
        <xref ref-type="bibr" rid="ref20">(Liu,
Ji, and Ye 2009)</xref>
        , for the Tree Structured Group
Regularization experiments. The proposed regularization method was
developed in Python.
7.1
      </p>
      <sec id="sec-7-1">
        <title>Comparing Different Regularization</title>
      </sec>
      <sec id="sec-7-2">
        <title>Strategies</title>
        <p>Here we compare the performance of different
regularization methods discussed in Section 5. The results are
summarized in Table 2 and Figure 4.</p>
        <p>For the MIMIC-III data set, the best performance, in terms
of AUC is obtained using the classical, l1 and l2
regularizations. However, the interpretability is highest for the
proposed tree-structured measure, followed by the earlier
published TSGL algorithm. On the other hand, the results for
(13)
(14)
the MDW data set show that the tree based regularizations
perform on par with the classical methods. However, the
interpretability is significantly higher for the proposed
treebased regularization scheme. The l2 regularizer, for obvious
reasons, does not produce a sparse solution (194.85 of MDW
and 716.48 of MIMIC-III), while the other three regularizers
induce significant sparsity. However, the structured
regularizers are able to achieve significantly low structured sparsity
(30.12 of MDW and 21.69 of MIMIC-III) which is
consistent with the ICD-9-CM hierarchy.</p>
      </sec>
      <sec id="sec-7-3">
        <title>Effect of Regularization Parameter, 1 The role of the</title>
        <p>regularization parameter, 1, in (10) is to control the penalty
on the tree-structure of the solution. Figure 4 shows how
the AUC score and the interpretability score vary with 1.
By increasing 1, we note significant improvement of
performance by leveraging more prior hierarchical information
as well as outstanding decrease on purification noise, which
indicates highly interpretability.
...
...
...
...
(451­459)  
Diseases Of Veins
&amp; Lymphatics, &amp;
Other Diseases Of</p>
        <p>Circulatory
System 
...</p>
        <p>453  
Other venous
embolism and
thrombosi 
453.8 
Acute venous
embolism &amp;
thrombosis of
other specified
veins 
458 
Hypotension
458.9 
Hypotension,
unspecified 
...</p>
        <p>(580­629)  
Diseases Of The
Genitourinary</p>
        <p>System
(580­589)  
Nephritis, Nephrotic</p>
        <p>Syndrome, And</p>
        <p>Nephrosis 
...</p>
        <p>584 
Acute kidney
failure 
584.5 
Acute kidney
failure with
lesion of tubular
necrosis 
584.9 
Acute kidney
failure,
unspecified </p>
        <p>585 
Chronic kidney
disease (ckd)
585.6 
End stage
renal disease </p>
        <p>585.9 
Chronic kidney
disease,
unspecified 
...</p>
        <p>(760­779) </p>
        <p>Certain Conditions
... Originating In The ...</p>
        <p>Perinatal Period
 
(800­999) 
 Injury And
Poisoning</p>
        <p>(V01­v91)  
Supplementary Class Of
Factors Influencing Health
Status &amp; Contact With</p>
        <p>Health Services 
...
...</p>
        <p>...
...</p>
      </sec>
      <sec id="sec-7-4">
        <title>7.2 Qualitative Interpretation of the Solution</title>
        <p>
          Figure 5 provides a graphical illustration of the non-zero
features learnt by the proposed model. The red nodes
denote the actual diagnosis codes that had non-zero weights.
The blue nodes are the ancestors of the selected leaf nodes.
We first note that the model consists of only 87 (out of
11567) features or disease codes. Additionally, the 87 codes
are concentrated within a few higher level disease
categories. Since the sub-trees containing the non-zero
features are relatively dense, it is easy to summarize them
with a higher order disease category. For instance, one can
determine that Disorders of fluid electrolyte
and acidbase balance is an important factor in
determining readmissions, which has been confirmed in actual
clinical studies
          <xref ref-type="bibr" rid="ref3">(Badawi and Breslow 2012)</xref>
          . Similarly,
pregnancy related complications are an important factor in
determining readmissions.
        </p>
        <p>On the other hand, a similar visualization for l1
regularization shows a highly non-interpretable result, as seen in
tracheal or bronchial tuberculosis ...,
which has a non-zero weight. It is unclear why only that
diagnosis code is selected whereas the other 35 siblings
(other forms of tuberculosis) are ignored.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>8 Discussion</title>
      <p>Section 7 shows that leveraging the hierarchical
information in the ICD-9CM classification improves the predictive
capability of logistic regression while gaining better
interpretability. In this section, we demonstrate the necessity
of high level interpretation to a medical record prediction
model from the healthcare perspective. The focus is to show
that an interpretable decision support tool for readmission
risk prediction can be effective, as shown in the following
real world case study.</p>
      <p>As shown in Figure 5, the model is well-informed by
the ICD-9-CM hierarchy. Interpretable learning grants the
model the ability to conclude high-level important category
that are more understandable to the healthcare providers in
the medical facilities. For example, the result of proposed
model suggests 428-Heart Failure as a important disease
category results in readmission instead of extremely specific
ICD-9-CM disease code like 428.33-Acute on Chronic
Diastolic Heart Failure. The high level important concept
ab...</p>
      <p>765 
Disorders relating
to short gestation
and low
birthweight 
765.18 
Other preterm
infants, 2,000­
2,499 grams </p>
      <p>765.19 
Other preterm
infants, 2,500
grams and over </p>
      <p>765.27 
33­34 completed
weeks of
gestation 
765.28 
35­36 completed
weeks of
gestation </p>
      <p>770 
Other respiratory
conditions of fetus
and newborn
770.6 
Transitory
tachypnea of
newbor 
770.81 
Primary apnea
of newborn 
770.89 
Other respiratory
problems after
birth 
...</p>
      <p>...</p>
      <p>995 
Certain adverse
effects not
elsewhere
classified 
995.91 
Sepsis 
995.92 
Severe sepsis 
...</p>
      <p>(996­999)  
Complications Of
Surgical And
Medical Care,
Not Elsewhere
Classified </p>
      <p>996 
Complications
peculiar to
certain specified
procedures 
996.62 
Infection and
inflammatory
reaction due to
other vascular</p>
      <p>device,
implant, and
graft 
997 
Complications</p>
      <p>affecting
specified body
system not
elsewhere
classified 
997.1 </p>
      <p>Cardiac
complications,
not elsewhere
classified 
998 </p>
      <p>Other
complications
of procedures
not elsewhere
classified 
998.59 </p>
      <p>Other
postoperative
infection 
...</p>
      <p>(V50­V59)  </p>
      <p>Persons
Encountering</p>
      <p>Health
Services For
Specific
Procedures
And Aftercare 
(V40­V49)  
Persons With
A Condition
Influencing
Their Health</p>
      <p>Status 
V44 
Artificial
opening
status </p>
      <p>V44.0 
Tracheostomy
status </p>
      <p>V44.1 
Gastrostomy
status </p>
      <p>V58 
Encounter for
other and
unspecified
procedures
and aftercare </p>
      <p>V58.61 
Long­term
(current) use</p>
      <p>of
anticoagulants </p>
      <p>V58.67 
Long­term
(current) use
of insulin 
...</p>
      <p>Other 14 categories
not selected 
(010­018) 
Tuberculosis 
(710­739)  
Diseases Of The
Musculoskeletal</p>
      <sec id="sec-8-1">
        <title>ConSnyesctteimve A Tnisdsue </title>
        <p>(710­­719)  
Arthropathies And
Related Disorders 
720­724 
(725­­729)  </p>
        <p>Rheumatism,
Excluding The Back 
730­739 </p>
        <p>719.44 
Pain in joint, hand </p>
      </sec>
      <sec id="sec-8-2">
        <title>Others e8l9e cctoedde  s not</title>
        <p>728 
Disorders of muscle
ligament and fascia 
Other 4 categories
not selected 
Other 8 categories
not selected 
717 </p>
        <p>Internal
derangement of
knee 
718 </p>
      </sec>
      <sec id="sec-8-3">
        <title>Other odfe rjoainngte ment</title>
        <p>719 
Other and
unspecified
disorders of joint 
Contruap7cpt1ue8rr.e 4a 2orm f j oint, lDateeurrana7nsl1pg m7eec.me4inf0eiien sdct u osf,
718.90 
deranUgnesmpeencitf ioefd joint, Others e1l9e cctoedde  s not
site unspecified </p>
      </sec>
      <sec id="sec-8-4">
        <title>Others e8l8e cctoedde  s not</title>
        <p>728.83 
Rupture of muscle,
nontraumatic </p>
        <p>Text</p>
      </sec>
      <sec id="sec-8-5">
        <title>Others e2l2e cctoedde  s not</title>
        <p>(800­999) 
&amp;nbsp;Injury
And Poisonin..g.</p>
        <p>(805­­809)  
Fracture Of Spine
And Trunk </p>
        <p>805 </p>
        <p>Fracture of
vertebral column
without mention of
spinal cord injury 
Other 4 categories
not selected </p>
        <p>805.14 
Open fracture of
fourth cervical</p>
        <p>vertebra </p>
      </sec>
      <sec id="sec-8-6">
        <title>Others e2l5e cctoedde  s not</title>
        <p>Other
15 categories
not selected 
(960­­979)  
Poisoning By Drugs,
Medicinals And</p>
        <p>Biological
Substances </p>
        <p>969 
Poisoning by
psychotropic</p>
        <p>agents 
Other 19 categories
not selected </p>
        <p>969.6 
Poisoning by
psychodysleptics
(hallucinogens) </p>
      </sec>
      <sec id="sec-8-7">
        <title>Others e1l9e cctoedde  s not</title>
        <p>Other 22 categories
not selected 
straction makes it easier to focus on important aspects while
removing excessive attention from specific disease codes,
that may distract the healthcare staff from the key factors
during the post-discharge phase.</p>
        <p>Case Study: M.J. was a 55-years-old white male
with medical history of Hypertension, Coronary
Artery Disease and High Cholesterol, and came to the
emergency department with left flank pain on 7/1/2018.
He was admitted to the hospital for left kidney stone
and treated with intravenous fluid and pain
medication. M.J. was also found of left hydro nephrosis due to
obstructing kidney stone and M.J. underwent urology
surgery to have the stone removed on 7/4/2018.
However, his hospital course was complicated with sepsis
(urinary tract infection) and electrolyte imbalance
(hyperkalemia) due to acute renal injury. M.J. finished the
course of antibiotics for urinary tract infection and his
renal function was back to the baseline. After a 2-week
hospital stay, the patient was debilitated so he was
discharged to a skilled nursing facility on 7/14/2018 for
rehabilitation.</p>
        <p>On 7/21/2018, M.J. was found unresponsive and
pulse-less by the staff in the skilled nursing facility. He
was resuscitated by emergency medical services and
readmitted to the intensive care unit for cardiac arrest
due to electrolyte imbalance (hyperkalemia). Staff
reported that he was not eating or drinking since being
discharged to the skilled nursing facility on 7/14/2018.
During the ICU, M.J. developed multi-organ failure
and died in the intensive care unit on 7/31/2018.
Given the patient’s medical history in the above case study,
the proposed model would assign a readmission risk to the
patient. But at the same time, the model would provide
additional factors that could be true for the patient. For instance,
electrolyte imbalance would be a factor for readmission,
along with the other chronic diseases. The discharge staff
would include that in the notes to ensure that it is monitored
in rehabilitation phase and possibly save the patient’s life.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>9 Conclusions</title>
      <p>In the last decade, there have been numerous studies that
link factors pertaining to a patient’s hospital stay to the risk
of readmission. However, most studies have been on a
focused cohort, limited to one or few hospitals. We show here
that similar results can be achieved using claims data, which
has fewer elements but provides a large population coverage;
the entire state of New York for this study and the wide
population coverage of MIMIC-III. Even with the large volume
of data, the predictive algorithms are not accurate enough to
be used as decision making tools. However, model
interpretation can reveal insights which can inform the strategies for
reducing and/or eliminating readmissions.</p>
      <p>A patient’s disease history is typically expressed using
diagnosis codes, which can take as many as 18000 possible
values, with many more possibilities in the next generation
ICD-10 disease classification. With so many possible
features, ensuring model interpretability is a challenge.
However, using structured sparsity inducing models, such as the
one proposed here, one can ensure that the truly important
factors can be identified within the hierarchy.</p>
      <p>10 Acknowledgment
This material is based upon work partially supported by
a seed grant from the University at Buffalo Germination
Space Program. We thank Chiahui Chen, MS, RN, FNP-BC,
School of Nursing, UB for developing the case study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Amarasingham</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>B. J.</given-names>
          </string-name>
          ; Tabak,
          <string-name>
            <given-names>Y. P.</given-names>
            ;
            <surname>Drazner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            ;
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. A.</surname>
          </string-name>
          ; Zhang,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. G.</surname>
          </string-name>
          ; Swanson,
          <string-name>
            <given-names>T. S.</given-names>
            ; Ma, Y.; and
            <surname>Halm</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. A.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>An automated model to identify heart failure patients at risk for 30-day readmission or death using electronic medical record data</article-title>
          .
          <source>Medical Care</source>
          <volume>48</volume>
          (
          <issue>11</issue>
          ):
          <fpage>981</fpage>
          -
          <lpage>989</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>F. R.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Consistency of the group lasso and multiple kernel learning</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>9</volume>
          :
          <fpage>1179</fpage>
          -
          <lpage>1225</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Badawi</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Breslow</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Readmissions and death after icu discharge: development and validation of two predictive models</article-title>
          .
          <source>PloS one</source>
          <volume>7</volume>
          (
          <issue>11</issue>
          ):e48758;
          <fpage>e48758</fpage>
          -
          <lpage>e48758</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          2010.
          <article-title>Graph-structured multi-task regression and an efficient optimization method for general fused lasso</article-title>
          .
          <source>CoRR abs/1005</source>
          .3579.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Choudhry</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Erdmann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sikka</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Sutariya</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>A public-private partnership develops and externally validates a 30-day hospital readmission risk prediction model</article-title>
          .
          <source>Online J Public Health Inform</source>
          .
          <volume>5</volume>
          (
          <issue>2</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Conway</surname>
            ,
            <given-names>P. H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Berwick</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Improving the rules for hospital participation in medicare and medicaid</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>JAMA</source>
          <volume>306</volume>
          (
          <issue>20</issue>
          ):
          <fpage>2256</fpage>
          -
          <lpage>2257</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Cox</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          <year>1958</year>
          .
          <article-title>The regression analysis of binary sequences</article-title>
          .
          <source>Journal of the Royal Statistical Society</source>
          . Series B (
          <year>Methodological</year>
          )
          <fpage>215</fpage>
          -
          <lpage>242</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          2013.
          <article-title>Potentially avoidable 30-day hospital readmissions in medical patients: Derivation and validation of a prediction model</article-title>
          .
          <source>JAMA Internal Medicine</source>
          <volume>173</volume>
          (
          <issue>8</issue>
          ):
          <fpage>632</fpage>
          -
          <lpage>638</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Futoma</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Morris</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Lucas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>A comparison of models for predicting early hospital readmissions</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>56</volume>
          :
          <fpage>229</fpage>
          -
          <lpage>238</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mathews</surname>
            ,
            <given-names>S. C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kalloo</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Hutfless</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Hewner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Seo</surname>
            ,
            <given-names>J. Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Gothard</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          ; and Johnson,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Aligning</surname>
          </string-name>
          population
          <article-title>-based care management with chronic disease complexity</article-title>
          .
          <source>Nursing Outlook</source>
          <volume>62</volume>
          (
          <issue>4</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Jamei</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nisnevich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wetchler</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sudat</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and Liu,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Predicting all-cause risk of 30-day hospital readmission using artificial neural networks</article-title>
          .
          <source>PLoS ONE</source>
          <volume>12</volume>
          :
          <year>e0181173</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chin</surname>
          </string-name>
          , K.-S.; Qu, G.; and
          <string-name>
            <surname>Tsui</surname>
            ,
            <given-names>K. L.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>An integrated machine learning framework for hospital readmission prediction</article-title>
          .
          <source>Knowledge-Based Systems</source>
          <volume>146</volume>
          :
          <fpage>73</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>A. E. W.</given-names>
          </string-name>
          ; Pollard,
          <string-name>
            <given-names>T. J.</given-names>
            ;
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Lehman</surname>
          </string-name>
          , L.- w. H.; Feng,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Ghassemi</surname>
          </string-name>
          , M.; Moody, B.;
          <string-name>
            <surname>Szolovits</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Anthony Celi, L.; and Mark,
          <string-name>
            <surname>R. G.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Mimic-iii, a freely accessible critical care database</article-title>
          .
          <source>Scientific Data</source>
          <volume>3</volume>
          :160035 EP -.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Kansagara</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Englander,
          <string-name>
            <given-names>H.</given-names>
            ;
            <surname>Salanitro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Kagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Theobald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Freeman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Kripalani</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2011</year>
          .
          <article-title>Risk prediction models for hospital readmission: a systematic review</article-title>
          .
          <source>Journal of American Medical Association</source>
          <volume>306</volume>
          (
          <issue>15</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Langford</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and Zhang, T.
          <year>2008</year>
          .
          <article-title>Sparse online learning via truncated gradient</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>10</volume>
          :
          <fpage>777</fpage>
          -
          <lpage>801</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.-W.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Faghri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shaw</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ; and Campbell,
          <string-name>
            <surname>R. H.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Analysis and prediction of unplanned intensive care unit readmission using recurrent neural networks with long short-term memory</article-title>
          .
          <source>bioRxiv.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Ji,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2009</year>
          .
          <article-title>SLEP: Sparse Learning with Efficient Projections</article-title>
          . Arizona State University.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          2010.
          <article-title>Solving structured sparsity regularization with proximal methods</article-title>
          .
          <source>In Proceedings of the 2010 European Conference on Machine Learning and Knowledge Discovery in Databases: Part II</source>
          ,
          <string-name>
            <surname>ECML</surname>
            <given-names>PKDD</given-names>
          </string-name>
          '
          <volume>10</volume>
          ,
          <fpage>418</fpage>
          -
          <lpage>433</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          2007.
          <article-title>Promoting greater efficiency in medicare</article-title>
          .
          <source>MPA Committee Report to Congress.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Niu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Regression Models for Readmission Prediction Using Electronic Medical Records</article-title>
          .
          <source>Master's thesis</source>
          , Wayne State University, Detroit, Michigan.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Pfuntner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wier</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Steiner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Costs for hospital stays in the united states</article-title>
          ,
          <year>2011</year>
          .
          <article-title>Healthcare Cost and Utilization Project (HCUP) Statistical Briefs 168</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nadkarni</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Guttag</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Bottinger</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <article-title>Leveraging hierarchy in medical codes for predictive modeling</article-title>
          .
          <source>In Proceedings of the 5th ACM Conference on Bioinformatics</source>
          , Computational Biology, and Health Informatics, BCB '
          <volume>14</volume>
          ,
          <fpage>96</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Tibshirani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Saunders,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Rosset</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Zhu</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Sparsity and smoothness via the fused lasso</article-title>
          .
          <source>Journal of the Royal Statistical Society Series B 91-108.</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Tibshirani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Regression shrinkage and selection via the lasso</article-title>
          .
          <source>Journal of the Royal Statistical Society, Series B</source>
          <volume>58</volume>
          :
          <fpage>267</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          2018.
          <article-title>Readmission prediction via deep contextual embedding of clinical concepts</article-title>
          .
          <source>PloS one 13</source>
          <volume>(4)</volume>
          :
          <fpage>e0195024</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Esbroeck</surname>
            ,
            <given-names>A. v.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Farooq</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fung</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Anand</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Krishnapuram</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Predicting readmission risk with institution specific prediction models</article-title>
          .
          <source>In Proceedings of the 2013 IEEE International Conference on Healthcare Informatics</source>
          ,
          <fpage>415</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Model selection and estimation in regression with grouped variables</article-title>
          .
          <source>JOURNAL OF THE ROYAL STATISTICAL SOCIETY, SERIES B</source>
          <volume>68</volume>
          :
          <fpage>49</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Rocha, G.; and
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>The composite absolute penalties family for grouped and hierarchical variable selection</article-title>
          .
          <source>The Annals of Statistics</source>
          <volume>37</volume>
          (
          <year>6A</year>
          ):
          <fpage>3468</fpage>
          -
          <lpage>3497</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>