<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Robust Lock-Down Optimization for COVID-19 Policy Guidance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ankit Bhardwaj</string-name>
          <email>bhardwaj@wadhwaniai.org</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Han Ching Ou</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haipeng Chen</string-name>
          <email>hpchen@seas.harvard.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shahin Jabbari</string-name>
          <email>jabbari@seas.harvard.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milind Tambe</string-name>
          <email>tambe@seas.harvard.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rahul Panicker</string-name>
          <email>rahul@wadhwaniai.org</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alpan Raval</string-name>
          <email>alpan@wadhwaniai.org</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Equal Contribution</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Harvard University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>As the COVID-19 outbreak continues to pose a serious worldwide threat, numerous governments choose to establish lockdowns in order to reduce disease transmission. However, imposing the strictest possible lock-down at all times has dire economic consequences, especially in areas with widespread poverty. In fact, many countries and regions have started charting paths to ease lock-down measures. Thus, planning efficient ways to tighten and relax lock-downs is a crucial and urgent problem. We develop a reinforcement learning based approach that is (1) robust to a range of parameter settings, and (2) optimizes multiple objectives related to different aspects of public health and economy, such as hospital capacity and delay of the disease. The absence of a vaccine or a cure for COVID to date implies that the infected population cannot be reduced through pharmaceutical interventions. However, non-pharmaceutical interventions (lock-downs) can slow disease spread and keep it manageable. This work focuses on how to manage the disease spread without severe economic consequences.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>While governments are responding to the spread of
COVID19 by imposing lock-downs of varying intensity to reduce
human-human contact, the situation cannot be maintained
indefinitely. Each day of lock-down brings severe economic
loss affecting the livelihood of billions. Thus, it is imperative
to use the available resources of interventions – lock-downs,
test-kits, ventilators etc., in an efficient manner. This work
aims to find optimal lock-down policies based on
epidemiological models and reinforcement learning.</p>
      <p>
        Reinforcement learning has shown promising results on
sequential decision making tasks like Go
        <xref ref-type="bibr" rid="ref28">(Silver et al. 2016)</xref>
        and autonomous driving
        <xref ref-type="bibr" rid="ref20">(Pan et al. 2017)</xref>
        . On these tasks,
training from real-world data directly is too expensive due
to costly data collection process. Learning the agent model
from simulations is thus necessary. However, simulations
don’t reflect the real world exactly due to uncertainties
transferred while fitting the simulation model
        <xref ref-type="bibr" rid="ref9">(Christiano et al.
2016)</xref>
        . It can be dangerous if the policy learned is unaware
of such uncertainties, which is especially true for our task.
      </p>
      <p>In addition to insufficient data and uncertainty, for our
problem, it is hard to specify a single objective that one
wants to achieve. It is likely that many objectives need to be
met for the task to be considered successful. For example,
we may want to delay the peak of infections while making
sure that our hospitals are not overburdened or our economy
is not affected too severely. The decision maker in this case
is looking at the problem from several perspectives leading
to many possible objectives that the model will be evaluated
on. Thus, in this work, we incorporate multi-objective
functions.</p>
      <p>The main contributions of this work can be summarized
as follows:
• We formulate the problem of lock-down implementation
as a Markov Decision Process (MDP). To solve this MDP,
we propose a Reinforcement Learning (RL) approach that
optimizes the trade-off between health objectives and
economic cost.
• We tackle the uncertainty in environment parameters that
might arise from the noise in the data and the process of
estimation by considering different robust approaches.
• We analyse different robust approaches including
uniform sampling and adversarial sampling during the
training phase. We find that there is a trade-off relation in the
average-case and worst-case between RL agents with
different degrees of risk aversion.
• We design different health objectives that might be of
interest to decision-makers and measure our performance
along these different objectives simultaneously.</p>
      <p>With this work, we aim to address the challenging task of
planning temporal resource allocation for lock-downs. The
models that we use for modelling the spread of COVID are
the SEIR class of epidemiological models.</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Work</title>
      <p>
        Since as early as the 17th century, when Bernoulli
proposed the first mathematical epidemic model for
smallpox
        <xref ref-type="bibr" rid="ref18 ref5">(Bernoulli and Blower 2004)</xref>
        , there have been
numerous efforts in the modeling and control of epidemics. One
important class of these models are called compartmental
models. These models, as their name suggests, divide the
population into different health states (compartments) and
model transitions of populations between these health states.
The underlying assumption is that these compartments have
homogeneously mixed populations.
Susceptible-ExposedInfected-Recovered (SEIR) family of models are
compartmental models with dynamics described by ordinary
differential equations. Recently, there have been advances in
fitting SEIR models with machine learning techniques
        <xref ref-type="bibr" rid="ref3">(Bannur
et al. 2021)</xref>
        . In this work, we use a
Susceptible-ExposedInfected-Recovered-Deceased (SEIRD) model (detailed
description in subsection Epidemic Model) to model the
COVID-19 data. However, the technique we propose is
applicable to any of the SEIR family models.
      </p>
      <p>
        Apart from epidemic modeling, the problem of
optimizing cure and control for preventing the spread of disease is
also of interest. However, most works in the computer
science literature usually assume an idealistic model, such as
every contact being known, no uncertainty in the disease
parameters or a strong cure/isolation that guarantees the
recovery of the individuals
        <xref ref-type="bibr" rid="ref1 ref12 ref19 ref29 ref33 ref33 ref34 ref34 ref37 ref37">(Ball, Knock, and O’Neill 2015;
Sun and Hsieh 2010; Wang 2005; Zhang and Prakash 2015;
Ganesh, Massoulie´, and Towsley 2005)</xref>
        . None of these are
true for most real-world diseases, such as the newly arisen
COVID-19 pandemic which has no cure as of the writing of
this paper. Even under most settings being ideal, a small
uncertainty could have serious implications on outcomes if not
handled properly. For example, the impact of curing
uncertainty under perfect observation is analyzed in Hoffman and
Caramanis
        <xref ref-type="bibr" rid="ref14">(Hoffmann and Caramanis 2018)</xref>
        by providing
non-constructive, algorithm-independent bounds. We aim to
address the challenging setting in which there are
uncertainties in most of the parameters in the model.
      </p>
      <p>
        Robust control is a branch of control theory that has a long
history. In particular, robustness toward parameter
uncertainty results in a performance drop from the model toward
its real-world application
        <xref ref-type="bibr" rid="ref15">(Mannor et al. 2004)</xref>
        . Numerous
works
        <xref ref-type="bibr" rid="ref18 ref19 ref33 ref35 ref5">(Nilim and El Ghaoui 2005, 2004; White III and
Eldeib 1994)</xref>
        have tried to tackle such uncertainty under the
robust MDP framework with different assumptions. In recent
years, Reinforcement learning has demonstrated promising
results on a variety of MDP problems
        <xref ref-type="bibr" rid="ref20 ref28">(Silver et al. 2016;
Pan et al. 2017)</xref>
        . For applications with a high safety
requirement, it is natural to combine robustness into reinforcement
learning
        <xref ref-type="bibr" rid="ref17 ref6">(Mihatsch and Neuneier 2002; Carpin, Chow, and
Pavone 2016; Chow et al. 2017)</xref>
        . Among these works, using
an adversarial agent to adjust the environment and discover
potential risk systematically has shown promising results in
many real-world tasks (Pinto, Davidson, and Gupta
        <xref ref-type="bibr" rid="ref8">2017;
Pattanaik et al. 2017</xref>
        ). A recent algorithm using an
adversarial framework is robust adversarial reinforcement learning
(RARL)
        <xref ref-type="bibr" rid="ref22 ref23">(Pinto et al. 2017)</xref>
        in which, two agents are trained,
one protagonist and other an adversary providing attacks on
input states and dynamics. In our work, similar to RARL,
we use an adversarial agent to systematically search risky
environmental parameters for the policy.
      </p>
      <p>
        Another important consideration for lock-down policy
makers might be to include different desirable objectives
in their decision-making. Multi-objective optimization has
tremendous practical importance in many real-life
applications
        <xref ref-type="bibr" rid="ref10">(Deb 2014)</xref>
        . Linear combinations of Pareto
optimali
      </p>
      <sec id="sec-2-1">
        <title>Notations</title>
        <p>S
E
I
R
D
t
Ro</p>
        <sec id="sec-2-1-1">
          <title>Tinc</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>Tinf</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>Trecover Tfatal</title>
          <p>
            ties are often considered and solved when the system is easy
to describe
            <xref ref-type="bibr" rid="ref7">(Censor 1977)</xref>
            . Multi-Objective Reinforcement
Learning however
            <xref ref-type="bibr" rid="ref25 ref31">(Roijers et al. 2013; Van Moffaert and
Nowe´ 2014)</xref>
            , is a relatively new research area that has been
actively studied only in recent years. In this work, we
designed a Quality Adjusted Life Year (QALY) value function
to calculate the suitable reward signal for any point of any
two objectives of the QALY variant. Such a function can
also be used to determine the difficulties of optimizing the
two objectives simultaneously.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Modelling</title>
      <sec id="sec-3-1">
        <title>Epidemic Model</title>
        <p>
          For modelling COVID-19, we adopt a discrete time SEIRD
model
          <xref ref-type="bibr" rid="ref34 ref37">(Weitz and Dushoff 2015)</xref>
          . The SEIRD class of
models is a part of compartmental models, as mentioned above.
An individual can be in one of the following health states:
S (a healthy individual susceptible to disease), E (the
individual has been exposed and has latent disease), or I (the
individual is infected), R (the individual is recovering and is
no longer infectious to others) and D (the individual is
deceased). Table 1 summarizes the symbols we use throughout
this paper.
        </p>
        <p>The discrete-time dynamics equations for our epidemic
model are:</p>
        <p>St+1 − St =
Et+1 − Et = (</p>
        <p>It+1 − It = (
Rt+1 − Rt =
Dt+1 − Dt =</p>
        <p>−StIt</p>
        <sec id="sec-3-1-1">
          <title>Ttrans(a, e)</title>
          <p>,</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Ttrans(a, e)</title>
          <p>St
−
Et</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Tinc</title>
          <p>It</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Trecover</title>
          <p>Rt
,
,
in which Ti1nf = Trec1over + Tfa1tal and the basic
reproductive number can be obtained from R0 = Tinf /Ttrans.
A typical SEIRD model described above starts from a
population being mostly susceptible and a small fraction of
infectious people. When R0 &gt; 1, each infected individual
will infect more than one susceptible individual in its
lifetime on average. Each susceptible individual will eventually
go through the exposed, infectious to finally recovered or
deceased states. A schematic diagram of the SEIRD model
is shown in Figure 1. Generally in compartmental models,
Ttrans is a constant. However, in a real-world setting, the
transmission time can be reduced through the deployment
of non-pharmaceutical lock-down interventions. Note that
there is no direct reduction in infected population as there is
no cure or vaccination available. We only consider the
lockdown interventions that increase the transmission time of the
virus based on their strength.</p>
          <p>
            Compartmental models and their variants are commonly
used in disease state forecasting and prediction. For
concreteness, in this work, state populations and numerical
values of the transmission parameters are based on the available
data for the city of Mumbai
            <xref ref-type="bibr" rid="ref13">(Group 2020)</xref>
            . However, it is
worth noting that both the intervention model and planning
algorithm we propose apply for most, if not all, of the SEIR
model variants.
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Intervention Modelling</title>
        <p>
          As there is no cure/vaccine available for COVID-19 till date,
models of pharmaceutical interventions are not applicable.
To manage the rapid disease spread, different lock-down
policies can be considered to limit individual contacts. For
example,
          <xref ref-type="bibr" rid="ref11">(Ferguson et al. 2020)</xref>
          considered different
levels of lock-down such as Case isolated at home, Voluntary
home quarantine, Social distancing of those over 70 years
of age, Social distancing of entire population and Closure
of schools and universities for non-pharmaceutical
interventions in British population, which all have different cost and
effectiveness.
        </p>
        <p>These lock-down policies should enjoy several
desiderata for real-world deployment. First, each type of
intervention needs to last for a minimum duration d. Second, since
the lock-down has economic cost, government bodies would
expect a trade-off between the total budget B spent for
planning and policy deployment and public health related gains.
To model such interventions, we considered lock-downs as a
series of action choices. The decision maker can plan
different policies in different time periods based on the limitations
mentioned above.</p>
        <p>
          During the lock-down period in India, the change in
estimated transmission time as the effect of interventions has
been observed based on
          <xref ref-type="bibr" rid="ref13">(Group 2020)</xref>
          . This corresponds to
Ttrans in the SEIRD model we proposed and the
effectiveness e varies in different regions. Thus we modeled the
action as extending the transmission time to different degrees
and different costs per day. Such a sequence of actions forms
an intervention vector a of length T as a planning schedule
with a total cost sum.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Multi-Objective Functions</title>
        <p>In the public health domain, governments and decision
makers may want to achieve different objectives when deploying
a policy. One direct objective could be eliminating the
disease which can be achieved by suppressing contacts so that
patients recover at a rate greater than the spread of infection.
This is equivalent to minimizing the area under the infection
curve. However, this is not achievable in many regions,
including cities in many of the developing countries due to the
huge economic cost of such strict lock-downs. Thus, we
focus on economically sustainable interventions that do not
reduce R0 below 1. In epidemic theory, this means the disease
cannot be eliminated within reasonable time no matter how
the government plans the lock-down in these regions. Every
susceptible individual will eventually go through the
recovered or deceased state. In other words, although the infection
curve will change, the area under the curve will remain the
same.</p>
        <p>To evaluate the effectiveness of lock-down policies
under these circumstances, we use indirect objectives that are
vital and achievable for sustainable interventions. For
example, as there is limited hospital capacity, a patient’s
quality of life will likely be better when the infected population
does not exceed that capacity and they can receive proper
treatment. Alternatively, we may want to delay the
infection to the point when we have better system preparedness,
medicines, resources etc. for handling the disease. These
different desired objectives can be described as a family of
objective functions that are variants of the Quality Adjusted
Life Year (QALY) score in our model, which is elaborated
below.</p>
        <p>
          QALY is a popular established metric to quantify the
effectiveness of health interventions. It is often used in the
public health literature
          <xref ref-type="bibr" rid="ref26">(Salomon et al. 2012)</xref>
          . It measures
the effectiveness of a certain intervention by combining
quantity and quality of health improvement. Specifically, a
person’s life quality at any given time is mapped to [0, 1],
with quality 1 corresponding to perfect health while 0
corresponding to death. QALY accumulates such measurements
over time as its final score. Naturally, different disease
conditions lie in the range [0, 1] depending on severity.
        </p>
        <p>In this work, we change the time scale from years to days
to adapt to the dynamics of the disease we are facing. We
mainly focus mainly on two objectives, burden and delay.
We define these two objective functions as:</p>
        <p>OBurden = X((I(t) − λ)1I(t)&gt;λ − δBurdenl(at)) (6)
t
ODelay = X(tI(t) − δDelayl(at)).</p>
        <p>t
(7)
where t refers to timestep. at, I(t) and l(at) refer to
action, infected population fraction and cost of action at time
t respectively. Also δ and λ refer to economy-health weight
and hospital capacity and 1 is the indicator function. Here,
we focus on optimizing a linear combination of these two
objective functions, written as:</p>
        <p>Omix(w) = wO¯Burden + (1 − w)O¯Delay,
(8)
where the weight w ranges from 0 to 1 and O¯ is O
normalized by the absolute value of no intervention, i.e., we divide
O by its absolute value in the absence of interventions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Formulation Using MDP</title>
      <p>
        Our lock-down control problem can be modeled as a Markov
Decision Process (MDP)
        <xref ref-type="bibr" rid="ref36">(Yang, Sun, and Narasimhan
2019)</xref>
        . Over the last two decades, reinforcement
learning
        <xref ref-type="bibr" rid="ref30">(Sutton et al. 1998)</xref>
        has provided an effective framework
for solving an MDP in both theory and application. This is
especially true when the system dynamics is either
complicated, unknown, or the state dimensionality is too high for
classical optimal control methods. In addition, the
environment parameters we estimate from real-world hospital data
involve uncertainty that cannot be ignored. Thus the output
policy needs to be robust to such uncertainty.
      </p>
      <p>We thus consider a parameter-wise robust reinforcement
learning model to solve the MDP framework. The MDP
framework can be written as:</p>
      <p>hS, A, P, Ri
with state space S, action space A, and transition
distribution and vector reward</p>
      <p>P(s′|s, a) for s, s′ ∈ S and a ∈ A</p>
      <p>r(s) ∈ R
and the preference weight w ∈ Rn.</p>
      <p>The states we consider are the fractions of population
present in S,E,I,R and D compartments at the given time.
Furthermore, we consider several discrete actions at each
time step corresponding to different strengths of lock-down
with different costs. It is natural to assume the strength to be
monotonically increasing with the cost as otherwise the
action choice will be dominated by actions with less cost but
more effectiveness. For simplicity, we adopt a linear
mapping for both cost and effectiveness, as:</p>
      <p>Ttrans(a, e) = (1 + ec(a)) Tinf</p>
      <p>R0
(9)
and c : A → [0, 1], in which e is the lock-down
effectiveness coefficient and R0 the basic reproduction number
when there are no lock-down interventions. Both of these
are estimated with data from the city of Mumbai, India.
Ttrans and Tinf are the transmission and infection time
periods.</p>
      <p>For the remaining tuple, the transition distribution
P(s′|s, a) is described as the disease transmission
equation 1 to 5. The total accumulated reward is exactly the
objective function O in equations 6, 7. The next section
describes the distribution of individual reward signals across
states.</p>
    </sec>
    <sec id="sec-5">
      <title>Reinforcement Learning Approach</title>
      <p>Multiple Objectives: We have defined the state, action,
transition probabilities and total reward in the MDP
section. The only missing piece for a complete reinforcement
learning framework is to design the reward signal at every
timestep. We have designed a framework that works not only
for the two example objectives we focus in this work, but on
most variants of QALY.</p>
      <p>Most variants of QALY, including our examples, are
related to time and population of certain health states. We
propose a function we call the QALY value V (x, t) which is a
function of the population x in a certain health state and time
t. We focus only on I or the Infectious state in these
experiments. However, the QALY value function can be
generalized to a vector form to include multiple states. For
controlling the hospital capacity, the function can be formulated
either as a constant penalty for x exceeding the capacity or
simply as a reward for x below the threshold, since the area
under the infection curve is a constant, as we elaborate in
section Multi-Objective Functions. We formalize this as:</p>
      <p>VBurden(x, t) = 1 for x &lt; λ
As for delay, we formalize this function as
(10)
(11)
VDelay(x, t) =
t
T</p>
      <p>Given that the QALY value V (x, t) of the objective
function is defined, the reward signal at any given time t can be
calculated by r(t) = R I(t) V (x, t)dt. We can thus apply the
0
reinforcement learning approach.</p>
      <p>One benefit of such a proposed approach is that the QALY
value function of the mixed objective can be easily
calculated as:</p>
      <p>Vmix(w) =</p>
      <p>wVBurden
OBurden(no action)
+</p>
      <p>
        (1 − w)VDelay
ODelay(no action)
(12)
Uncertainty: Another important aspect other than having a
multi-objective function in the lock-down application is the
uncertainty of the parameters (e, Tinf , Tinc), which are
related to the infection curve directly or indirectly. We
experiment with three approaches to analyze the effect of
uncertainty in a reinforcement learning setup:
(1)Fixed RL(FRL): Train the RL agent using only the mean
of the uncertain parameters.
(2)Distributed RL(DRL): Train the RL agent using
samples of uncertain parameters from the estimated range.
(3)Adversarial RL(ARL): Inspired by
        <xref ref-type="bibr" rid="ref22 ref23">(Pinto et al. 2017)</xref>
        ,
train the RL agent with another adversarial RL agent that
will maliciously pick the worst possible parameter set for the
RL agent during training. Note that the worst case
parameter is not trivial to find as the policy changes. The action of
the adversarial RL agent is set to be the discrete uncertain
parameters in the disease model.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>In this section, we describe the application of our method
to a specific location – the city of Mumbai, India. In
subsequent subsections, we describe how we estimate the model
parameters as well as the uncertainty in these parameters.
We also report the results of our method when used on the
estimated parameter ranges.</p>
      <sec id="sec-6-1">
        <title>Parameter and Uncertainty Estimation</title>
        <p>
          We fit our SEIRD model to the time-series data from the
COVID19-India API
          <xref ref-type="bibr" rid="ref13">(Group 2020)</xref>
          for the city of Mumbai.
The data is aggregated in fields called Recovered, Deceased,
Hospitalized and Total Infected, where Total Infected =
Recovered + Hospitalized + Deceased. In our SEIRD model,
we fit the I compartment to Total Infected, D compartment
to Deceased and R compartment to Hospitalized +
Recovered. In this sense, the R compartment in our model
estimates people who are either under recovery or have
recovered, and thus are no longer infectious.
        </p>
        <p>
          We decided the initial search space for the model
parameters based on the estimates given by public health
experts and those cited in literature. We process the data with
smoothing techniques to reduce the effect of bulk data
entry. Then, we search over the parameter space for parameter
sets that have a small aggregated RMSE loss between
predicted numbers and actual numbers using the Hyperopt
library
          <xref ref-type="bibr" rid="ref4">(Bergstra, Yamins, and Cox 2013)</xref>
          . The parameter set
giving the least loss value is taken to be the best-fit
parameter set for the purposes of this experiment.
        </p>
        <p>We found that there are diverse parameter sets that have
loss close to the best-fit parameter set. Thus, we picked all
parameter sets that have a loss within a certain range of the
best loss (within 10%). Among all picked parameter sets,
we find the range of values taken by individual parameters.
These ranges for individual parameters give us a measure
of uncertainty for these parameters. We assume a uniform
distribution over these ranges as our parameter distribution.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Analysis and Results</title>
        <p>Robustness: Robustness of policy to uncertainty in
parameters is an important aspect. Over the estimated uniform
distribution range, we find the worst-case parameters for
different methods using a fine grid-search. Then, we measure</p>
      </sec>
      <sec id="sec-6-3">
        <title>Model</title>
        <p>Random
FRL
DRL
ARL
the performance of different methods on their corresponding
worst-case parameters. We also find the corresponding
average performance over the parameter distribution. The results
are tabulated in Tables 2 and 3. As shown in these tables, the
ARL helps the reinforcement learning discover risky
parameters and thus performs best in its worst case scenario. For
average case, however, ARL performs worse than the best
method (FRL and DRL respectively). This has shown the
trade-off between performance and robustness in our
lockdown problem - at the cost of average performance, we can
obtain better worst-case performance.</p>
        <p>Different objectives: We use different weights between
Burden and Delay objectives and compare the results to the
case when we individually focus on Delay and Burden in
Table 4. The objective function we use is (1 − w) ∗ Delay +
w ∗ Burden for different values of w. The aim is to maximize
normalized objective for both Burden and Delay.</p>
        <p>From Table 4, we observe that, as expected, as the weight
on Burden increases, the Burden objective becomes larger
for all methods in general. Similar behaviour is observed for
Delay as well. When applying this method for policy
guidance, we can tune w to achieve the required objectives for
both Burden and Delay.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and Future Work</title>
      <p>We implemented reinforcement learning on the lock-down
policy optimization problem for COVID-19 while
considering important real-world aspects like robustness and
multiobjective optimization. Robustness can be achieved by
introducing an adversarial agent for parameter discovery, but at
the cost of sacrificing some performance on average. For the
multi-objective mixture, we study the trade-off between
controlling hospital capacity and delaying the infection spread.
We proposed a reward distribution framework for the
reinforcement learning agent to shift from one objective to
another in the lock-down problem. One point to note is that our
epidemiological model (SEIRD) is a homogeneous model
and is being used to optimize the policy keeping the trade-off
between economy and health for the community as a whole.
The model does not discriminate between two infected
individuals based on their economic contribution and neither
is it capable for the same. This makes sure that we generate
lockdown policy as fairly as possible.</p>
      <p>The future direction of this work is to gather more data
on both cost and effectiveness of the real-world lock-down
policies on community scale so that a more complex model
can be used to better estimate the real-world scenarios. For
example, transmission times are known to not be
homogeneous and several super-spreader events have been identified
in many different spreading routes. Collecting data on such
cases and modifying the model to have different
transmission times for different cases of spread would give us a more
holistic view of the entire scenario. Another important
direction of extension would be estimating the reporting rate from
other sources of data and normalizing the reported numbers
to estimate parameters that are closer to the real-world.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This study is made possible by the generous support of
the American People through the United States Agency
for International Development (USAID) and Army
Research Office (ARO). The work described in this article
was implemented under the TRACETB Project, managed
by WIAI under the terms of Cooperative Agreement
Number 72038620CA00006 and by Teamcore, CRCS, Harvard
University under Multidisciplinary University Research
Initiative grant number W911NF1810208. The contents of this
manuscript are the sole responsibility of the authors and do
not necessarily reflect the views of USAID, ARO or the
United States Government.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ball</surname>
            ,
            <given-names>F. G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Knock</surname>
            ,
            <given-names>E. S.</given-names>
          </string-name>
          ; and
          <string-name>
            <given-names>O</given-names>
            <surname>'Neill</surname>
          </string-name>
          ,
          <string-name>
            <surname>P. D.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Stochastic epidemic models featuring contact tracing with delays</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>Mathematical biosciences</source>
          <volume>266</volume>
          :
          <fpage>23</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bannur</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Maheshwari</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ; Jain,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Shetty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Merugu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Raval</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2021</year>
          .
          <article-title>Adaptive COVID-19 Forecasting via Bayesian Optimization</article-title>
          .
          <source>In Proceedings of the ACM India Joint International Conference on Data Science and Management of Data</source>
          , CoDS-COMAD '
          <fpage>21</fpage>
          . New York, NY, USA:
          <article-title>Association for Computing Machinery</article-title>
          .
          <source>doi:10.1145/ 3430984</source>
          .3431047. URL https://doi.org/10.1145/3430984.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bergstra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yamins</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Cox</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures</article-title>
          .
          <source>In International conference on machine learning</source>
          ,
          <fpage>115</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Bernoulli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Blower</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>An attempt at a new analysis of the mortality caused by smallpox and of the advantages of inoculation to prevent it</article-title>
          .
          <source>Reviews in medical virology 14</source>
          <volume>(5)</volume>
          :
          <fpage>275</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Carpin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Chow, Y.-L.; and Pavone,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Risk aversion in finite Markov Decision Processes using total cost criteria and average value at risk</article-title>
          .
          <source>In 2016 ieee international conference on robotics and automation (icra)</source>
          ,
          <fpage>335</fpage>
          -
          <lpage>342</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Censor</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>1977</year>
          .
          <article-title>Pareto optimality in multiobjective problems</article-title>
          .
          <source>Applied Mathematics and Optimization</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <fpage>41</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          2017.
          <article-title>Risk-constrained reinforcement learning with percentile risk criteria</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>18</volume>
          (
          <issue>1</issue>
          ):
          <fpage>6070</fpage>
          -
          <lpage>6120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Christiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mordatch</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Blackwell,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Tobin</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ; Abbeel,
          <string-name>
            <given-names>P.</given-names>
            ; and
            <surname>Zaremba</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Transfer from simulation to real world through learning deep inverse dynamics model</article-title>
          .
          <source>arXiv preprint arXiv:1610</source>
          .
          <fpage>03518</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Deb</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Multi-objective optimization</article-title>
          .
          <source>In Search methodologies</source>
          ,
          <fpage>403</fpage>
          -
          <lpage>449</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ferguson</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Laydon</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nedjati-Gilani</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ; Imai,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Ainslie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Baguelin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Bhatia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Boonyasiri</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          ; Cucunuba´,
          <string-name>
            <given-names>Z.</given-names>
            ;
            <surname>Cuomo-Dannenburg</surname>
          </string-name>
          , G.; et al.
          <year>2020</year>
          .
          <article-title>Report 9: Impact of non-pharmaceutical interventions (NPIs) to reduce COVID19 mortality and healthcare demand</article-title>
          .
          <source>Imperial College London</source>
          <volume>10</volume>
          :
          <fpage>77482</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Ganesh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; Massoulie´, L.; and
          <string-name>
            <surname>Towsley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>The effect of network topology on the spread of epidemics</article-title>
          .
          <source>In Proceedings IEEE 24th Annual Joint Conference of the IEEE Computer and Communications Societies</source>
          ., volume
          <volume>2</volume>
          ,
          <fpage>1455</fpage>
          -
          <lpage>1466</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Group</surname>
            ,
            <given-names>C.-. I. O. D. O.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Accessed on yyyy-mm-dd from https://api</article-title>
          .covid19india.org/.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hoffmann</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Caramanis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>The Cost of Uncertainty in Curing Epidemics</article-title>
          .
          <source>Proceedings of the ACM on Measurement and Analysis of Computing Systems</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <fpage>31</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Mannor</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Simester</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Sun,
          <string-name>
            <given-names>P.</given-names>
            ; and
            <surname>Tsitsiklis</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. N.</surname>
          </string-name>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>Bias and variance in value function estimation</article-title>
          .
          <source>In Proceedings of the twenty-first international conference on Machine learning</source>
          ,
          <volume>72</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Mihatsch</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ; and Neuneier,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2002</year>
          .
          <article-title>Risk-sensitive reinforcement learning</article-title>
          .
          <source>Machine learning 49(2-3)</source>
          :
          <fpage>267</fpage>
          -
          <lpage>290</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Nilim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>El Ghaoui</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Robustness in Markov decision problems with uncertain transition matrices</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>839</volume>
          -
          <fpage>846</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Nilim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>El Ghaoui</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Robust control of Markov decision processes with uncertain transition matrices</article-title>
          .
          <source>Operations Research</source>
          <volume>53</volume>
          (
          <issue>5</issue>
          ):
          <fpage>780</fpage>
          -
          <lpage>798</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>You</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Virtual to real reinforcement learning for autonomous driving</article-title>
          .
          <source>arXiv preprint arXiv:1704</source>
          .
          <fpage>03952</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Pattanaik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Bommannan, G.; and Chowdhary,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Robust deep reinforcement learning with adversarial attacks</article-title>
          .
          <source>arXiv preprint arXiv:1712</source>
          .
          <fpage>03632</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Davidson</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Supervision via competition: Robot adversaries for learning tasks</article-title>
          .
          <source>In 2017 IEEE International Conference on Robotics and Automation (ICRA)</source>
          ,
          <fpage>1601</fpage>
          -
          <lpage>1608</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Sukthankar, R.; and
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <article-title>Robust adversarial reinforcement learning</article-title>
          .
          <source>arXiv preprint arXiv:1703</source>
          .
          <fpage>02702</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Roijers</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vamplew</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Whiteson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and Dazeley,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>A survey of multi-objective sequential decisionmaking</article-title>
          .
          <source>Journal of Artificial Intelligence Research</source>
          <volume>48</volume>
          :
          <fpage>67</fpage>
          -
          <lpage>113</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Salomon</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vos</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          ; Gagnon,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Naghavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Mokdad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Begum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Karyana,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Kosen</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; et al.
          <year>2012</year>
          .
          <article-title>Common values in assessing health outcomes from disease and injury: disability weights measurement study for the Global</article-title>
          <source>Burden of Disease Study</source>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <source>The Lancet</source>
          <volume>380</volume>
          (
          <issue>9859</issue>
          ):
          <fpage>2129</fpage>
          -
          <lpage>2143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Silver</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Maddison</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Guez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sifre</surname>
          </string-name>
          , L.; Van Den Driessche, G.;
          <string-name>
            <surname>Schrittwieser</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Antonoglou</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Panneershelvam</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lanctot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; et al.
          <year>2016</year>
          .
          <article-title>Mastering the game of Go with deep neural networks and tree search</article-title>
          .
          <source>nature</source>
          <volume>529</volume>
          (
          <issue>7587</issue>
          ):
          <fpage>484</fpage>
          -
          <lpage>489</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and Hsieh,
          <string-name>
            <surname>Y.-H.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Global analysis of an SEIR model with varying population size and vaccination</article-title>
          .
          <source>Applied Mathematical Modelling</source>
          <volume>34</volume>
          (
          <issue>10</issue>
          ):
          <fpage>2685</fpage>
          -
          <lpage>2697</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Sutton</surname>
            ,
            <given-names>R. S.</given-names>
          </string-name>
          ; et al.
          <year>1998</year>
          .
          <article-title>Introduction to reinforcement learning</article-title>
          , volume
          <volume>135</volume>
          . MIT press Cambridge.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Van Moffaert</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and Nowe´,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Multi-objective reinforcement learning using sets of pareto dominating policies</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <source>The Journal of Machine Learning Research</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3483</fpage>
          -
          <lpage>3512</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Modeling and analysis of massive social networks</article-title>
          .
          <source>Ph.D. thesis</source>
          , UMD.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Weitz</surname>
            ,
            <given-names>J. S.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Dushoff</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Modeling post-death transmission of Ebola: challenges for inference and opportunities for control</article-title>
          .
          <source>Scientific reports</source>
          <volume>5</volume>
          :
          <fpage>8751</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>White</surname>
            <given-names>III</given-names>
          </string-name>
          ,
          <string-name>
            <surname>C. C.</surname>
          </string-name>
          <article-title>;</article-title>
          and Eldeib,
          <string-name>
            <surname>H. K.</surname>
          </string-name>
          <year>1994</year>
          .
          <article-title>Markov decision processes with imprecise transition probabilities</article-title>
          .
          <source>Operations Research</source>
          <volume>42</volume>
          (
          <issue>4</issue>
          ):
          <fpage>739</fpage>
          -
          <lpage>749</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Narasimhan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>A generalized algorithm for multi-objective reinforcement learning and policy adaptation</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          ,
          <volume>14636</volume>
          -
          <fpage>14647</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.; and
          <string-name>
            <surname>Prakash</surname>
            ,
            <given-names>B. A.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Data-aware vaccine allocation over large networks</article-title>
          .
          <source>ACM Transactions on Knowledge Discovery from Data (TKDD) 10</source>
          (
          <issue>2</issue>
          ):
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>