<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Do Androids Dream of Electric Fences? Safety-Aware Reinforcement Learning with Latent Shielding</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Peter He</string-name>
          <email>peter.he.21@ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Borja G. Leo´ n</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Belardinelli</string-name>
          <email>francesco.belardinelli@ic.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Apricity</institution>
          ,
          <addr-line>London</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computing, Imperial College London</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Wellcome / EPSRC Centre for Interventional and Surgical Sciences, University College London</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The growing trend of fledgling reinforcement learning systems making their way into real-world applications has been accompanied by growing concerns for their safety and robustness. In recent years, a variety of approaches have been put forward to address the challenges of safety-aware reinforcement learning; however, these methods often either require a handcrafted model of the environment to be provided beforehand, or that the environment is relatively simple and low-dimensional. We present a novel approach to safetyaware deep reinforcement learning in high-dimensional environments called latent shielding. Latent shielding leverages internal representations of the environment learnt by modelbased agents to “imagine” future trajectories and avoid those deemed unsafe. We experimentally demonstrate that this approach leads to improved adherence to formally-defined safety specifications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The steady trickle of reinforcement learning (RL) systems
making their way out of the lab and into the real world
has cast a spotlight on the safety and robustness of RL
agents. The motivation behind this should be relatively easy
to grasp: when training an agent in real-world settings, it is
desirable that some states are never reached as they could,
for instance, cause permanent damage to the hardware the
agent is controlling. We can thus informally define the
notion of safety-aware RL in terms of the classical RL setup
with the added requirement that the number of unsafe states
visited be minimised. Under this definition, however, it has
been found that many state-of-the-art RL algorithms
unnecessarily enter unsafe states despite safe alternatives
being available and there being a positive correlation between
avoiding such states and reward (Giacobbe et al. 2021).</p>
      <p>
        The field of safety-aware RL encompasses a multitude
of approaches ranging from constrained policy optimisation
        <xref ref-type="bibr" rid="ref1 ref7">(Chow et al. 2017; Achiam et al. 2017; Yang et al. 2020)</xref>
        to safety critics
        <xref ref-type="bibr" rid="ref26 ref6">(Srinivasan et al. 2020; Bharadhwaj et al.
2021; Thananjeyan et al. 2021)</xref>
        to meta-learning
        <xref ref-type="bibr" rid="ref27">(Turchetta
et al. 2020)</xref>
        . In this work, we focus on a particular family of
approaches known as shielding
        <xref ref-type="bibr" rid="ref2 ref21 ref3 ref8">(Alshiekh et al. 2018;
Anderson et al. 2020; Giacobbe et al. 2021; ElSayed-Aly et al.
2021; Pranger et al. 2021)</xref>
        . Central to shielding is the
notion of a shield, a filter that checks actions proposed by the
agent’s existing policy with reference to a model of the
environment’s dynamics and some formal safety specification.
The shield overrides actions that may lead to an unsafe state
using some other safe (but by no means optimal) policy. A
key advantage of many shielding approaches is that the
resulting shielded policies are formally verifiable; however,
a shortcoming is that they require a model of
environmental dynamics - typically handcrafted - to be provided in
advance. Providing such a model may prove difficult for
complex real-world environments, with inaccuracies and human
biases creeping into handcrafted models.
      </p>
      <p>In this work, we propose a safe RL agent that makes uses
of latent shielding, an approach to shielding in environments
where a formally-specified dynamics model is not available
in advance. At an intuitive level, the agent uses a data-driven
approach to learn its own latent world model (a component
of which is a dynamics model) which is then leveraged by a
shield. The shield then uses the agent’s model to “imagine”
trajectories arising from different actions, forcing the agent
to avoid those it foresees leading to unsafe states. In
addition, the agent can be trained within its own latent world
model thus reducing the number of safety violations seen
during training.</p>
      <p>Contributions The main contribution of this work is a
framework for shielding agents in complex, stochastic and
high-dimensional environments without knowledge of
environmental dynamics a priori. We further introduce a new
method to aid exploration when training shielded agents.
Though our framework loses the formal safety guarantees
associated with traditional symbolic shielding approaches,
our experiments illustrate that latent shielding reduces
unsafe behaviour during training and achieves testing
performance comparable to previous symbolic approaches.</p>
    </sec>
    <sec id="sec-2">
      <title>Preliminaries</title>
      <p>In this section, we cover some relevant background topics.
We begin by introducing our problem setup for safety-aware
RL and give an overview of the specification language used
in this work. This is followed by an outline of the latent
world model we make use of in this work as well as a
discussion on shielding.</p>
      <sec id="sec-2-1">
        <title>Problem Setup</title>
        <p>We consider an agent interacting with an environment E
modelled as a partially observable Markov decision
process (POMDP) with states s ∈ SE , observations ot ∈ OE ,
agent-generated actions at ∈ AE and scalar rewards rt ∈ R
over discrete time steps t ∈ [0, 1, ..., T − 1]. We assume the
environment has been augmented with a labelling function
uLsϕE w:hSeEth→er {asvaifoel a,tuionnsahfaes}otchcaut,rraetdeawchithti mreespsetecpt,tionfsoormmes
formal safety specification ϕ . For the avoidance of doubt,
we define a violation to have occurred whenever ϕ does not
hold. This is a weaker assumption than previous works in
shielding (which assume access to an abstraction of the
environment) and can be thought as a secondary safety-focused
reward function with a binary output. Intuitively, the goal of
the agent is to learn a policy π that maximises its expected
cumulative reward while minimising the number of
violations of ϕ .</p>
      </sec>
      <sec id="sec-2-2">
        <title>Syntactically Co-Safe Linear Temporal Logic</title>
        <p>
          In this work, we use syntactically co-safe Linear Temporal
Logic (scLTL)
          <xref ref-type="bibr" rid="ref18">(Kupferman and Vardi 2001)</xref>
          as our
speciifcation language. Valid scLTL formulae over some set of
atomic propositions AP can be constructed according to the
following grammar:
        </p>
        <p>ϕ ::= true | d | ¬d | ϕ ∨ ϕ | ϕ ∧ ϕ | ⃝ ϕ | ϕ ∪ ϕ | ⋄ ϕ (1)
where d ∈ AP , ¬ (negation), ∨ (disjunction), ∧
(conjunction) are the familiar operators from propositional logic, and
⃝ (next), ∪ (until) and ⋄ (eventually) are temporal operators.</p>
        <p>
          The main rationale behind our choice of specification
language is the fact that we can efficiently monitor a
system’s adherence to an scLTL specification using a technique
known as progression
          <xref ref-type="bibr" rid="ref5">(Bacchus and Kabanza 2000)</xref>
          . This is
highly advantageous as it means that, given an scLTL
speciifcation, Lϕ can be straightforwardly synthesised.
        </p>
        <p>E</p>
      </sec>
      <sec id="sec-2-3">
        <title>Recurrent State-Space Models</title>
        <p>
          We refer to the predictive model of an environment
maintained by a model-based agent as its world model. World
models can be learnt from experience and be used both as
a substitute for the environment during training
          <xref ref-type="bibr" rid="ref12 ref15">(Ha and
Schmidhuber 2018; Hafner et al. 2021)</xref>
          and for planning at
run-time
          <xref ref-type="bibr" rid="ref13 ref14">(Hafner et al. 2019b)</xref>
          . Though many realisations
of the notion of a world model exist, the world model used
in this work is based on the recurrent state-space model
(RSSM) proposed by
          <xref ref-type="bibr" rid="ref13 ref14">(Hafner et al. 2019b)</xref>
          .
        </p>
        <p>An RSSM is composed of three key components: a latent
dynamics model, a reward model, and an observation model.
These components act on compact states formed from the
concatenation of a deterministic latent state ht and
stochastic latent state zt.</p>
        <p>Latent Dynamics Model The latent dynamics model is
made up of a number of smaller models. First, the recurrent
model ht = f (ht− 1, zt− 1, at− 1) is used to compute the
deterministic latent state based on the previous compact state
and action. From ht and the current observation ot, a
distribution q(zt|ht, ot) over posterior stochastic latent states zt
is computed by the representation model. At the same time,
a distribution p(zˆt|ht) over prior stochastic latent states zˆt
is computed by the transition model, based only on ht.
During training, the transition model attempts to minimise the
Kullback Leibler (KL) divergence between the prior and
posterior stochastic latent state distributions. In doing this, the
RSSM learns to predict future latent states (using the
recurrent and transition models) without access to future
observations.</p>
        <p>Observation Model The observation model computes the
distribution p(oˆt|ht, zt) over observations oˆt for a
particular state. Though not strictly needed, the observation model
can prove useful for visualising predicted future states and
providing a richer training signal.</p>
        <p>Reward Model The reward model computes the
distribution p(rˆt|ht, zt) over rewards rˆt for a particular state.</p>
        <p>
          In practice, the distributions p and q are implemented with
neural networks pθ and qθ respectively, parameterised by
some set of parameters θ . These latent dynamics models
deifne a fully-observable Markov decision process (MDP) as
the latent states in the agent’s own internal model can
always be observed by the agent
          <xref ref-type="bibr" rid="ref13 ref14">(Hafner et al. 2019a)</xref>
          . We
denote the state space of this MDP (comprised of compact
latent states) as SI .
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Shielding</title>
        <p>
          The classical formulation of shielding in RL is given by
          <xref ref-type="bibr" rid="ref2">(Alshiekh et al. 2018)</xref>
          . It assumes access to two ingredients: an
LTL safety specification and abstraction (a MDP model of
the environment that captures the aspects of the environment
relevant for planning ahead with respect to the safety
speciifcation). These ingredients are used to construct a formally
verifiable reactive system that monitors the agent’s actions,
overriding those which lead to violation states.
        </p>
        <p>Proposed by (Giacobbe et al. 2021), bounded prescience
shielding (BPS) avoids the need for hand-crafted
abstractions by exploiting the fact that some agents are trained in
computer simulations. The shield operates by leveraging
access to the program underlying the simulation to look ahead
into future states within some finite horizon. Using BPS
over classical shielding does, however, come with a few
disadvantages. Firstly, it requires access to the simulation at
run-time which may prove difficult to provide (especially in
cases where running the simulation is computationally
expensive). Moreover, an agent using BPS, even when
starting from a safe state, can find itself entering unsafe states in
cases where the number of steps between a violation being
caused by an action and the violation state itself exceeds the
shield’s look-ahead horizon. This is not the case for classical
shielding which resembles BPS with an infinite horizon.</p>
      </sec>
      <sec id="sec-2-5">
        <title>Bounded Safety</title>
        <p>The notion of safety used by BPS is defined over MDPs. For
an arbitrary MDP with states S and actions A, a bounded
trajectory ρ of length H is a sequence of states and actions
a0 a1 an− 1
s0 →− s1 →− . . →.−− − sn comprised of no more than
H states and with the final state sn either being a terminal
state or n = H − 1. We further denote the set of all finite
trajectories starting from some arbitrary state s ∈ S by ϱ(s)
and the set of all bounded trajectories of length H that start
from s by ϱH (s).</p>
        <p>We say a bounded trajectory ρ of length H satisfies H
bounded safety with respect to safety specification ϕ , written
SH (ρ, ϕ ), if and only if for all si ∈ ρ , LϕE (si) = saf e.
Moreover, we can extend the notion of H -bounded safety
over the set of policies: a policy π is H -bounded safe with
respect to ϕ , denoted as SH (π, ϕ ), if and only if for all s
∈ S ,
• either there exists some ρ ∈ ϱH (s) such that SH (ρ, ϕ )
and π (s0) = a0;
• or for all ρ ∈ ϱH (s), ¬SH (ρ, ϕ ).</p>
        <p>In other words, the policy will choose a safe trajectory as
long as one exists. Finally, we formally define a violation
of ϕ to be inevitable in state s0 ∈ S if and only if for all
ρ ∈ ϱ(s0), ¬SH (ρ, ϕ ).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Shielded Dreams</title>
      <p>
        We introduce the notion of latent shielding, a novel class
of shielding approaches that replace abstractions used by
shields with learned latent dynamics models thus allowing
the enforcement of formal safety specifications while
avoiding the need for an explicitly-defined abstraction of the
environment. We further introduce the first such approach,
approximate bounded prescience shielding, a framework for
latent shielding that leverages latent world models learnt
by model-based deep RL (DRL) agents. In this work, our
model-based agent of choice is Dreamer1
        <xref ref-type="bibr" rid="ref15">(Hafner et al.
2021)</xref>
        , which we modify to incorporate shielding into its
data collection, training and deployment phases.
      </p>
      <sec id="sec-3-1">
        <title>Safety RSSM</title>
        <p>We augment the standard RSSM with a labelling function
LϕI : SI → {saf e, unsaf e} which maps latent states st ∈
SI to whether they correspond to states in violation of ϕ .
As with the other models, Lϕ is implemented with a neural
network with a categorical oIutput lθ also parameterised by
θ . This yields an enhanced RSSM (illustrated in Figure 1)
which we will refer to as a safety RSSM (SRSSM).</p>
        <p>
          We train lθ along with the other components of the
SRSSM with the objective
mθ in Lmodel = Lobservation + Lreward + LKL + Lviolation
(2)
where the first three terms are as described in
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref20 ref26 ref6">(Hafner et al.
2019a, 2021)</xref>
          . Lviolation is a new term that we introduce that
1In practice, any model-based agent with a latent dynamics
model can be used.
acts as a weighted binary cross-entropy loss over predictions
by lθ .
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Approximate Bounded Prescience Shielding</title>
        <p>We now integrate the SRSSM as part of a latent shielding
approach which we shall refer to as approximate bounded
prescience shielding (ABPS). Though ABPS is inspired by
BPS, it differs in two key aspects: (1) we approximate the
labelling function Lϕ and environmental dynamics using an</p>
        <p>E
SRSSM; and (2) we sample a fixed number of potential
future trajectories as opposed to exhaustively exploring all
possibilities.</p>
        <p>Thus, our approach can be thought of as an
approximation of some “ideal” bounded prescience shield that uses the
true environmental dynamics and labelling function. The
advantage of the first difference should be obvious: it enables
the shield to learn its own abstraction, removing the need
for hand-crafting or access to a digital environment’s
underlying program. Why the second difference is advantageous
may be slightly less obvious - it’s a heuristic that allows us
to increase the horizon H . By directing the sampling of
trajectories in accordance with states and actions the policy is
biased towards (as opposed to uniformly), it may be
possible to achieve comparable performance to a standard BPS in
foreseeing unsafe states. The intuition behind this is that by
not sampling trajectories the agent is unlikely to take, more
of our computational budget can be dedicated into
ensuring that the agent’s most likely trajectories are safe. In other
words, we will not spend time planning to correct actions
that the agent is unlikely to take.</p>
        <p>More formally, our shielded policy π ∗ can be written:
π ∗ (st) =
π alt(st), otherwise
π (st), if P (lθ (st+1) = 1|π (st)) &lt; ϵ
(3)
where π (st) is the agent’s policy without shielding; st ∈ SI
is some compact latent state; and π alt is an alternative policy
that ensures that, if a violation isn’t already inevitable, the
Algorithm 1: Approximate Bounded Prescience
Shielding in latent space.</p>
        <p>Input: Current compact latent state (h, z),
unshielded action a, horizon H, number of
trajectories to sample N , random noise η ,
alternative policy π alt and threshold ϵ .</p>
        <p>Output: A tuple containing an action and whether or
not the shield had to interfere.
1 Initialise array of zeros Λ with length N .
2 for n = 1..N do
/* Sample a trajectory and check
if it is unsafe. */
3 Imagine trajectory</p>
        <p>ρ = (h, z→)− a (hˆ1, zˆ1) →−a1 ..→.−−− aH− 1 (hˆH , zˆH )
4
where ai = π (hˆi, zˆ ) + η .</p>
        <p>i
if ∃i ∈ [1, 2, ..., H ] such that
lθ (hi, zi) = unsaf e then Λ[ n] := 1.
8
9 end
10 return ⟨a,interfered⟩.
5 end
/* Interfere if probability of</p>
        <p>violation exceeds ϵ .
6 interf ered := f alse.
7 if N1 Pλ ∈Λ λ &gt; ϵ then</p>
        <p>
          a, interf ered := π alt(hi, zi), true.
agent avoids a predicted unsafe state. Though more complex
candidates for π alt (such as selecting the unshielded policy’s
highest-ranked safe action
          <xref ref-type="bibr" rid="ref2">(Alshiekh et al. 2018)</xref>
          ) do exist,
the implementation of π alt in this work simply considers all
the other actions until a safe trajectory is found. In the event
that no safe action is found, the agent takes a random action.
        </p>
        <p>Since the SRSSM models stochasticity, we can sample
multiple futures arising from an action being taken and
derive probabilistic estimates of whether an action will lead
to a violation. In this way, we can estimate the safety of an
action taken in a given state by checking whether the
probability of a violation occurring (inferred by sampling)
exceeds a fixed threshold ϵ (see Algorithm 1). Moreover, the
sampling process can, in practice, be augmented to sample
a wider range of trajectories (less likely to be taken by the
agent) by adding a small amount of noise η ∼ N (0, σ 2) to
actions suggested by the policy during sampling.
Data Collection In this phase, the agent interacts with the
real environment to collect a dataset D ⊆ S × A × O ×
R × { 0, 1} of states, actions, observations, rewards and
violations with which a latent world model can be learned. At
the very start of training, we collect trajectories from S seed
episodes using a random policy; at all other times, we use
the agent’s shielded policy.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Latent World Model Training The goal of this phase is</title>
        <p>to improve the model of the world so that the policy has an
accurate imagined environment to train in. To this end, we
train the latent world model with respect to the objective in
Equation 2 on B data sequences of length L sampled from
D.</p>
        <p>
          Agent Training Agent training is composed of two steps.
First, starting from states from the same B data sequences
from the latent world model training phase, the agent
imagines I trajectories of length H with actions chosen from its
current policy. Next, the unshielded policy π is updated in
an actor-critic fashion as in
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref20 ref26 ref6">(Hafner et al. 2019a, 2021)</xref>
          .
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Key Changes and Contributions We describe and dis</title>
        <p>cuss the key changes we have made that differentiate our
agent from previous approaches.</p>
        <p>Experience Dataset. Elements in the experience dataset
D now contain a binary variable representing whether a
violation has occurred λ t.</p>
        <p>Latent Shielding. Before being sent to the environment
E , actions at generated by the unshielded policy π are routed
through our newly-proposed latent shield (described in
Algorithm 1). The shield returns a new approximately
Hbounded safe action a′t.</p>
        <p>
          Intrinsic Punishment. Whenever a violation occurs
(whether detected by the latent shield or by the
environment), we override the environment’s reward function by
instead assigning a reward p &lt; 0. This discourages the agent’s
unshielded policy π from taking actions that lead to unsafe
states through the standard RL setup. Over time, this means
that the shield will have to interfere less as π becomes biased
away from unsafe states. This is a standard practice in the
shielding literature
          <xref ref-type="bibr" rid="ref2">(Alshiekh et al. 2018)</xref>
          , however new to
the Dreamer family of agents
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref20 ref21 ref26 ref6">(Hafner et al. 2019a,b, 2021)</xref>
          .
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Shield Introduction Schedule We introduce the novel</title>
        <p>notion of a shield introduction schedule Π which can enable
and disable shielding during training. The rationale behind
this is that a shield backed by an inaccurate world model will
incorrectly label some safe states as unsafe, and vice versa.
In some cases, this may prove detrimental to the training
process: suppose lθ happens to be initialised in such a way
that all states are labelled as unsafe. As a result, the shield
can prevent the agent from exploring the environment and
improving its internal model of the environment. This, in
turn, can prevent the labelling function from learning to
correctly differentiate safe states from unsafe states.</p>
        <p>To give the agent time to learn a good world model before
restricting exploration through a learned shield, we augment
the training procedure with Π which aims to gradually
introduce shielding. To our knowledge, this is the first time
such a system has been proposed. Though there exist a wide
*/</p>
      </sec>
      <sec id="sec-3-6">
        <title>Training Regime</title>
        <p>
          We extend the training regime proposed in
          <xref ref-type="bibr" rid="ref13 ref14">(Hafner et al.
2019a)</xref>
          to include the training and application of ABPS.
Though the full training procedure is detailed in Algorithm
2, we provide an overview below.
        </p>
        <p>Overview The training procedure can be split into three
phases: data collection, latent world model training and
agent training (lines 10-22, 4-6 and 7-8 in Algorithm 2
respectively). These phases are cycled through until
convergence.
range of possible implementations of Π , in this work we use
a simple schedule that seems to work well empirically: start
training with shielding disabled and enable shielding once
the world model loss (in particular, the violation loss)
begins to plateau. A detailed comparison of different shield
introduction schedules, however, is beyond the scope of this
study and left for further work.
23 end
end
Add experience to dataset</p>
        <p>T</p>
        <p>D := D ∪ {(ot, at, rt, λ t)t=1}.</p>
        <p>Algorithm 2: Training Dreamer with Approximate
Bounded Prescience Shielding.</p>
        <p>Input: Punishment for violation rpunish, shield
imagination horizon H, number of steps to
imagine ahead for training I, number of seed
episodes S, number of training steps per
episode C, number of real environment
interaction steps per episode T , sequence
length L, batch size B, environment E ,
labelling function Lϕ and shield introduction
schedule Π . E
1 Initialise dataset D with S seed episodes and neural
network parameters θ randomly.
2 while not converged do
/* Learning the internal world</p>
        <p>model and agent policy.
3 for c = 1..C do
4 Draw B data sequences
{(at, ot, rt, λ t)tk=+kL} ∼ D .</p>
        <p>Compute model states ht, zt and zˆt.</p>
        <p>Update θ using representation learning.</p>
        <p>Imagine a trajectory ρ of length I for each
state in each data sequence.</p>
        <p>Update the unshielded policy π based on the
imagined trajectories.
end
/* Collecting data from the</p>
        <p>environment.
for t = 1..T do</p>
        <p>Compute model states ht, zt from history.</p>
        <p>Select action at with the unshielded policy,
adding exploration noise if desired.
if Π has enabled shielding then</p>
        <p>Pass ht, zt and at into Algorithm 1 to
obtain the tuple ⟨a′t, interf ered⟩.
*/
*/
else
end
rt, ot := E .step (a′t).</p>
        <p>a′t, interf ered := at, f alse.</p>
        <p>Check if Lϕ has detected a violation and store</p>
        <p>E
the result as 0 or 1 in λ t.</p>
        <p>
          if λ t = 1 or interf ered then rt := rpunish.
In this section, we compare our ABPS agent against
Dreamer without shielding
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref20 ref26 ref6">(Hafner et al. 2019a, 2021)</xref>
          ,
Dreamer with BPS (Giacobbe et al. 2021), and CPO
          <xref ref-type="bibr" rid="ref1">(Achiam et al. 2017)</xref>
          . We also empirically investigate some
aspects of the internal workings of our agent. A summary of
the environments used can be found below.
        </p>
        <p>Visual Grid World The Visual Grid World (VGW)
environment is a simple deterministic navigation benchmark
with high-dimensional (64 × 64 × 3) visual observations
(see Figure 2) and discrete actions (up, down, left, right and
staying still). The agent’s (in green) task is to navigate to
randomly-placed targets (in black) while avoiding any
unsafe locations (in red). This yields a relatively
straightforward safety specification:</p>
        <p>¬agent in red square ∪ episode ended
where agent in red square is true if and only if the agent
is in an unsafe location, and episode ended is only true at
the end of an episode.</p>
        <p>The reward function used is also quite simple, with a
small penalty term at each time step to encourage movement
(though indeed different forms of intrinsic motivation may
be also used):
100, agent reaches target,


− 40, agent enters unsafe state,
− 10, agent does not move,
− 1, otherwise.</p>
        <p>We experiment with both fixed and procedurally
generated grids over episodes consisting of 500 steps.
Cliff Driver The Cliff Driver (CD) environment is a
symbolic benchmark with stochastic dynamics and continuous
actions. The agent controls the forward acceleration of a
car and is tasked with driving to the edge of a cliff as
quickly as possible without overshooting and falling into
the sea. The car exists on a one-dimensional road with
actions at ∈ [− 1, 1] which correspond to the acceleration of
the car. The car cannot move backwards and its speed vt is
lower-bounded at 0 (thus at &lt; 0 corresponds to braking as
opposed to reversing). The agent starts each episode
stationary at a fixed distance x0 from the edge of the cliff with its
distance at subsequent time steps being written xt.
Stochasticity comes from the fact that at each time step there is a
probability pstick that the car’s controls get “stuck”
meaning that the current action is ignored and replaced with the
previous action.</p>
        <p>Observations from the environment are given as
twodimensional vectors encoding the distance from the edge of
the cliff in one component and the speed of the car in the
other. The safety specification can thus be written:
¬agent fallen off cliff ∪ episode ended
where agent fallen off cliff is true if and only if the
agent has overshot the cliff, and episode ended is only true
at the end of an episode.
(1 − xx0t , agent has not fallen off cliff,
− 5,</p>
        <p>otherwise.</p>
        <p>We experiment with pstick = 0.1 and pstick = 0.5 on
roads 10 units long over episodes consisting of 20 steps.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Performance Evaluation</title>
        <p>We evaluate our agents on three metrics: (1) average reward
per episode at test-time; (2) average number of violations
per episode at test-time; and (3) total number of violations
during training. Results are calculated by averaging over vfie
versions of each agent trained from different seeds.
Training Details Experiments were carried out on a
machine with a single NVIDIA RTX 2080 Ti GPU, an AMD
Ryzen 7 2700X processor and 64GB of RAM. For each
environment, the model-based agents were trained for the same
number of steps; the model-free CPO agents were allowed
to train for longer (2× longer for the VGW environment and
5× longer for the CD environments). The latent shield
horizon H was 2 and 6 for the VGW and CD environments
respectively.</p>
        <p>
          We used the network architectures proposed in
          <xref ref-type="bibr" rid="ref13 ref14">(Hafner
et al. 2019a)</xref>
          . The number of nodes in each layer varied
depending on the environment and can be found in Table 2. We
implemented our encoder for symbolic observations from
the CD environment as a feed-forward network with three
fully-connected hidden layers and ReLU activations. For the
CPO agent, we used the same observation encoders and
policy networks as mentioned above. Moreover, since the
action space of the CD environment is continuous, we
discretised the actions into four bins when performing BPS (this
boiled down to rounding the continuous actions proposed by
the agent to the nearest action in the set {− 1, − 0.1, 0.1, 1}.
Such modifications, however, were not needed for our agent.
        </p>
        <p>Shield introduction schedules were kept relatively
simple. For the VGW environment, the agent started with
shielding disabled. After 10 episodes (including 5 seed
episodes), shielding was enabled every third episode.
After 20 episodes, shielding was enabled every other episode.
Shielding was fully enabled after 30 episodes. In addition,
the unsafe threshold ϵ was decayed linearly from 0.5 to
0.125 over the course of training. For the CD environment,
shielding was enabled after 60 episodes (including 50 seed
episodes).</p>
        <p>Results As can be seen in Table 1, our agent collected
more reward at test-time than both the unshielded and BPS
agents in the VGW environment. Moreover, our agent saw
a seven-fold reduction in test-time violations compared to
the unshielded agent across both the fixed and
procedurally generated environments, averaging at less than one
violation per episode for the fixed VGW. This reduction in
training violations is even more dramatic when compared
to the model-free CPO agent. Our agent outperformed the
other agents at test-time in the most stochastic CD
environment (pstick = 0.5). This is possibly due to the SRSSM
used by our latent shield being better able to capture
nondeterminism than BPS. Nevertheless, the result is still rather
impressive given that the agent using BPS sampled 1024
trajectories at every time step whereas our agent only sampled
20. Plots of training reward and violations can be seen in
Figure 3.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Qualitative Evaluation of Learned Dynamics</title>
        <p>We evaluate the quality of each agent’s latent dynamics
model by observing its trajectory predictions given a
particular starting state and action sequence. Each agent was given
the same starting observation and predefined sequence of 10
actions. The agents took these actions in their respective
latent world models with the states traversed being decoded
for inspection. Results can be seen in Figure 2 and indicate
that, qualitatively, our agent’s decoder performed the best,
accurately predicting 10 frames into the future. One
possible reason for this observation is that the inclusion of the
violation detection objective in the SRSSM encourages the
model to focus on accurately capturing violation states.
Another potential factor is that the BPS agent never actually
enters unsafe states and thus finds it difficult to represent
them.</p>
        <p>Visual
Grid World</p>
        <p>Cliff
Driver</p>
        <p>Flavour</p>
        <p>Fixed
Procedural
pstick = 0.1
pstick = 0.5</p>
        <p>Metric
Testing Reward
Testing Violations
Training Violations</p>
        <p>Testing Reward
Testing Violations
Training Violations</p>
        <p>Testing Reward
Testing Violations
Training Violations</p>
        <p>Testing Reward
Testing Violations
Training Violations</p>
        <p>Latent
We compare the first 100 training episodes of our agent in
the fixed VGW environment with and without a shield
introduction schedule. As with the performance evaluation,
results are calculated over vfie trained agents and a plot of
training reward can be seen in Figure 4. Though both agents
start at roughly the same reward, the agents with a shield
introduction schedule consistently outperforms their
counterparts without shield introduction schedules.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Related Work and Discussion</title>
      <p>
        Latent World Models Learning latent world models from
visual observations has seen growing interest from the RL
community
        <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15 ref22 ref23 ref28 ref30 ref30">(Wahlstro¨m, Scho¨n, and Deisenroth 2015;
Watter et al. 2015; Racanie`re et al. 2017; Ha and
Schmidhuber 2018; Hafner et al. 2019a,b; Schrittwieser et al. 2020;
Hafner et al. 2021)</xref>
        . One key trend in the literature is that of
training agents in their own learned world models
        <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15 ref20 ref26 ref6">(Ha and
Schmidhuber 2018; Hafner et al. 2019a, 2021)</xref>
        . In this work,
we extend the world model formulation used in
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref20 ref21 ref26 ref6">(Hafner et al.
2019a,b, 2021)</xref>
        to encode a notion of safety into state
representations.
      </p>
      <sec id="sec-4-1">
        <title>Safety By Filtering Actions Overriding unsafe actions</title>
        <p>
          with safe ones has been a popular approach to safety-aware
RL. Introduced into the RL scene by
          <xref ref-type="bibr" rid="ref2">(Alshiekh et al. 2018)</xref>
          ,
shielding has received much research interest and has seen
applications in real-world settings
          <xref ref-type="bibr" rid="ref20">(Nikou et al. 2021)</xref>
          .
Various works have attempted to address some of its
shortcomings of the original formulation. These include allowing the
shield to be updated to aid exploration or correct model
imperfections
          <xref ref-type="bibr" rid="ref21 ref3">(Anderson et al. 2020; Pranger et al. 2021)</xref>
          ;
improving performance in non-deterministic environments
          <xref ref-type="bibr" rid="ref16 ref19">(Jansen et al. 2020; Li and Bastani 2020)</xref>
          ; and extending
shielding to multi-agent RL
          <xref ref-type="bibr" rid="ref8">(ElSayed-Aly et al. 2021)</xref>
          . To
our knowledge, few works
          <xref ref-type="bibr" rid="ref26 ref6">(Srinivasan et al. 2020;
Thananjeyan et al. 2021; Bharadhwaj et al. 2021; Giacobbe et al.
2021)</xref>
          focus on removing the need for handcrafting an
abstraction (among the most time-consuming and error-prone
aspects of shielding), and only one of them (Giacobbe et al.
2021) achieves this without getting rid of the abstraction
altogether, albeit by providing the agent with access to the
program that controls the environment.
        </p>
        <p>
          In contrast, latent shielding tackles the problem head-on
by directly learning an abstraction for use by the shield.
In this work, the abstraction we use is an SRSSM, which
captures stochasticity by design, a useful property for
nondeterministic environments. Though our latent shield
satisifes an approximation of H-bounded safety with respect to
the learned abstraction, its safety with respect to the true
environmental dynamics is not guaranteed and instead relies
on the fidelity with which the true dynamics are captured.
Furthermore, unlike for its formally-verified predecessors,
it is a necessary sacrifice that the agent visits unsafe states
during training in order to learn a notion of safety (unless, of
course, pre-training is possible). Nevertheless, a learned
abstraction may be advantageous in settings where
handcrafting an abstraction is not feasible and privileged access to
some simulation (as in (Giacobbe et al. 2021)) cannot be
assumed. Moreover, latent shielding may provide a greater
degree of explainability over model-free methods which get rid
of the abstraction altogether
          <xref ref-type="bibr" rid="ref26 ref6">(Srinivasan et al. 2020;
Thananjeyan et al. 2021; Bharadhwaj et al. 2021)</xref>
          - if the shield
overrides an action, one can reconstruct the imagined unsafe
trajectories that led to the interference. Finally, it should be
noted that the problem setup used in this work assumes no
prior knowledge on how safe behaviour might be achieved
in the environment. This may not be the case in many
realworld settings and it is conceivable that combining latent
shield learning with curriculum learning based on human
knowledge (in a system such as that proposed by
          <xref ref-type="bibr" rid="ref27">(Turchetta
et al. 2020)</xref>
          ) may lead to improved safety during training.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper, we have presented latent shielding, a new
framework for shielding DRL agents without the need for
a handcrafted abstraction of the environment. Using this
framework, we have designed a novel shield and
demonstrated that this method not only leads to improved
adherence to safety specifications on two benchmark
environments with respect to an unshielded agent, but also works
out-of-the-box on both continuous and discrete
environments (unlike its predecessor, BPS). Furthermore, we have
demonstrated for the first time that shielding at
inappropriate times may adversely impact the performance of
modelbased DRL agents and showed how this phenomenon can be
counteracted using our novel notion of shield introduction
schedules.</p>
      <p>Our work opens several exciting avenues for future work.
For instance, this work uses a very simple shield
introduction schedule; future work may provide a richer
investigation into the properties of different schedules. Moreover,
though demonstrating promising empirical results, our
realisation of latent shielding loses the formally verifiable safety
guarantees enjoyed by many symbolic shielding approaches
- whether it is possible to construct a verifiable
implementation of latent shielding, or compensate for the loss of formal
guarantees, are open problems.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The authors would like to thank Claudio Elgueta Karstegl
for turning our GPU machine off and on again during the
national lockdowns.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Achiam</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Held</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tamar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and Abbeel,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Constrained Policy Optimization</article-title>
          . In Precup, D.; and Teh, Y. W., eds.,
          <source>Proceedings of the 34th International Conference on Machine Learning</source>
          , volume
          <volume>70</volume>
          <source>of Proceedings of Machine Learning Research</source>
          ,
          <volume>22</volume>
          -
          <fpage>31</fpage>
          . PMLR.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Alshiekh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bloem</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Ehlers,
          <string-name>
            <given-names>R.</given-names>
            ; Ko¨nighofer, B.;
            <surname>Niekum</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; and Topcu,
          <string-name>
            <surname>U.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Safe Reinforcement Learning via Shielding</article-title>
          .
          <source>In Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>32</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Verma</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dillig</surname>
            ,
            <given-names>I.;</given-names>
          </string-name>
          and
          <string-name>
            <surname>Chaudhuri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>Neurosymbolic Reinforcement Learning with Formally Verified Exploration</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <fpage>6172</fpage>
          -
          <lpage>6183</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Bacchus</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Kabanza</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2000</year>
          .
          <article-title>Using Temporal Logics to Express Search Control Knowledge for Planning</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>116</volume>
          (
          <issue>1</issue>
          ):
          <fpage>123</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bharadhwaj</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Rhinehart</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Levine</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shkurti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2021</year>
          .
          <article-title>Conservative Safety Critics for Exploration</article-title>
          .
          <source>In International Conference on Learning Representations.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Chow</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ghavamzadeh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Janson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and Pavone,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Risk-Constrained Reinforcement Learning with Percentile Risk Criteria</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>18</volume>
          (
          <issue>1</issue>
          ):
          <fpage>6070</fpage>
          -
          <lpage>6120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>ElSayed-Aly</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ; Bharadwaj,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Amato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Ehlers</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Topcu,
          <string-name>
            <given-names>U.</given-names>
            ; and
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2021</year>
          .
          <article-title>Safe Multi-Agent Reinforcement Learning via Shielding</article-title>
          .
          <source>In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems</source>
          , AAMAS '
          <volume>21</volume>
          ,
          <fpage>483</fpage>
          -
          <lpage>491</lpage>
          . Richland,
          <source>SC: International Foundation for Autonomous Agents and Multiagent Systems. ISBN 9781450383073.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          2021.
          <article-title>Shielding Atari Games with Bounded Prescience</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems</source>
          ,
          <volume>1507</volume>
          -
          <fpage>1509</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>ISBN</surname>
          </string-name>
          <year>9781450383073</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Ha</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Recurrent World Models Facilitate Policy Evolution</article-title>
          .
          <source>In Proceedings of the 32nd International Conference on Neural Information Processing Systems</source>
          ,
          <volume>2455</volume>
          -
          <fpage>2467</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Hafner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Lillicrap,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Ba</surname>
          </string-name>
          , J.; and Norouzi,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2019a</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hafner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Lillicrap,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          ; Villegas,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ; and
            <surname>Davidson</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2019b</year>
          .
          <article-title>Learning Latent Dynamics for Planning from Pixels</article-title>
          .
          <source>In Proceedings of the 36th International Conference on Machine Learning</source>
          ,
          <fpage>2555</fpage>
          -
          <lpage>2565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Hafner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Lillicrap,
          <string-name>
            <given-names>T. P.</given-names>
            ;
            <surname>Norouzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Ba</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ; Ko¨nighofer, B.;
          <string-name>
            <surname>Junges</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Serban</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and Bloem,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>Safe Reinforcement Learning Using Probabilistic Shields (Invited Paper)</article-title>
          . In Konnov, I.; and Kova´cs, L., eds.,
          <source>31st International Conference on Concurrency Theory (CONCUR</source>
          <year>2020</year>
          ), volume
          <volume>171</volume>
          <source>of Leibniz International Proceedings in Informatics (LIPIcs)</source>
          ,
          <volume>3</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          :
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Dagstuhl</surname>
          </string-name>
          , Germany: Schloss
          <string-name>
            <surname>Dagstuhl-Leibniz-Zentrum</surname>
            <given-names>fu</given-names>
          </string-name>
          ¨r Informatik.
          <source>ISBN 978-3-95977-160-3.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Kupferman</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ; and Vardi,
          <string-name>
            <surname>M. Y.</surname>
          </string-name>
          <year>2001</year>
          . Formal Methods in System Design,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>291</fpage>
          -
          <lpage>314</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Bastani</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Robust Model Predictive Shielding for Safe Reinforcement Learning with Stochastic Dynamics</article-title>
          .
          <source>In 2020 IEEE International Conference on Robotics and Automation (ICRA)</source>
          ,
          <fpage>7166</fpage>
          -
          <lpage>7172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Nikou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mujumdar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; Orlic´,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and
          <string-name>
            <given-names>Vulgarakis</given-names>
            <surname>Feljan</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2021</year>
          .
          <article-title>Symbolic Reinforcement Learning for Safe RAN Control</article-title>
          .
          <source>In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems</source>
          , AAMAS '
          <volume>21</volume>
          ,
          <fpage>1782</fpage>
          -
          <lpage>1784</lpage>
          . Richland,
          <source>SC: International Foundation for Autonomous Agents and Multiagent Systems. ISBN 9781450383073.</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Pranger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Ko¨nighofer, B.;
          <string-name>
            <surname>Tappler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Deixelberger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Jansen</surname>
            , N.; and Bloem,
            <given-names>R.</given-names>
          </string-name>
          <year>2021</year>
          .
          <article-title>Adaptive Shielding under Uncertainty</article-title>
          .
          <source>In 2021 American Control Conference (ACC)</source>
          ,
          <fpage>3467</fpage>
          -
          <lpage>3474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          Racanie`re, S.; Weber,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Reichert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            ;
            <surname>Buesing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Guez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Rezende</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Badia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            ;
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ;
            <surname>Heess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Pascanu</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Battaglia,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Hassabis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ; and
            <surname>Wierstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Imagination-Augmented Agents for Deep Reinforcement Learning</article-title>
          .
          <source>In Proceedings of the 31st International Conference on Neural Information Processing Systems</source>
          , NIPS'
          <volume>17</volume>
          ,
          <fpage>5694</fpage>
          -
          <lpage>5705</lpage>
          . Red Hook,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA: Curran Associates Inc. ISBN 9781510860964.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Schrittwieser</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Antonoglou</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ; Hubert,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Sifre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Schmitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Guez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Lockhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Hassabis</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          ; Graepel,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ; and
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model</article-title>
          .
          <source>Nature</source>
          ,
          <volume>588</volume>
          (
          <issue>7839</issue>
          ):
          <fpage>604</fpage>
          -
          <lpage>609</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          2020.
          <article-title>Learning to be Safe: Deep RL with a Safety Critic</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>ArXiv</surname>
          </string-name>
          , abs/
          <year>2010</year>
          .14603.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Thananjeyan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Balakrishna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nair</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Luo,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Srinivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            ;
            <surname>Ibarz</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ; Finn,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Goldberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <year>2021</year>
          .
          <string-name>
            <surname>Recovery</surname>
            <given-names>RL</given-names>
          </string-name>
          :
          <article-title>Safe Reinforcement Learning With Learned Recovery Zones</article-title>
          .
          <source>IEEE Robotics and Automation Letters</source>
          ,
          <volume>6</volume>
          (
          <issue>3</issue>
          ):
          <fpage>4915</fpage>
          -
          <lpage>4922</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Turchetta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kolobov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Safe Reinforcement Learning via Curriculum Induction</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          ,
          <volume>33</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          Wahlstro¨m, N.; Scho¨n, T. B.; and Deisenroth,
          <string-name>
            <surname>M. P.</surname>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <article-title>From Pixels to Torques: Policy Learning with Deep Dynamical Models</article-title>
          . arXiv:
          <volume>1502</volume>
          .
          <fpage>02251</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Watter</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Springenberg</surname>
            ,
            <given-names>J. T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Boedecker</surname>
            , J.; and Riedmiller,
            <given-names>M.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images</article-title>
          .
          <source>In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume</source>
          <volume>2</volume>
          ,
          <fpage>2746</fpage>
          -
          <lpage>2754</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          2020.
          <article-title>Projection-Based Constrained Policy Optimization</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>