<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Domain Shifts in Reinforcement Learning: Identifying Disturbances in Environments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tom Haider</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felippe Schmoeller Roza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dirk Eilers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karsten Roscher</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stephan Gu¨ nnemann</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer IKS</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Technical University of Munich</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>A significant drawback of End-to-End Deep Reinforcement Learning (RL) systems is that they return an action no matter what situation they are confronted with. This is true even for situations that differ entirely from those an agent has been trained for. Although crucial in safety-critical applications, dealing with such situations is inherently difficult. Various approaches have been proposed in this direction, such as robustness, domain adaption, domain generalization, and out-of-distribution detection. In this work, we provide an overview of approaches towards the more general problem of dealing with disturbances to the environment of RL agents and show how they struggle to provide clear boundaries when mapped to safety-critical problems. To mitigate this, we propose to formalize the changes in the environment in terms of the Markov Decision Process (MDP), resulting in a more formal framework when dealing with such problems. We apply this framework to an example real-world scenario and show how it helps to isolate safety concerns.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Deep Reinforcement Learning (RL) has been successfully
applied to many different domains, achieving state-of-the-art
performance on a wide range of high-dimensional control
problems [Mnih et al., 2013; Silver et al., 2017; Levine et al.,
2016; Lillicrap et al., 2015]. However, RL systems are still
not commonly applied to safety-critical scenarios. A relevant
factor for this is that RL agents are highly sensitive to
disturbances in the environment they were trained in [Eysenbach
and Levine, 2021]. Even slight changes in the specifications
of the environment can already lead to severe malfunctions or
failures of many RL-based systems.</p>
      <p>Take as an example a humanoid robot that is trained to walk
on a rigid body surface, such as concrete or wood. While
state-of-the-art algorithms can be deployed to train policies
that solve this task [Haarnoja et al., 2018; Schulman et al.,
2017], it is still difficult to ensure that the robot will
function properly if tested on a softer surface such as grass or
sand. Even though the task essentially stays the same, most
RL algorithms struggle with disturbances like these, leading
to potential failures.</p>
      <p>In the classic supervised or unsupervised learning
paradigm, we could frame this problem as out-of-distribution
(OOD). In that context, training samples are typically
assumed to be Independent and Identically Distributed (IID)
whereas for OOD samples, the ‘identically’ assumption is
violated. Strictly applying this definition we could interpret
soft and rigid surfaces to be sampled from non-identical
distribuions and thus a normal operation of the robot can not
be expected. However, solid and soft surfaces could
theoretically come from the same distribution (the distribution of
surfaces) but during training, only a small part of the
distribution was experienced (the part of the distribution where
surfaces are solid). In this case, this is not an OOD instance but
rather a generalization problem. The robot should be able to
‘generalize’ to the entire distribution. In Control Theory and
RL, this is also known as ‘robustness to external influences’.
Note that even this simple example can become ambiguous
and one could think that OOD and robustness are dealing with
the same phenomena.</p>
      <p>In this work, we provide insights towards a better
formalization of disturbances to the environment of RL agents. We
propose a simple but powerful decomposition of any given
disturbance into its aspects of the Markov Decision Process
(MDP). We compare this decomposition to a set of commonly
used terms, that also describe disturbances to the environment
and, by that, aim to disambiguate any commonalities and
differences. To illustrate our method, we apply it to a series of
potential disturbances in an exemplary real-world scenario.
We show that this method helps at identifying potential risks
that can occur during the deployment of RL agents. We see
this as a necessary step towards guaranteeing the safety of RL
agents in non-idealistic problem settings. We do not attempt
at building a unifying framework, that combines all existing
work in this direction. Rather, we want to provide an
alternative view on disturbances in the environments of RL agents,
that should help when approaching this problem and allow to
compare approaches from different related fields.
In this section we formalize the RL framework and present
different concepts related to the problem of dealing with
changes from the training domain to the testing domain.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Preliminaries</title>
      <p>In RL, we consider an agent that sequentially interacts
with an environment modeled as a Markov Decision
Process (MDP) [Puterman, 2014]. An MDP is a tuple M :=
(S; A; R; P; 0), where S is the set of states, A is the set
of actions, R : S A S 7! R is the reward function,
P : S A S 7! [0; 1] is the transition probability function
which describes the system dynamics, where P (st+1jst; at)
is the probability of transitioning to state st+1, given that the
previous state was st and the agent took action at, and 0 :
S 7! [0; 1] is the starting state distribution. At each timestep
the agent observes the current state st 2 S, takes an action
at 2 A, transitions to the next state st+1 drawn from the
distribution P (st; at), and receives a reward R(st; at; st+1).
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Distributional Shift and OOD</title>
      <p>Distributional shift describes changes in data distributions.
When only the input distribution changes while the
conditional distribution of outputs remains unchanged, it is called
covariate shift. When the testing distribution differs from the
training distribution, machine learning systems may not only
exhibit poor performance but also wrongly assume that their
performance is good [Amodei et al., 2016].</p>
      <p>Out-of-distribution (OOD) defines the data that lay outside
the training distribution. More formally, consider that PX and
QX are two distinct data distributions defined on the input
space X . If a model is trained on a dataset drawn from the
distribution PX , samples from this set will be defined as
indistribution while QX is part of the OOD domain [Liang et
al., 2017].</p>
      <p>Although formally defined, mapping real in- and
out-ofdistributions into training and testing sets is challenging for
high-dimensional feature spaces. Using datasets for testing
that contain classes different from those composing the
training set is an established practice in image classification
[DeVries and Taylor, 2018]. Despite being an interesting
approach, it can induce a bias towards a chosen OOD set and
can, nevertheless, only be applied to classification problems.</p>
      <p>In RL, there is not a consensus on how to frame OOD into
practical examples yet. [Sedlmeier et al., 2019] provide one
of the first publications on this topic and define OOD as every
state-action tuple not experienced during training. The
difficulty in generating OOD scenarios due to a lack of controlled
and reproducible datasets in RL is also mentioned.
[Mendonca et al., 2020] use a more relaxed interpretation, defining
OOD as tasks never seen by the agent.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Transfer Learning, Domain Adaptation and</title>
    </sec>
    <sec id="sec-5">
      <title>Domain Randomization</title>
      <p>The objective of transfer learning (TL) is to learn a task TB
that belongs to a target domain B, by utilizing experience
(prior knowledge) from some source task TA from source
domain A, i.e. to transfer knowledge from A to B [Yosinski et
al., 2014].</p>
      <p>Domain adaptation (DA) is a sub-field of transfer learning.
DA can be defined as the capability to deploy a model trained
in one or more source domains into a different target
domain. While DA encompasses cases where the source and
target domains have the same feature space, TL includes cases
where the feature space of the target domain differs from the
source domain [Redko et al., 2019]. Typically, invariance is
assumed, meaning that features that differ between two
domains are irrelevant for the task. [Tzeng et al., 2014] use
this assumption to train a CNN architecture to learn a domain
invariant feature representation that transfers well between
two domains. [Ganin et al., 2016] deploy domain adversarial
training to learn features that are both discriminative for the
main learning task on the source domain and indiscriminate
with respect to the shift between the domains.</p>
      <p>In domain randomization, the source domain is
manipulated at random over a set of parameters in order to train a
more generalizable model. [Sadeghi and Levine, 2016] and
[Tobin et al., 2017] show that randomizing the rendering
settings for a simulator can be used to train policies that
generalize to the real world without requiring the simulator to
be extremely realistic. [Rajeswaran et al., 2016] show that
randomizing over a set of system dynamics can lead to more
robust policies that can generalize to a broad range of
possible target domains, including effects that are not modeled in
the training distribution.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Novelty Detection and Intrinsic Motivation</title>
      <p>Novelty is a known mechanism used by human beings to
explore and learn new things. Many studies in the neuroscience
field bring the converging evidence that novelty can activate
the reward region of the brain in both humans and animals,
playing an important role in their reinforcement learning
process [Houillon et al., 2013].</p>
      <p>In machine learning, novelty describes data that belong to
an unknown pattern [Hajer et al., 2020]. Novelty detection
is oftentimes used to identify outliers and faulty operation
states. In RL, however, identifying undesired operations is
more closely related to distributional shift and OOD detection
while novelty is commonly utilized as a mechanism for
exploration purposes. As shown by [Burda et al., 2018], this
mechanism is also referred to as intrinsic motivation or
curiositydriven exploration. [Lehman and Stanley, 2008] show how
exploration based on novelty can be an efficient way to avoid
getting stuck at local optima regions, a recurrent issue in
problems where climbing the stepping stones that ultimately
lead to the goal is not rewarded by the objective function.</p>
      <p>Despite its importance, seeking novelty brings a
paradoxical consequence: exploring the unknown may lead to a
completely unexpected outcome, including a bad one. Although
learning from bad experiences is valuable, in safety-critical
applications some states are dangerous and must not be
visited. In such cases, safe exploration methods are paramount
to avoid violating the safety constraints.
2.5</p>
    </sec>
    <sec id="sec-7">
      <title>Robust RL</title>
      <p>Robustness is an important topic in ML since trained
models are expected to work despite small deviations in the input
features. Recently, discussions regarding adversarial attacks
have shown that neural networks are vulnerable to
specifically designed small changes in the input which result in
drastic changes in their predictions [Madry et al., 2017]. The
very existence of adversarial attacks may suggest an inherent
weakness of deep learning models, reigniting the interest of
researchers in the search for more robust models. This same
problem can affect RL models, as shown by [Pinto et al.,
2017]. [Nilim and El Ghaoui, 2003] argue that in RL,
robustness is usually related to uncertainties in the transition matrix
of the MDP while [Pinto et al., 2017] shows that RL models
should also be robust to model initializations and modeling
errors.</p>
      <p>Robustness regarding RL can also be viewed by the
optics of control theory since RL can be described as a
dynamical closed-loop system. In control theory, robustness is a
widely studied field that focuses on the control of systems
subject to uncertainty, noise, and disturbances, consisting of
the synthesis (i.e., designing robust controllers) and
analysis (i.e., evaluating the robustness of a controller) problems
[Zhou and Doyle, 1998; Bhattacharyya and Keel, 1995].
Robust control can be used to increase operational safety and
performance on safety-critical applications, such as the
control of nuclear power plants [Jin et al., 2010]. [Calafiore
and Campi, 2006] argue that robust control has received
criticism for being too rigid in describing uncertainty, resulting in
overly conservative controllers and opening the path to
alternative approaches. However, formalizing robustness in RL is
not a trivial task and robust RL algorithms are not yet widely
accepted as a replacement to classical robust control methods.
2.6</p>
    </sec>
    <sec id="sec-8">
      <title>Operational Design Domain</title>
      <p>Operational Design Domain (ODD) is defined by [NHTSA,
2017], in the context of automated driving systems. ODD
is used to describe the specific conditions on which the
system is intended to function safely. Information like roadway
types, geographic area, speed range, among others, should be
included in the ODD and documented accordingly to describe
the desired operation domain. How the system should react
when getting outside of its defined ODD is also a relevant
related topic.</p>
      <p>[Koopman and Fratrik, 2019] show how to use ODD in ML
problems and extend ODD by including object and event
detection and response, vehicle maneuvers, and fault
management. A list including examples relevant for the automated
driving scenario, which fall in each of these categories, is then
provided as a starting point when designing systems based on
ML.</p>
      <p>Despite being originally described for automated driving
applications, ODD provides an explicit framework to specify
the intended deployment domain regarding safety-relevant
aspects and is a powerful tool that can be helpful when
designing RL systems and defining their operation domain.
3</p>
      <sec id="sec-8-1">
        <title>How can these concepts help in RL problems?</title>
        <p>The concepts presented in the previous section are of
relevance when dealing with uncertainties and changes in the
environment. However, the distinction between them is
subTraining Environment
tle, and bringing these concepts down to RL problems is not
straightforward. One approach that could help in that
endeavor is shown in Figures 1 and 2.</p>
        <p>Figure 1 shows the whole problem domain, i.e., the
distribution over environments reachable by the agent during
training and testing. The Training Environment considers
all the situations the agent can encounter during training. It
can alternatively be described as the in-distribution domain.
Designing this region using the ODD framework is helpful
when defining safe boundaries that the agent must respect.
The region corresponding to the already explored portion of
the training domain is hereby defined as the Known
Environment. The last domain is the Testing Environment, which
encompasses the whole region the agent can come across after
deployment. There can be overlaps between the different
domains, as depicted in the figure. The optimal scenario
happens when the whole testing domain is included in the
training region and the agent is able to explore the domain
completely during training, which is often not feasible for large
problems.</p>
        <p>When dealing with the problem of domain shifts in RL,
mapping the regions of Figure 1 into the concepts described
in the previous section can help with its characterization, as
depicted in Figure 2. This mapping is not unique and is not
intended to reflect the whole theory behind domain shifts. It is
rather, in our perspective, the most appropriate for RL
problems subject to changes in the environment. Analyzing the
Known Domain is not the focus of this paper and is
therefore ignored here. However, we would like to point out that
unsafe states can also occur within the known environment
of a system. The region inside the training domain that is
still unknown to the agent is mapped as the Novelty Region,
i.e., the region that can be explored during training. After
deployment, the agent should be robust to novel states, which
belong to the same distribution of the training examples, here
defined as the Robustness Region. The last region is defined
as the OOD region, composed of samples from a different
distribution to the training distribution.</p>
        <p>A drawback of this approach is that it is difficult to draw
specific boundaries to differentiate these regions, especially
because it is not trivial to list the changes to the environment
that can lead to situations outside of the training domain. To
make it even worse, a consensus on how to apply some of
the concepts in RL is still missing. Therefore, different
interpretations may arise based on how the ideas are ported
from more general ML problems into the RL framework. For
instance, if we follow the definition from [Sedlmeier et al.,
2019], everything outside the known environment is
considered OOD. On the other hand, based on the interpretation
from [Mendonca et al., 2020], OOD is composed of tasks
semantically different from the training ones. However,
formally defining semantic differences in the tasks is difficult.
4</p>
      </sec>
      <sec id="sec-8-2">
        <title>Formalizing Domain differences as Aspects of the MDP</title>
        <p>The methods described in section 2 all share a common
characteristic: they deal with problems in which the train/source
task differs from the test/target task. In other words, the
test/target task is disturbed. The difference in these methods
arises from how this disturbance manifests, i.e., how training
and testing tasks relate to each other. As a first step towards
safe RL, we believe it is necessary to gain a better
understanding of this relationship. For this, we borrow an idea that
is commonly used in the Transfer Learning literature [Taylor
and Stone, 2009; Zhu et al., 2021], which we coin MDP
decomposition. That is, any disturbance is sub-divided into the
individual components that build up the MDP. This
decomposition should come naturally since MDPs are used to describe
RL problems in the first place. We want to point out that Safe
RL has more many more dimensions, such as safety during
training or robustness against adversarial attacks. Safety
concerns that arise from differences between training and
deployment environment are only one aspect, which is the aspect
tackled by the hereby proposed approach.</p>
        <p>Two tasks MA and MB, i.e., the training scenario A and
the testing scenario B, can thus be different only in the
following aspects (or a combination thereof):</p>
        <p>S (State-space): Enlargements, reductions or changes to
the set of possible states;
A (Action-space): Additional or fewer actions, changed
ranges of continuous actions;
P (Transition Dynamics): Different transition
probabilities, given the same state-action pairs;
R (Reward Function): Modified reward function that
reinforces different behaviors;</p>
        <p>0 (Initial states): Different initial state distributions.</p>
        <p>Although this may seem like a relatively trivial matter at
first, the benefits of this simple method quickly become
apparent in practice. Training an RL agent to solve tasks while
reliably meeting necessary safety constraints is an intricate
problem. While many safety-critical issues can be addressed
by accurate modeling in the training environment, via domain
randomization or hand-crafted safety controllers, accounting
for every possible safety-critical situation is intractable given
the complexity of the real world. We, therefore, have to make
assumptions on the relationship between training and testing
environments, otherwise guaranteeing safety becomes
impossible. To formulate this relationship we can turn to the
traditional problem formulations mentioned in 2. We find,
however, that framing even simple examples into definitions like
distributional shift, novelty detection, out-of-distribution or
robustness is not only complicated but provides limited
insights towards an appropriate safety argumentation. MDP
decomposition on the other hand is straightforward and allows
us to locate the problem in the individual components of the
MDP.</p>
        <p>It is important to mention that every element that
composes an MDP is formally defined. Therefore, it is possible
to outline operational limits, boundaries, or distributions over
each element. As an example, one could train a robust agent
where, due to degradation on the actuators that might occur
with time, the maximum action can be reduced by up to 20%
of the nominal value. Another example is an agent able to
deal with some sort of noisy sensor, that will change the
observation of the states. In that sense, it becomes clear how to
design test cases that cover the range of operation, solely by
changing the MDP element that is affected by these changes.
This naturally helps with separating concerns, a crucial
component when arguing about safety. Moreover, it is an initial
step towards better comparability in the regime of Safe RL as
it allows us to compare approaches that currently run under
different terms and formulations on a more uniform scale.</p>
        <p>Mapping which aspect of the MDP is affected by a change
in the system relies on a good comprehension of the process.
The state-space can be altered due to novel elements inserted
in the surroundings of the agent. Also, changes in the sensors
(noise, wear and tear, replacement, preprocessing changes)
can lead to a different perception of the state-space. Novel
elements can also have dynamics unseen by the agent,
affecting the transition function. The underlying behavior of
novel elements will also affect the transition function, as they
can have a collaborative or adversarial behavior for instance.
Changes in the action-space are usually related to failures
or degradation of the actuators. Upgrading the system (e.g.,
the attachment of a new actuator in a robot) can also lead
to a different action-space. Mapping changes to the reward
function is a particular case since the function is designed to
translate the task into an objective function that allows the
agent to learn. Therefore, changes in the task itself will cause
changes in the reward function. Also, the reward function can
be changed to guide the agent while learning (reward
shaping). Isolating initial state changes is more straightforward
and self-explanatory.
5</p>
      </sec>
      <sec id="sec-8-3">
        <title>Example Scenario</title>
        <p>To demonstrate the expressiveness of the decomposition
described above, we apply it to an exemplary real-world
scenario, depicted in Figure 3.</p>
        <p>Consider an Automated Guided Vehicle (AGV) navigating
in a warehouse. To determine its location and to identify its
surroundings, it is equipped with an array of sensors (e.g.,
camera, lidar, accelerometer). The primary task goal for the
AGV is to reach the destination position. The safety
constraint is to avoid any collisions with its surroundings. In a
realistic setting, there are several challenges the AGV might
face, such as workers or other robots interacting with the
AGV (collaborative and non-collaborative), multiple goals,
malfunctions of AGV (flat tire, low battery, etc.), or noisy
sensors. For the purpose of this example, we consider all of
these hazards absent during training, i.e., they represent
disturbances to the environment during deployment.</p>
        <p>Table 1 shows that each of these hazards can be traced
back to only a single or at most two components of the MDP.
This mapping is straightforward and directly helps to isolate
safety-relevant issues. Once isolated, these issues can be
detected and handled more easily.</p>
        <p>Conversely, applying the existing formulations covered in
section 2 provides limited insights, as we show in the
following. Apart from noisy sensors, all problems can be framed
as an instance of domain shift, OOD, or novelty,
depending on the definition, as described in 1. Therefore, domain
shift/OOD/novelty as a proxy for safety-critical situations are
not particularly helpful. If the severity of disturbances is
not addressed explicitly, minor disturbances, which are
essentially irrelevant to the safety of the system, are already
Workers interacting
with the AGV
Other robots interacting
with the AGV
Changed warehouse layout
Multiple goals
Unusual starting position
Malfunctions of the AGV
Noisy sensors
X
X
X
X
(X)
X
X
X
X
X
deemed as a safety concern.</p>
        <p>Robustness is typically used when dealing with problems
such as noisy sensors, slight malfunctions of the robot, or
minor external disturbances. However, as soon as the
disturbances are more severe, e.g., workers interacting with the
environment, achieving robustness is usually infeasible.</p>
        <p>ODD can help to define the boundaries of a system more
clearly but it essentially requires anticipating all potential
safety threads during the design of the system. MDP
decomposition can help in this process, by giving deeper insights
into potential causes of safety threads.</p>
        <p>Transfer Learning/Domain adaptation as described in
section 2 actually tackles another problem. They assume the
changes from source to target domain to be static, and not
ad hoc like sudden malfunctions of the robot or interactions
with humans. If we can assume, however, that disturbances
are static, e.g., changes to the layout of the warehouse,
transfer learning approaches might be the tool of choice.
6</p>
      </sec>
      <sec id="sec-8-4">
        <title>Conclusion</title>
        <p>Handling disturbances in the environment is an essential step
towards safe RL systems. We showed that previous work
approaches this problem from various different angles, though
often lacking a clear problem formalization within the RL
domain. Thus, applying existing approaches from other areas to
RL problems is not straightforward. To disambiguate how
changes in the environment of an RL agent can manifest in
a formal task definition, we proposed to decompose complex
problems into the aspects that build up the MDP. This simple
trick allows us to isolate different concerns and treat each of
them separately. We applied this method to a simple obstacle
avoidance task, where a wheeled robot has to navigate in a
warehouse. By listing potential disturbances and analyzing
how they affect the MDP, we laid ground for this framework
to help when dealing with complex real-world scenarios.</p>
        <p>We also see some clear limitations to this approach. Even
if potential disturbances can be located in a single part of
the MDP, it is still a problem of its own to properly handle
them. For the above scenario, consider for example that we
introduce novel obstacles that look the same as the cardboard
boxes but which are also able to move. This would manifest
as a change to the transition function only, since individual
states stay the same and only sequences of states are different
from before. Although posing a severe thread to safety,
detecting such a change is complicated, even when decomposed
into MDP components.</p>
        <p>Future work has to detail further how the MDP
decomposition and structuring of the problem can be utilized for
adding robustness and ensuring safety under environmental
disturbances. We intend to use this formalization and
analyze how we can both detect and handle disturbances in
individual components of the MDP. We believe that this change
of perspective results in a more accurate characterization of
safety critical applications targeted for RL systems and
naturally helps with formalizing safety objectives, bringing us
closer to enable an appropriate validation of RL systems.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Amodei et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Dario</given-names>
            <surname>Amodei</surname>
          </string-name>
          , Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mane´.
          <article-title>Concrete problems in ai safety</article-title>
          .
          <source>arXiv preprint arXiv:1606.06565</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Bhattacharyya and Keel</source>
          , 1995]
          <string-name>
            <given-names>Shankar P</given-names>
            <surname>Bhattacharyya and Lee H Keel</surname>
          </string-name>
          .
          <article-title>Robust control: the parametric approach</article-title>
          .
          <source>In Advances in control education 1994</source>
          , pages
          <fpage>49</fpage>
          -
          <lpage>52</lpage>
          . Elsevier,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Burda et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Yuri</given-names>
            <surname>Burda</surname>
          </string-name>
          , Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and
          <article-title>Alexei A Efros</article-title>
          .
          <article-title>Large-scale study of curiosity-driven learning</article-title>
          .
          <source>arXiv preprint arXiv:1808.04355</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Calafiore and Campi</source>
          , 2006]
          <article-title>Giuseppe C Calafiore and Marco C Campi. The scenario approach to robust control design</article-title>
          .
          <source>IEEE Transactions on automatic control</source>
          ,
          <volume>51</volume>
          (
          <issue>5</issue>
          ):
          <fpage>742</fpage>
          -
          <lpage>753</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[DeVries and Taylor</source>
          , 2018]
          <article-title>Terrance DeVries</article-title>
          and Graham W Taylor.
          <article-title>Learning confidence for out-ofdistribution detection in neural networks</article-title>
          .
          <source>arXiv preprint arXiv:1802.04865</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Eysenbach and Levine</source>
          , 2021]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Eysenbach</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Levine</surname>
          </string-name>
          .
          <article-title>Maximum entropy rl (provably) solves some robust rl problems</article-title>
          .
          <source>arXiv preprint arXiv:2103.06257</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Ganin et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Yaroslav</given-names>
            <surname>Ganin</surname>
          </string-name>
          , Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Franc¸ois Laviolette, Mario Marchand, and
          <string-name>
            <given-names>Victor</given-names>
            <surname>Lempitsky</surname>
          </string-name>
          .
          <article-title>Domain-adversarial training of neural networks</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ),
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Haarnoja et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Tuomas</given-names>
            <surname>Haarnoja</surname>
          </string-name>
          , Aurick Zhou, Pieter Abbeel, and
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Levine</surname>
          </string-name>
          .
          <article-title>Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor</article-title>
          .
          <year>2018</year>
          .
          <article-title>Milestone Algorithm; Stochastic Actor-Critic.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Hajer et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Jan</given-names>
            <surname>Hajer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ying-Ying</surname>
            <given-names>Li</given-names>
          </string-name>
          , Tao Liu, and
          <string-name>
            <given-names>He</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Novelty detection meets collider physics</article-title>
          . Physical
          <string-name>
            <surname>Review</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <volume>101</volume>
          (
          <issue>7</issue>
          ):
          <fpage>076015</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Houillon et al.,
          <year>2013</year>
          ]
          <string-name>
            <given-names>A</given-names>
            <surname>Houillon</surname>
          </string-name>
          , RC Lorenz, W Boehmer, MA Rapp,
          <string-name>
            <given-names>A</given-names>
            <surname>Heinz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Gallinat</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K</given-names>
            <surname>Obermayer</surname>
          </string-name>
          .
          <article-title>The effect of novelty on reinforcement learning</article-title>
          .
          <source>Progress in brain research</source>
          ,
          <volume>202</volume>
          :
          <fpage>415</fpage>
          -
          <lpage>439</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Jin et al.,
          <year>2010</year>
          ]
          <string-name>
            <given-names>Xin</given-names>
            <surname>Jin</surname>
          </string-name>
          , Asok Ray, and
          <string-name>
            <surname>Robert M Edwards.</surname>
          </string-name>
          <article-title>Integrated robust and resilient control of nuclear power plants for operational safety and high performance</article-title>
          .
          <source>IEEE Transactions on Nuclear Science</source>
          ,
          <volume>57</volume>
          (
          <issue>2</issue>
          ):
          <fpage>807</fpage>
          -
          <lpage>817</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[Koopman and Fratrik</source>
          , 2019]
          <string-name>
            <given-names>Philip</given-names>
            <surname>Koopman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Fratrik</surname>
          </string-name>
          .
          <article-title>How many operational design domains, objects</article-title>
          , and events?
          <source>SafeAI@ AAAI, 4</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[Lehman and Stanley</source>
          , 2008]
          <string-name>
            <given-names>Joel</given-names>
            <surname>Lehman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kenneth O</given-names>
            <surname>Stanley. Exploiting</surname>
          </string-name>
          open
          <article-title>-endedness to solve problems through the search for novelty</article-title>
          .
          <source>In ALIFE</source>
          , pages
          <fpage>329</fpage>
          -
          <lpage>336</lpage>
          . Citeseer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Levine et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Levine</surname>
          </string-name>
          , Chelsea Finn, Trevor Darrell, and
          <string-name>
            <given-names>Pieter</given-names>
            <surname>Abbeel</surname>
          </string-name>
          .
          <article-title>End-to-end training of deep visuomotor policies</article-title>
          .
          <volume>17</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1334</fpage>
          -
          <lpage>1373</lpage>
          ,
          <year>2016</year>
          . Publisher: JMLR. org.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Liang et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Shiyu</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Yixuan</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Rayadurgam</given-names>
            <surname>Srikant</surname>
          </string-name>
          .
          <article-title>Enhancing the reliability of out-of-distribution image detection in neural networks</article-title>
          .
          <source>arXiv preprint arXiv:1706.02690</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Lillicrap et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Timothy P.</given-names>
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jonathan J.</given-names>
            <surname>Hunt</surname>
          </string-name>
          , Alexander Pritzel, Nicolas Heess, Tom Erez,
          <string-name>
            <given-names>Yuval</given-names>
            <surname>Tassa</surname>
          </string-name>
          , David Silver,
          <string-name>
            <given-names>and Daan</given-names>
            <surname>Wierstra</surname>
          </string-name>
          .
          <article-title>Continuous control with deep reinforcement learning</article-title>
          .
          <year>2015</year>
          .
          <article-title>continuous Q learning with actor network for approximate maximizatio</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Madry et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Aleksander</given-names>
            <surname>Madry</surname>
          </string-name>
          , Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Vladu</surname>
          </string-name>
          .
          <article-title>Towards deep learning models resistant to adversarial attacks</article-title>
          .
          <source>arXiv preprint arXiv:1706.06083</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Mendonca et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Russell</given-names>
            <surname>Mendonca</surname>
          </string-name>
          , Xinyang Geng, Chelsea Finn, and Sergey Levine.
          <article-title>Meta-reinforcement learning robust to distributional shift via model identification and experience relabeling</article-title>
          .
          <source>arXiv preprint arXiv:2006.07178</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Mnih et al.,
          <year>2013</year>
          ]
          <string-name>
            <given-names>Volodymyr</given-names>
            <surname>Mnih</surname>
          </string-name>
          , Koray Kavukcuoglu, David Silver,
          <string-name>
            <given-names>Alex</given-names>
            <surname>Graves</surname>
          </string-name>
          , Ioannis Antonoglou, Daan Wierstra, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Riedmiller</surname>
          </string-name>
          .
          <article-title>Playing atari with deep reinforcement learning</article-title>
          .
          <source>arXiv preprint arXiv:1312.5602</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[NHTSA</source>
          ,
          <year>2017</year>
          ] NHTSA.
          <source>Automated driving systems 2</source>
          .
          <article-title>0: A vision for safety</article-title>
          . Washington, DC: US Department of Transportation,
          <source>DOT HS</source>
          ,
          <volume>812</volume>
          :
          <fpage>442</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>[Nilim and El Ghaoui</source>
          ,
          <year>2003</year>
          ]
          <article-title>Arnab Nilim and Laurent El Ghaoui. Robustness in markov decision problems with uncertain transition matrices</article-title>
          .
          <source>In NIPS</source>
          , pages
          <fpage>839</fpage>
          -
          <lpage>846</lpage>
          . Citeseer,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [Pinto et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Lerrel</given-names>
            <surname>Pinto</surname>
          </string-name>
          , James Davidson, Rahul Sukthankar, and
          <string-name>
            <given-names>Abhinav</given-names>
            <surname>Gupta</surname>
          </string-name>
          .
          <article-title>Robust adversarial reinforcement learning</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          , pages
          <fpage>2817</fpage>
          -
          <lpage>2826</lpage>
          . PMLR,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Puterman</source>
          , 2014] Martin L Puterman.
          <article-title>Markov decision processes: discrete stochastic dynamic programming</article-title>
          . John Wiley &amp; Sons,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Rajeswaran et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Aravind</given-names>
            <surname>Rajeswaran</surname>
          </string-name>
          , Sarvjeet Ghotra, Balaraman Ravindran, and
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Levine</surname>
          </string-name>
          .
          <article-title>Epopt: Learning robust neural network policies using model ensembles</article-title>
          .
          <source>arXiv preprint arXiv:1610.01283</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Redko et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Ievgen</given-names>
            <surname>Redko</surname>
          </string-name>
          , Emilie Morvant, Amaury Habrard,
          <string-name>
            <given-names>Marc</given-names>
            <surname>Sebban</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Younes</given-names>
            <surname>Bennani</surname>
          </string-name>
          .
          <article-title>Advances in domain adaptation theory</article-title>
          .
          <source>Elsevier</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>[Sadeghi and Levine</source>
          , 2016]
          <string-name>
            <given-names>Fereshteh</given-names>
            <surname>Sadeghi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Levine</surname>
          </string-name>
          . Cad2rl:
          <article-title>Real single-image flight without a single real image</article-title>
          .
          <source>arXiv preprint arXiv:1611.04201</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Schulman et al.,
          <year>2017</year>
          ] John Schulman, Sergey Levine, Philipp Moritz,
          <string-name>
            <given-names>Michael I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Pieter</given-names>
            <surname>Abbeel</surname>
          </string-name>
          .
          <article-title>Trust region policy optimization</article-title>
          .
          <year>2017</year>
          .
          <article-title>deep RL with natural policy gradient and adaptive step size</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [Sedlmeier et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Sedlmeier</surname>
          </string-name>
          , Thomas Gabor, Thomy Phan, Lenz Belzner, and
          <string-name>
            <surname>Claudia</surname>
          </string-name>
          Linnhoff-Popien.
          <article-title>Uncertainty-Based Out-of-Distribution Classification in Deep Reinforcement Learning</article-title>
          .
          <source>December</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Silver et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>David</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Julian</given-names>
            <surname>Schrittwieser</surname>
          </string-name>
          , Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai,
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Bolton</surname>
          </string-name>
          , et al.
          <article-title>Mastering the game of go without human knowledge</article-title>
          .
          <source>nature</source>
          ,
          <volume>550</volume>
          (
          <issue>7676</issue>
          ):
          <fpage>354</fpage>
          -
          <lpage>359</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [Taylor and Stone, 2009]
          <string-name>
            <given-names>Matthew E</given-names>
            <surname>Taylor and Peter Stone</surname>
          </string-name>
          .
          <article-title>Transfer Learning for Reinforcement Learning Domains: A Survey</article-title>
          .
          <source>page 53</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Tobin et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Josh</given-names>
            <surname>Tobin</surname>
          </string-name>
          , Rachel Fong, Alex Ray, Jonas Schneider,
          <string-name>
            <given-names>Wojciech</given-names>
            <surname>Zaremba</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Pieter</given-names>
            <surname>Abbeel</surname>
          </string-name>
          .
          <article-title>Domain randomization for transferring deep neural networks from simulation to the real world</article-title>
          .
          <source>In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          . IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Tzeng et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Eric</given-names>
            <surname>Tzeng</surname>
          </string-name>
          , Judy Hoffman, Ning Zhang, Kate Saenko, and
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Darrell</surname>
          </string-name>
          .
          <article-title>Deep domain confusion: Maximizing for domain invariance</article-title>
          .
          <source>arXiv preprint arXiv:1412.3474</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [Yosinski et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Jason</given-names>
            <surname>Yosinski</surname>
          </string-name>
          , Jeff Clune, Yoshua Bengio, and
          <string-name>
            <given-names>Hod</given-names>
            <surname>Lipson</surname>
          </string-name>
          .
          <article-title>How transferable are features in deep neural networks</article-title>
          ?
          <source>arXiv preprint arXiv:1411.1792</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>[Zhou and Doyle</source>
          , 1998]
          <string-name>
            <given-names>Kemin</given-names>
            <surname>Zhou</surname>
          </string-name>
          and John Comstock Doyle.
          <article-title>Essentials of robust control</article-title>
          , volume
          <volume>104</volume>
          .
          <article-title>Prentice hall Upper Saddle River</article-title>
          , NJ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [Zhu et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Zhuangdi</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Kaixiang</given-names>
            <surname>Lin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jiayu</given-names>
            <surname>Zhou</surname>
          </string-name>
          .
          <article-title>Transfer Learning in Deep Reinforcement Learning: A Survey</article-title>
          . arXiv:
          <year>2009</year>
          .07888 [cs, stat],
          <year>March 2021</year>
          . arXiv:
          <year>2009</year>
          .07888.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>