<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Fear Field: Adaptive constraints for safe environment transitions in Shielded Reinforcement Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haritz Odriozola-Olalde</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nestor Arana-Arexolaleiba</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maider Zamalloa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jon Perez-Cerrolaza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jokin Arozamena-Rodríguez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ikerlan Technology Research Centre</institution>
          ,
          <addr-line>José María Arizmendiarrieta 2, Arrasate-Mondragón, Gipuzkoa, 20500</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mondragon Unibertsitatea</institution>
          ,
          <addr-line>Loramendi 4, Arrasate-Mondragón, Gipuzkoa, 20500</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Shielding methods for Reinforcement Learning agents show potential for safety-critical industrial applications. However, they still lack robustness on nominal safety, a key property for safety control systems. In the case of a significant change in the environment dynamic, shielding methods cannot guarantee safety until their inherent dynamics model is updated to the new scenario. The agent could reach risky states because the model cannot predict well. These situations could lead to catastrophic outcomes, such as damage to the cyber-physical system or loss of human lives, which are not allowed on safety-critical applications. The novel method presented in this paper, Fear Field, replicates human behaviour in those scenarios, adapting safety constraints whenever a drastic environmental change is introduced. Fear Field reduces safety violations by one order of magnitude compared to an RL agent implementing only a shield.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Reinforcement Learning</kwd>
        <kwd>Shielding</kwd>
        <kwd>Adaptive constraints</kwd>
        <kwd>Robustness</kwd>
        <kwd>Safe AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>to minimise unsafe situations that can be risky for the
cyber-physical system, during the exploration process
The design of controllers for autonomous systems has carried out during learning and subsequent execution
emerged in a new era with the remarkable evolution process [1, 5, 7]. Among the diferent methods proposed
of Machine Learning (ML). Techniques such as Super- for this purpose is Shielded Reinforcement Learning.
vised Learning, Unsupervised Learning and Reinforce- In this method, each action proposed by the agent is
ment Learning have shown extraordinary value both for checked, so the Shield only allows the action to be
extheir great adaptability to highly complex problems and ecuted if the environment transit to a safe state. For
reduced computational cost in inference. this, most of the proposed methods use a model of the</p>
      <p>One of the emerging techniques within ML is Rein- environment.
forcement Learning (RL), linked to optimal control the- The shielded RL algorithms proposed in the literature
ory [4]. In RL, an agent interacts with its environment focus mainly on nominal safety, while functional safety is
through the paradigm of trial and error, and exploration relegated to a second level. The nominal safety focuses on
and exploitation [20]. Depending on the performance making safe logical decisions, and it is part of the wider
shown by the agent in a given task, it receives a reward. functional safety considerations that must also consider
During a series of trial and error, the agent learns how underlying hardware and software failures. This work
to maximise the accumulated discounted reward. The focuses on nominal safety, and the complete functional
resulting agent is expected to be able to select the optimal safety considerations must be studied in future research
action for a given state, which will be called the policy. work.</p>
      <p>Safe Reinforcement Learning (SRL) research methods A drawback of the Shielded RL methodology is that
an eventual change in the dynamics of the environment
Macao’23: AISafety-SafeRL 2023 Workshop (IJCAI), August 19–21, can make its transition model obsolete; therefore, the
2023, Macao, SAR, China Shield would not act correctly until the model is updated
* Corresponding author. to the new environment. Another scenario where the
n$arhaondar@iomzoolnad@raikgeornla.end.eus((NH..AOrdarniao-zAorlae-xOollaalledieb)a;); agent may take risky actions due to an outdated model
mzamalloa@ikerlan.es (M. Zamalloa); jmperez@ikerlan.es is transferring a policy trained in the simulator to the
(J. Perez-Cerrolaza); jarozamena@ikerlan.es target or real scenario, known as the Sim2real gap [9, 13].
(J. Arozamena-Rodríguez) In order to improve the safety of Shielded RL
meth0000-0001-6597-0225 (H. Odriozola-Olalde); ods, this paper presents the Fear Field (FF) framework,
(0M00.0Z-0a0m0a2l-l3o3a0)5;-08010008-0(N00.1A-6r3a8n9a--6A4r8eXxo(lJa.lPeiebrae)z;-0C0e0r0r-o0l0a0z3a-);3872-9908 which aims to reduce the number of unsafe states reached
0000-0002-5687-7325 (J. Arozamena-Rodríguez) when a significant modification is produced in the
en© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License vironment dynamic. As humans do, when faced with
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
previously unknown scenarios, Fear Field acts cautiously 2.1. Markov Decision Process
by dynamically adapting the constraints to the situation.</p>
      <p>OpenAI Gym’s [3] slightly modified Frozen Lake A Markov Decision Process is called a sequential
decisionbenchmark and the open-source Skrl library [19] are making system, in which the action  taken in the state
used for testing and validating the Fear Field framework.  determines the immediate reward , the next states
Frozen Lake benchmark consists of a tile-based discrete to transit, and the future rewards to receive. Let it be
environment, where a robot has to learn the path to a tuple ⟨, , , ,  ⟩, where  is the space of states,
reach the goal while avoiding holes (unsafe states) in  is the space of actions,  :  ×  ×  → [0, 1]
its way. Initially, the RL agent is trained with normal is the probabilistic transition function associated to the
environment conditions, where each action makes the environment,  :  × ×  → R determines the reward
robot move only one tile. When the environment’s dy- function and the discount factor  determines the present
namic changes, an action taken leads the robot to move value of future rewards [1].
an additional tile following the same direction. The agent
has been trained using tabular Q-Learning, a temporal- 2.2. Reinforcement Learning
diference-based algorithm. It has been observed that Several learning techniques are found within Machine
an agent trained with only Shielded RL avoids the holes Learning, such as Supervised Learning, Unsupervised
while there are no changes in the environment dynamic, Learning, Imitation Learning and Dimensionality
Reducbut it fails just when the environment dynamic changes. tion [4]. One technique that has recently been showing
After some episodes, it adapts to the new situation. Thus, potential is Reinforcement Learning (RL). RL difers from
an agent trained with shielded RL does not guarantee to other subfields of ML in that it does not require a
sufulfil safety constraints after the environment dynamic pervisor or complete models of the environment [20]. It
changes. is, therefore, of interest in complex problems that span</p>
      <p>The contributions of this work can be summarised as diferent engineering domains and where the only way
follows: to learn the properties of the environment is to interact
• The Fear Field framework is proposed, which with it. RL algorithms are primarily linked to optimal
conidentifies the change in the environment, adapts trol theory [4], as they are based on the interaction of an
the imposed safety constraints to the new sce- agent with its environment through the paradigm of
trialnario and acts accordingly. error and exploitation-exploration [23]. Reinforcement
• The Fear Field framework is defined and validated Learning use ranges from object manipulation control
in the OpenAI Gym’s Frozen Lake environment. problems to the compression of search-based schedulers
• The robustness of the Fear Field to significant en- or optimal controllers, where neural networks reduce
vironmental changes and the improvement over their computational cost [2].
a Shielded RL-based control system safety is eval- The learning process of the agent over the
environuated in an experimental setup. ment, i.e. mapping the states to the actions [16], is carried
out because the agent receives a reward for each action</p>
      <p>The rest of the paper proceeds as follows. Section 2 performed and ultimately tries to maximise the sum of
briefly defines the Markov Decision Process, Q-Learning the rewards received [8, 23]. Such interaction allows
and Shielded Reinforcement Learning. In section 3, re- the agent to learn and ideally generalise the knowledge
lated works to the problem studied in this work are re- gained to deal with inexperienced situations, obtaining
sumed. In section 4, constraints and Shield implemen- a policy determining the agent’s behaviour. The policy
tation are defined. The Fear Field framework is defined can range from a simple relationship, such as a lookup
in section 5. In section 6, the experimentation methodol- table, to complex relationships that require high
comogy is given. Following, results obtained in testing are putation, being in general stochastic relationships. The
discussed in section 7. Finally, in section 8, the main policy can be identified as the core of an RL agent because
conclusions and future work are summarised. it determines the behaviour of the agent [20].</p>
      <p>While the reward is a short-term indication of how
good an action performed by the agent is, the Q-value
2. Preliminaries function  (, ) defines the expected discounted
accumulated reward following policy  after taking action 
Reinforcement Learning is one of the most popular tech- in state . Because this long-term estimation is
computaniques of Machine Learning, where an agent that inter- tionally expensive, the correct method choice is
considacts with its environment learns through the paradigm of ered a key component of most RL algorithms [20].
trial-error and exploration-exploitation [23]. The
mathematical idealisation of Reinforcement Learning
algorithms is the Markov Decision Process (MDP).</p>
      <sec id="sec-1-1">
        <title>2.3. Q-Learning</title>
        <p>Q-Learning is a temporal-diference-based control
algorithm in which the Q-value function (, ) directly
approximates its optimal value * , maximising it regardless
of the policy (of-policy) being applied. It should be noted
that, even so, the policy determines which action-state
is visited and updated at each instant [23]. For a time
instant  the Q-Learning algorithm is defined as follows:
(, ) ←</p>
      </sec>
      <sec id="sec-1-2">
        <title>2.4. Shielded RL</title>
        <p>[︁
(, ) +  +1 + 
· max (+1, ) − (, )]︁

(1)</p>
        <p>Learning</p>
        <p>Agent
State</p>
        <p>Reward
Environment
Action</p>
        <p>Reward
Safe
action</p>
        <sec id="sec-1-2-1">
          <title>The reactive Shielded RL methods are among the various</title>
          <p>methods proposed to ensure the safety of the controlled
cyber-physical system. Shielded RL was proposed by
Alshiek et al. [1]; it defines the use of a shield, which
acts as a filter to block those actions that transit the
environment to an unsafe state (See Figure 1).</p>
          <p>Through safety specifications a safety automaton   =
(, 0, , ) is defined where 0 correspond to the
initial state and  is the set of safe states [1, 12].</p>
          <p>The Shield monitors the action  proposed by the
agent at each time step. It checks the expected state +1
after applying an action  in state . If the expected
state +1 is unsafe with respect to  , the Shield will
block the action and ofer another safe action using the
safe policy  . This safe policy   is defined in advance
using particularly severe constraints.</p>
          <p>
            Another aspect to consider is whether the Shield
should penalise the agent’s proposal of an unsafe action
in the reward. As summarised by Odriozola-Olalde [
            <xref ref-type="bibr" rid="ref11">17</xref>
            ],
several authors [6, 11] see favourable to introducing a
penalty so that a bias is generated in the agent’s policy
that reduces the need for shield intervention in the future.
          </p>
          <p>Still, other authors [1, 2, 14, 21] find the introduction of
the penalty mentioned above detrimental.
the shield defined for an initial environment, they
propose the synthesis of a new shield for a modified
environment. But they do not analyse what happens to the agent
during the time needed to synthesise the new Shield, as
the agent is partially safe or not guaranteed to be safe.</p>
          <p>Bastani [2] modifies the cart-pole environment
increasing the time horizon of the controller to demonstrate
that Model Predictive Shielding (MPS) still fulfils the
safety guarantees obtained on the initial environment.</p>
          <p>Although, the cumulative reward obtained decreases
significantly, reducing the performance of the controller at
the cost of assuring safety.</p>
          <p>Thumm and Althof [21] proposed method, Failsafe
Planner, is the only one that considers Functional Safety
standards in the Shielded RL field. They propose a
framework to guarantee the speed and separation monitoring
for human-robot interaction environments defined in
DIN EN ISO 10218-1 2021, 5.10.3. A combination of a
lowfrequency RL agent and a formal high-frequency safety
verification algorithm is used to synthesise a safety shield.</p>
          <p>
            Even though Failsafe Planner is able to avoid all
humanrobot collisions, it has a goal-reaching success rate of
65%. Modelled human movements on experimentation
3. Related Work are quite limited, so safety is not guaranteed in more
complex environments. Also, they do not study how
Although Shielded RL is a relatively recently developed Failsafe Planner could behave in a dynamic changing
framework, many studies [1, 2, 21] propose diferent environment, thus the robustness of Failsafe Planner is
methodologies for implementing shielding in decision- not assured.
making control systems. These methodologies show Lazarus et al. [15] propose a similar approach to Fear
promising results for ensuring nominal safety but lack Field for Runtime Safety Assurance (RTSA) in order to
experimentation in environments that may sufer drastic ensure the nominal safety of an Unmanned Aerial Vehicle
dynamic changes [
            <xref ref-type="bibr" rid="ref11">17</xref>
            ]. This situation emphasises the (UAV). Using Safety Envelopes, a subspace of the state
analysis of the robustness of proposed controllers in en- space defined through safety constraints, they shrink
vironments that may sufer dynamic changes. it specifying a distance  . While the UAV is inside the
          </p>
          <p>Zhu et al. [22] consider the efectiveness of their pro- shrunk subspace, a   black-box nominal policy is used.
posed method in diferent environments. Starting from But once the UAV exits the shrunk subspace, a   simple
and safe recovery policy is deployed. A Reinforcement the predicted state deviation is lower than the
threshsLweaiprneifnrgomagoennet pisolticrayintoedanionthoerrd. eTrhteodcrhawoobsaeckwshoefnthtios oulndti⃒⃒⃒lno− meanin⃒⃒⃒g≤ful  diifenreanlclestbeeptswoefenthempoevreimode,nit.e.
proposition are that  is a hyperparameter that may be predicted by the model and the actual movement of the
dificult to tune, it has no adaptability at all and that the agent is found.
recovery policy   consists of turning of the UAV rotors At each step, the shield studies the feasibility of each
and deploying a parachute. possible action, predicting the next state +1 using the</p>
          <p>Therefore, most works do not study how they perform dynamics model and observing if it belongs to the
ensemin significantly changing environments, so it is neces- ble of safe states . Following, it orders the actions
sary to study the lack of safety guarantees that could by its expected reward. The selected action to be applied
be shown when the environment dynamic changes and,  is selected if and only if it is safe and its expected
rethus, during the time the new shield becomes available. ward is the highest of all safe actions. The environment
Also, the verifiability of the proposed algorithms must transition is predicted over a finite time horizon ℎ:
be studied in order to obtain formal safety guarantees.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Problem Setup</title>
      <sec id="sec-2-1">
        <title>This section defines the techniques used to define the constraints and the methodologies used to synthesise the Shield.</title>
        <sec id="sec-2-1-1">
          <title>4.1. Constraints</title>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>In this work, it has been decided to establish the safety constraints using a tabular table or constraint table (4). The table has the exact dimensions as the GridWorld used in the experimentation.</title>
        <p>() =
︂{ 0
1
 ∈/ 
ℎ
(2)
where  () is the constraint table,  is the state and
 is the aforementioned set of safe states.</p>
        <p>This implies that it is necessary to know in advance the
constraints of the environment, which sometimes cannot
be known due to the complexity of the cyber-physical
system to be controlled and changes in the environment’s
dynamics. In this work, it is assumed that  and
 are known, as it is a reach-avoid problem.</p>
        <sec id="sec-2-2-1">
          <title>4.2. Shield</title>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>The shield implementation contains a model of the dy</title>
        <p>namics of the environment  (+1|, ) [2, 6, 11, 14,
21]. The model input is the proposed action, and the
output is the predicted state that the robot will reach on
the next state.</p>
        <p>Each time a deviation further than a predefined
threshold  is detected between the next state prediction and
the real one ⃒⃒⃒  − ⃒⃒⃒ &gt;  , the environment dynamic
model is updated. This threshold  is defined to avoid that
insignificant deviation that the model can have relative to
the real environment dynamic will afect in runtime. New
data  (′|, ) is first collected in , and then it is
used to update the model. This process is repeated until
+ℎ ←
 (+ℎ|, +1, ..., +ℎ− 1, )
(3)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Fear Field</title>
      <p>As humans adapt the caution measures taken in our
activities according to our confidence and knowledge of the
environment at a specific moment, Fear Field proposes
adapting the safety constraints depending on the Shield’s
confidence in the environment’s model accuracy. The
diference between the model’s predicted state and the
real one quantifies the model’s accuracy.</p>
      <p>
        In an initial environment, for an agent with a shield
and where there is an accurate model of the environment,
an action  taken in the state  transits the
environment to state +1. In this case, the Shield can predict
whether this transition is safe since the associated model
matches the initial environment. Suppose now that the
dynamics of the environment have changed so that the
model associated with the Shield does not match reality.
In the second case, the same action , let’s call ′, taken
in state  causes the environment to transit to a state
′+1 which can be an unsafe or a hidden unsafe state
[
        <xref ref-type="bibr" rid="ref11">17</xref>
        ].
      </p>
      <p>Therefore, the use of the Fear Field (See Figure 2) is
proposed, which is reflected as an extension of the safety
constraints defined in the problem. Fear Field has been
designed in order to support the Shield when the model
predictions are not accurate. For this purpose, a new
constraint table is defined such that:
 ′() =
⎧ 0
⎨0.5
⎩ 1</p>
      <p>∈/ 
 ∈   
ℎ</p>
      <p>(4)</p>
      <p>The width of the Fear Field  () ∝ ⃒⃒⃒ − 1 − −1⃒⃒⃒
is directly proportional to the distance between the
pre</p>
      <p>and the real state . Thus, the width
dicted state 
 () is a dynamic variable linked to the diference
between the predicted state and the real one. In the case
of  () having diferent values through timesteps, the
highest value is taken as the worst-case scenario. Fear
Field states   is the set of states where each state  
is within a maximum Euclidean distance  () from each
unsafe state such that:
table  ′() is generated (Line 4). Since the predicted
state does not match the state reached, the environment
dynamic model is outdated. If  number of steps
has been passed from the last timestep   that
the model was updated (Line 5), then, the model is
updated with the new dataset (Line 6) and   is
 ∈   ⇔ ∃ : | − | ≤  () (5) restored to the current timestep (Line 7).</p>
      <p>Only when the model is updated such that it is capable</p>
      <p>Once the new constraints have been defined, the shield
analyses if the action ‘ would transit the environment porfemdiacttcehdisntgataelsl (lLaisnte9),the isnteitpiasl ocfonsstatrtaeinretatacbhleeda(nd)
to either an unsafe state or a state that is part of the Fear is loaded again. This allows the shield to check the safety
Field  ′(+1) = 0.5. Even though the model is outdated of the actions proposed by the agent (Line 12) and choose
due to that the environment dynamic has changed, the the action that has the highest Q-value and is safe (Line
Shield will predict the transition to the state +1 which 13). Finally, once the action is applied to the environment
is now within Fear Field so that it will block that action and it transits to the new state +1 (Line 14), the last
′ and will look for an action ′′ that does not lead to   of states reached and predicted states are saved
an unsafe state or a state within the Fear Field, i.e. it in a bufer (Line 15).
will look for  ′(′′+1) = 1, and has the highest Q-Value Note that the Fear Field algorithm is being executed
Q(s,a). when the RL agent is trained, and also when it is in</p>
      <p>Only when the model is retrained, and the diference execution; thus, the model updating process is conducted
between the model and the reality is non-existent af- independently of the agent’s learning process.
ter  steps, the width  of the Fear Field will be
zero again. It is necessary to be noted that during the
Fear Field intervention, the agent keeps following the 6. Experiments
same previous policy and gathers  steps data as
a retraining dataset. The algorithm used for Fear Field As mentioned above, the benchmark environment used
implementation is shown on Algorithm 1, which is exe- for validating the Fear Field method was an OpenAI
Gridcuted continuously, and the schematic representation is World [3]. Specifically, a modified version of Frozen
shown in Figure 3. Lake was used. This benchmark environment consists</p>
      <p>The Fear Field algorithm works as follows: In every of a robot that starts in a box in one corner of the
twoiteration, the previous timestep state reached and the dimensional state space and must reach the goal in the
predicted state are compared (Line 2). If they are not opposite corner. If the robot falls into a hole (), it
equal, Fear Field’s width  +1() is calculated using the has to start the episode again. So, the Frozen Lake
envidiference existing between the state reached and the ronment is a reach-avoid problem. The environment grid
predicted state (Line 3). Using  +1(), a new constraint size is 10x10 blocks, with 16 of them being holes. The
unique feature of the modified version of Frozen Lake movement is not taken into account for retraining the NN
is that periodically, the world is slippery, meaning that model, and also while applying the relative movement
the robots will move one additional square for the same calculated by the NN to the current state, it is taken into
action. It is necessary to be noted that the environment account if the predicted state will be out of the border.
used for experimentation is deterministic. The testing set, composed of 50 trials of 700 episodes</p>
      <p>The open-source modular library Skrl, which inte- each, has been performed to obtain a significant number
grates several RL algorithms and supports OpenAI Gym, of results in order to validate the algorithm presented
has been used to implement the benchmark [19]. The in this paper. Each episode is terminated if the agent
software code provided by Skrl has been modified to in- falls into a hole, reaches the goal or if it takes more than
corporate both the Shield and the Fear Field. Q-Learning 100 steps. In the first 350 episodes (Non-slippery), the
algorithm has been used as an RL learning algorithm, robot moves only one square for each action. The
followspecifically a tabular Q-learning algorithm, as it matches ing 150 episodes correspond to the slippery world
(Slipthe discrete nature of the benchmark environment. pery). Finally, in the last 150 episodes, the robot moves</p>
      <p>In order to the environment dynamic model, a Neu- only one square again for each action (Non-slippery).
ral Network (NN) based model has been chosen due to The testing procedure consists of analysing the
perforits capacity to adapt to the changes in the dynamics of mance and nominal safety level of the non-shield RL
the environment, its accuracy and the reduced compu- algorithm, a shielded RL baseline, and the Fear Field
intetational cost in inference. Also, training over the past grated shielded RL algorithm. The first one (Q-Learning)
model reduces the computational cost associated with is the tabular Q-Learning [20] without any safety
meathe relearning process [22]. sures. The second (Shield) corresponds to the previous</p>
      <p>The output of the model has been defined as the rel- Q-Learning algorithm with a shield integrated. Finally,
ative movement of the robot. This way, the required the third one (FF) is based on the second one but
incorminimum NN topology has been reduced to one hidden porates the Fear Field framework.
layer with 12 neurons on it. Also, working with relative The values of the hyperparameters are shown in Table
movement helps to reduce the problem’s complexity and, 1. If a mismatch is detected between the state predicted
consequently, the needed data and training time. The by the NN model and the actual state, 1000 samples are
ifnite time horizon used is ℎ = 1. This way, the shield collected, and the network is retrained using batches of
is capable of predicting the next step state adding to the 6 samples for 100 epochs. The equation used to calculate
actual state the predicted relative movement: the width  () of Fear Field is the following one:
 =  +   ()
+1</p>
      <p>(6)</p>
      <sec id="sec-3-1">
        <title>It is necessary to be noted that grid world border states</title>
        <p>+1() = ⃒⃒⃒  − ⃒⃒⃒
(7)</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>7. Results</title>
      <sec id="sec-4-1">
        <title>The testing results are shown in Figure 4. Three phases</title>
        <p>are shown in the x-axis, corresponding to the
environment state: Slippery or non-slippery. Also, a pink interval
is shown on the y-axis, indicating an unsafe state has
been reached. It is observed that both the Shield and Fear</p>
      </sec>
      <sec id="sec-4-2">
        <title>Field methods accelerate the convergence of the agent</title>
        <p>by almost five times in the first training process. This
is because the Shield allows the agent to explore safely,
significantly improving the number of steps performed
in each episode, thus gaining more knowledge of the
environment per episode.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Another significant advantage of shield-assisted learn</title>
        <p>
          ing is that no safety breaches (red zone) are committed
during the initial (first 350 episodes) learning process.
In the Q-Learning test, the robot persists in an insecure
state for the first 150 episodes.
served that both Q-Learning and Shielded Q-Learning
sufer from reaching unsafe states in the episodes
immediately after the change made in the environment’s
dynamic (episode 350). As hypothesised in the shielded
reinforcement learning review [
          <xref ref-type="bibr" rid="ref11">17</xref>
          ], the prediction made
by the model does not match the actual movement made
by the agent, so it cannot predict the state to which the
agent will transit correctly, and the Shield will not block
the unsafe action; obtaining a 0.0156% probability of
reaching an unsafe state. As can be seen, the agent adapts
to the new environment over time because the model
associated with the shield is updated. Despite this, the
average number of unsafe states reached compared to
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Q-Learning without the Shield is approximately 50 times lower (see Table 2).</title>
      </sec>
      <sec id="sec-4-5">
        <title>For the Fear Field method (FF), it can be observed that</title>
        <p>In the slippery period (episodes 350-550), it can be ob- to predict future states that the environment will transit,
the transition from the normal environment to the
slippery one is performed with a significant reduction of the
unsafe state visited. However, it should be noted that in
some trials, the Fear Field approach still encountered
unsafe states (see Table 2). Specifically, in</p>
        <p>60% of the trials
performed, FF obtained a null number of unsafe states.</p>
        <p>This is because sometimes the retrained NN model is not
accurate enough and therefore fails to predict the state
transition, obtaining a 0.00179% probability of reaching
an unsafe state in total. Thus, the reason behind visiting
unsafe states is the inaccuracies associated with the NN
model. Despite this, one order of magnitude is reduced
compared to the Shielded Q-Learning case.</p>
        <p>One of the main shortcomings of the Fear Field is that
after transitioning to a previously unknown environment,
the convergence time of the policy is increased. This
behaviour is related to the agent being more constrained
in taking action and taking fewer risks. Due to that, the
previously learned path cannot be taken, so the agent
must learn a new path to reach the goal. The learning
process to obtain a safe path to the goal can take some
episodes to be learned.</p>
        <p>Also, no policy convergence has been observed in 2 of
the 50 trials performed when the environment changes
to non-slippery after being slippery. This phenomenon
is due to the model not being updated properly, keeping
the robot transiting to an unsafe state constantly. Thus,
no useful dataset needed to update the model properly is
obtained, and the agent enters a non-ending cycle.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Thus, the dataset obtained does not</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>8. Conclusions and future work</title>
      <sec id="sec-5-1">
        <title>Shielded Reinforcement Learning is a method of great</title>
        <p>interest in control and decision-making fields because it
drastically reduces the number of unsafe states reached.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Since Shield uses a dynamic model of the environment</title>
        <p>the Shield’s efectiveness is low when faced with changes
in the environment’s dynamics. This problem persists
until the model associated with the Shield adapts to the
new environment.</p>
      </sec>
      <sec id="sec-5-3">
        <title>In order to reduce this impact, the Fear Field method</title>
        <p>has been proposed and validated in this paper.
Adapting the safety constraints in proportion to the diference
between the model prediction and the real transitions,</p>
      </sec>
      <sec id="sec-5-4">
        <title>Fear Field is able to reduce the unsafe states reached by order of magnitude when compared to the Shielded RL</title>
        <p>Non-slippery
Slippery
method.</p>
        <p>Looking ahead, the retraining of the model needs to be
improved to reduce the number of unsafe states further
reached. Also, procedures must be developed to address
the observed increase in convergence time, such as
introducing a small value of exploration rate  when the
robot cannot find a new path to the goal.</p>
        <p>Regarding the applicability of the Fear Field
framework, a use case of an autonomous guided vehicle (AGV)
in a warehouse scenario is proposed as future work. The
behaviour of the AV will be observed in a scenario with
significant changes in the wheels grip, and therefore Fear
Field will be tested.</p>
        <p>For the experimentation conducted in this paper, a
deterministic environment is used. In future work, the
effects of the introduction of stochasticity must be studied.
the Fear Field framework shows potential for more
complex scenarios, where model capacity might be limited
in order to capture the environment dynamics perfectly.
In this case, the Fear Field framework could be helpful in
reducing safety constraint violations.</p>
        <p>Shielded RL is an AI-based safety-focused algorithm
that must still be developed in compliance with
applicable safety standards (e.g., IEC 61508, ISO 5469).
Therefore, the integration of methodologies such as Safety
Envelopes, certified according to the required safety
standard, may be interesting in order to provide the
decisionmaking controller with formal safety guarantees. For
many research areas, there are methodologies developed
to define the Safety Envelope through safety standards,
e.g. Responsibility-Sensitive Safety (RSS) for autonomous
driving vehicles [10, 18].</p>
        <p>Finally, further research must be conducted to find
how the values of the hyperparameters regarding the
Shielded RL and Fear Field frameworks afect the safety
constraint violation rate.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>9. Acknowledgments</title>
      <sec id="sec-6-1">
        <title>We want to thank Alexander Lenail for allowing the use of NN SVG to generate NN topology diagrams.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Mohammed</given-names>
            <surname>Alshiekh</surname>
          </string-name>
          et al. “
          <article-title>Safe Reinforcement Learning via Shielding”</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          . Vol.
          <volume>32</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          IEEE.
          <year>2021</year>
          , pp.
          <fpage>3488</fpage>
          -
          <lpage>3494</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Greg</given-names>
            <surname>Brockman</surname>
          </string-name>
          et al. “
          <article-title>OpenAI Gym”</article-title>
          .
          <source>In: arXiv preprint arXiv:1606.01540</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Carleo</surname>
          </string-name>
          et al. “
          <article-title>Machine learning and the physical sciences”</article-title>
          .
          <source>In: Reviews of Modern Physics</source>
          <volume>91</volume>
          (
          <issue>4</issue>
          ) (
          <year>2019</year>
          ), p.
          <fpage>045002</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Steven</given-names>
            <surname>Carr</surname>
          </string-name>
          et al. “
          <article-title>Safe Reinforcement Learning via Shielding for POMDPs”</article-title>
          . In: arXiv preprint (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Ingy ElSayed-Aly</surname>
          </string-name>
          et al. “
          <article-title>Safe Multi-Agent Reinforcement Learning via Shielding”</article-title>
          .
          <source>In: Proceedings of the International Joint Conference on Au- [18] tonomous Agents and Multiagent Systems, AAMAS</source>
          <volume>1</volume>
          (
          <year>2021</year>
          ), pp.
          <fpage>483</fpage>
          -
          <lpage>491</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Javier</given-names>
            <surname>García</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Fernández</surname>
          </string-name>
          .
          <article-title>“A Comprehensive Survey on Safe Reinforcement Learning”</article-title>
          .
          <source>In: Journal of Machine Learning Research</source>
          <volume>16</volume>
          (
          <year>2015</year>
          ), pp.
          <fpage>1437</fpage>
          -
          <lpage>1480</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Mirco</given-names>
            <surname>Giacobbe</surname>
          </string-name>
          et al. “
          <article-title>Shielding Atari Games with Bounded Prescience”</article-title>
          .
          <source>In: Proceedings of the International Joint Conference on Autonomous Agents and Multiagent Systems, AAMAS</source>
          <volume>3</volume>
          (
          <year>2021</year>
          ), pp.
          <fpage>1495</fpage>
          -
          <lpage>1497</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>In: AIAA Scitech 2020 Forum</source>
          . Vol.
          <volume>1</volume>
          PartF. American Institute of Aeronautics and Astronautics Inc, AIAA,
          <year>2020</year>
          , pp.
          <fpage>386</fpage>
          -
          <lpage>403</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Ichiro</given-names>
            <surname>Hasuo</surname>
          </string-name>
          . “
          <article-title>Responsibility-Sensitive Safety: an Introduction with an Eye to Logical Foundations and Formalization”</article-title>
          .
          <source>In: arXiv:2206.03418</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [17] [19] [20] [21] [22]
          <string-name>
            <given-names>Peter</given-names>
            <surname>He</surname>
          </string-name>
          , Borja G León, and
          <string-name>
            <surname>Francesco</surname>
          </string-name>
          Belar- [23]
          <fpage>dinelli</fpage>
          . “
          <article-title>Do Androids Dream of Electric Fences? Safety-Aware Reinforcement Learning with Latent Shielding”</article-title>
          . In: SafeAI@ AAAI (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Floris den Hengst</surname>
          </string-name>
          et al. “
          <article-title>Planning for potential: eficient safe reinforcement learning”</article-title>
          .
          <source>In: Machine Learning 111.6</source>
          (
          <issue>2022</issue>
          ), pp.
          <fpage>2255</fpage>
          -
          <lpage>2274</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Kai-Chieh Hsu</surname>
          </string-name>
          et al. “
          <article-title>Sim-to-Lab-to-Real: Safe Reinforcement Learning with Shielding and Generalization Guarantees”</article-title>
          .
          <source>In: Artificial Intelligence</source>
          <volume>314</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Nils</given-names>
            <surname>Jansen</surname>
          </string-name>
          et al. “
          <article-title>Safe Reinforcement Learning Using Probabilistic Shields”</article-title>
          .
          <source>In: 31st International Conference on Concurrency Theory (CONCUR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Lazarus</surname>
          </string-name>
          , James G Lopez, and
          <string-name>
            <surname>Mykel J Kochenderfer.</surname>
          </string-name>
          “
          <article-title>Runtime Safety Assurance Using Reinforcement Learning”</article-title>
          .
          <source>In: AIAA/IEEE 39th Digital Avionics Systems Conference (DASC)</source>
          .
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>In: American Control Conference (ACC)</source>
          (
          <year>2022</year>
          ), pp.
          <fpage>1808</fpage>
          -
          <lpage>1813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Haritz</given-names>
            <surname>Odriozola-Olalde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Maider</given-names>
            <surname>Zamalloa</surname>
          </string-name>
          , and
          <string-name>
            <surname>Nestor</surname>
          </string-name>
          Arana-Arexolaleiba.
          <article-title>“Shielded Reinforcement Learning: A review of reactive methods for safe learning”</article-title>
          .
          <source>In: IEEE/SICE International Symposium on System Integrations</source>
          .
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Mateusz</given-names>
            <surname>Orłowski</surname>
          </string-name>
          et al. “
          <article-title>Safe and Goal-Based Highway Maneuver Planning with Reinforcement Learning”</article-title>
          .
          <source>In: Advanced, Contemporary Control: Proceedings of KKA - The 20th Polish Control Conference</source>
          . Springer,
          <year>2020</year>
          , pp.
          <fpage>1261</fpage>
          -
          <lpage>1274</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Antonio</given-names>
            <surname>Serrano-Muñoz</surname>
          </string-name>
          et al. “
          <article-title>skrl: Modular and Flexible Library for Reinforcement Learning”</article-title>
          .
          <source>In: arXiv:2202.03825</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Richard S</given-names>
            <surname>Sutton and Andrew G Barto</surname>
          </string-name>
          .
          <article-title>Reinforcement Learning: An Introduction</article-title>
          .
          <source>Tech. rep</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Jakob</given-names>
            <surname>Thumm</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Althof</surname>
          </string-name>
          . “
          <article-title>Provably Safe Deep Reinforcement Learning for Robotic Manipulation in Human Environments”</article-title>
          .
          <source>In: International Conference on Robotics and Automation (ICRA)</source>
          .
          <year>2022</year>
          , pp.
          <fpage>6344</fpage>
          -
          <lpage>6350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          New York, NY, USA: Association for Computing Machinery,
          <year>2019</year>
          , pp.
          <fpage>686</fpage>
          -
          <lpage>701</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>In: IEEE Transactions on Intelligent Transportation Systems</source>
          (
          <year>2021</year>
          ), pp.
          <fpage>14043</fpage>
          -
          <lpage>14065</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>