<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applying Action Masking and Curriculum Learning Techniques to Improve Data Eficiency and Overall Performance in Operational Technology Cyber Security using Reinforcement Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alec Wilson</string-name>
          <email>Alec.Wilson@uk.bmt.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William Holmes</string-name>
          <email>William@adsp.ai</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryan Menzies</string-name>
          <email>Ryan.Menzies@uk.bmt.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kez Smithson Whitehead</string-name>
          <email>Kez.SmithsonWhitehead@uk.bmt.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ADSP</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>BMT</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In previous work, the IPMSRL environment (Integrated Platform Management System Reinforcement Learning environment) was developed with the aim of training defensive RL agents in a simulator representing a subset of an IPMS on a maritime vessel under a cyber-attack. This paper extends the use of IPMSRL to enhance realism including the additional dynamics of false positive alerts and alert delay. Applying curriculum learning, in the most dificult environment tested, resulted in an episode reward mean increasing from a baseline result of -2.791 to -0.569. Applying action masking, in the most dificult environment tested, resulted in an episode reward mean increasing from a baseline result of -2.791 to -0.743. Importantly, this level of performance was reached in less than 1 million timesteps, which was far more data eficient than vanilla PPO which reached a lower level of performance after 2.5 million timesteps. The training method which resulted in the highest level of performance observed in this paper was a combination of the application of curriculum learning and action masking, with a mean episode reward of 0.137. This paper also introduces a basic hardcoded defensive agent encoding a representation of cyber security best practice, which provides context to the episode reward mean ifgures reached by the RL agents. The hardcoded agent managed an episode reward mean of -1.895. This paper therefore shows that applications of curriculum learning and action masking, both independently and in tandem, present a way to overcome the complex real-world dynamics that are present in operational technology cyber security threat remediation.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Reinforcement Learning</kwd>
        <kwd>Cyber Security</kwd>
        <kwd>Artificial Intelligence</kwd>
        <kwd>Operational Technology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In previous work, the IPMSRL environment [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] was developed with the aim of training defensive RL
agents in a simulator representing a subset of an IPMS on a maritime vessel under a cyber-attack. This
work explores the impact of changing the dificulty of the simulator through the manipulation of values
representing real world dynamics, e.g. False negative rate of alerts.
      </p>
      <p>RL agents will often be significantly limited if they are not exposed to the environment in which
they are intended to be deployed. This also translates to when trained agents are deployed in a real
scenario. Consequently, the environment needs to replicate the real scenario as closely as possible.</p>
      <p>This paper extends the use of IPMSRL to enhance realism including the additional dynamics of false
positive alerts and alert delay. Additionally, three configurations of the environment of varying degrees
of dificulty are defined and tested to understand the diferent levels of performance a trained RL agent
can reach.</p>
      <p>This paper also applies curriculum learning and action masking as ways to mitigate the increased
levels of dificulty, showing that using these techniques is data eficient and leads to higher mean episode
reward.</p>
      <p>
        Curriculum learning alone is shown to increase mean episode reward [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Action masking alone
is shown to similarly improve mean episode reward [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Action masking also has the additional
benefit of significantly faster training and the ability to constrain an agent’s available action space to a
set that meets user-defined criteria.
      </p>
      <p>Finally, both curriculum learning and action masking are applied together. This training method
resulted in the highest level of mean episode reward observed in this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Reinforcement Learning</title>
        <p>
          RL is a training method where an agent learns to interact with an environment to complete a task.
The environment is a Markov Decision Process (MDP) and consists of a state space, action space, reward
function, and a transition model [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. At each timestep an agent will take an action on the environment.
The agent will then receive a reward from the reward function and the updated state of the environment.
The goal of RL is for the agent to learn to maximise the reward signal. We choose to use RL over other
AI methods as it allowed the agent to learn without the need for existing datasets.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. IPMSRL</title>
        <p>
          IPMSRL [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is a Gymnasium-based [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] Reinforcement Learning (RL) environment that simulates an
Integrated Platform Management System (IPMS) on a vessel under a cyber-attack. An IPMS controls
and monitors many ship systems across propulsion, power, steering, stability, auxiliary and ancillary
systems. To achieve this, IPMS utilises a distributed control system architecture that facilitates interfaces
with sensors, equipment, plants, software-based control systems and network-based data [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. A physical
representation of the IPMSRL environment has been shown in Figure 2. The configuration of IPMSRL
used in this paper controls a subset of these systems and focuses on the propulsion and chilled water
system. The defending RL agent receives intrusion detection alerts on components following simulated
infection by a cyber attacker. These alerts are based upon the MITRE ATT&amp;CK ICS framework1 [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
These alerts are then passed onto the defending agent through the observation space which the agent
uses to represent the environment state, , as shown on Figure 1. Subsequently, the agent chooses
a discrete action,  contain, eradicate, recover, or wait, for a given node. Contain, eradicate, and
recover represents an action space modified from NIST SP-800-61 guidance, which was adapted to
an Operational Technology (OT) scenario [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. An instantaneous timestep, , takes place and the
environment produces a reward, +1, and updated state +1.
        </p>
        <p>
          This sequence is then repeated until all critical nodes are compromised which would result in a
negative reward or all infections are completely removed which would result in a positive reward. We
also implemented an early stopping criterion of 50 timesteps.
1© 2024 The MITRE Corporation. This work is reproduced and distributed with the permission of The MITRE Corporation.
2.3. PPO
Proximal Policy Optimisation (PPO) is one of the current state-of-the-art algorithms used within RL and
Multi-Agent Reinforcement Learning (MARL) [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. PPO applies the policy gradient method within
the actor-critic architecture. The algorithm was chosen as it is robust to hyperparameter tuning and
has been shown to be performant in previous work [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The actor, or namely the policy network,
chooses the action given an observation. The critic, or namely the value network, produces an estimate
of the sum of future rewards given the action from the policy network and state. This can be simplified
to the actor chooses an action and the critic assesses its quality. The hyperparameters and architecture
used in our experiments have been provided in the appendix.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.4. Curriculum Learning and Action Masking</title>
        <p>
          Curriculum Learning and Action Masking are two popular guided RL methods to address data eficiency
concerns in RL [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. This paper explores both forms of guided RL for the IPMSRL cyber
security environment [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], initially individually and subsequently in combination, which we show further
improves performance.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.5. Curriculum Learning</title>
        <p>
          Curriculum Learning (CL) in RL is the process of increasing the dificulty of a task by periodically
shaping aspects of the MDP throughout training [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. This type of guided RL is often implemented
by altering aspects of the environment such as the complexity of the transition function to increase the
dificulty of optimising the reward function [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. As a result, CL can be considered a form of transfer
learning and has been shown to be beneficial for sim-to-real applications [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          CL has been shown to reduce learning time and improve performance of the trained agent for
applications including robotics and games [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. In this paper, we explore if CL ofers similar
benefits in the domain of cyber security by testing it on the IPMSRL environment [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. We introduce
three stages to the curriculum (Easy, Medium and Hard) where dificulty is defined as the uncertainty
in the transition function. We show CL outperforms training directly on the Hard environment
configuration.
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>2.6. Action Masking</title>
        <p>
          Action Masking (AM) is a guided RL method which allows the integration of additional human knowledge
into the learning process [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. AM limits the range of actions by introducing guardrails which can
prevent undesirable actions from being chosen by the agent. These guardrails require action space
shaping which has similar disadvantages to reward shaping, such as increased manual set up time and
susceptibility to human bias [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. However, a key advantage of guardrails is that they can provide both
increased data eficiency and user-defined constraints which are both essential considerations for cyber
security applications [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>
          Action masking limits the set of actions that can be executed per timestep based on the current
environment state. In the context of masking for discrete spaces, this is the process of reducing the
choice of actions so that undesirable or impossible actions are unavailable to the agent. Specifically, this
is achieved by setting the probabilities of selecting the undesirable actions to zero or near zero during
stochastic learning [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          Previous work has shown action masking can simplify learning for the agent and result in reduced
training times in video games including StarCraft II and DOTA 2 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The focus of the work
was primarily to address the high dimensionality of the action space as opposed to implementing safety
critical constraints, but existing work [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] has shown how action masking can be applied in RL to
improve safety for trafic-based applications.
        </p>
        <p>The data eficiency benefits of action masking in cyber security applications with discrete action
spaces were explored in this paper. Specifically, we mask invalid actions on the IPMSRL environment
and show that the learning process provides a higher average return, as the agent focuses on learning
only from the valid set of actions. This helped prevent the agent from wasting time exploring trajectories
and taking undesirable actions that would not be applicable for deployment. In future work, the use of
action masking could be further extended to restrict available actions based on safety-critical criteria.</p>
        <p>We show below how action masking can also be applied to improve the realism of the IPMSRL
simulation, including full details of how masking was applied in our experiments. For example, in the
real world if the agent has sent a command to contain a node, then the user would likely have to wait
for this process to complete before sending a diferent command to the given node.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.7. Combined Curriculum and Action Masking</title>
        <p>
          AM and CL often aim to achieve the same benefits of safer learning and reduced training time. We show
below how these methods can be combined to improve performance over each method individually.
Existing work has shown how automatic action masking can be used as a type of CL to alter the action
space [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. In contrast, we show action masking can be applied to the action space and used with vanilla
CL as a method to address the uncertainty in the environment’s transition function.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments</title>
      <sec id="sec-3-1">
        <title>3.1. Environment Dificulty Configurations</title>
        <p>In a real scenario, Security Information and Event Management (SIEM) systems are used to give increased
visibility of an OT system and flag any potential malicious activity. SIEMs are not perfect and sufer
from False Positives (FP) and False Negatives (FN). The IPMSRL environment used in this paper has the
additional feature of FP alerts.</p>
        <p>
          Three defined environment dificulties were explored: easy, medium and hard: Table 1 shows the
values chosen for each parameter and dificulty. The easy configuration has no FP or FN alerts, no
alert delays and a 100% action success rate. The medium and hard configurations add dificulty to the
environment by amending the FP, FN, action success probabilities and the alert delay. In previous work,
it was demonstrated that as alert and action success probabilities decreased (FN increasing), a reduction
to the agent’s performance was observed [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          The alerts in the IPMSRL environment are based on the MITRE ATT&amp;CK ICS Tactics [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], with each
tactic representing a diferent type of alert. These tactics are given a tactic level to represent them
with the first tactic listed given a level of 1, with the following tactics increasing in level incrementally.
The tactics are: Initial Access, Execution, Persistence, Privilege Escalation, Evasion, Discovery, Lateral
Movement, Collection, Command and Control, Inhibit Response Function, Impair Process Control and
Impact.
        </p>
        <p>The alert delay parameter breaks the 12 tactics into 3 sections, displayed in Table ??. In the easy
dificulty environment there is no delay present, in the medium dificulty environment there is a small
delay to alerts earliest in the attack on a given node, the hard dificulty environment has a larger delay
for the earliest stage tactic alerts and a small delay for tactics level 5-8. The rationale behind earlier
stage tactics receiving a larger delay is that, generally, these tactics are likely to have a higher threshold
which will need to be reached before malicious activity on a given node or network can be detected
and reported as an alert, as the activity associated with these tactics is harder to diferentiate from user
activity. Therefore, as the environment increases in dificulty, the alert delay increases towards a more
realistic scenario.</p>
        <p>It is necessary to point out that although IPMSRL has added support for more realistic and complex
dynamics, it is still an abstract representation of an IPMS and an attack on this system. The diferent
dificulties of environment add this realism to test the performance of diferent training approaches and
algorithms, but further work needs to be completed on IPMSRL before it can be considered representative
of a real-world system.</p>
        <p>
          All of the experiments conducted in this paper use a PPO [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] based agent, trained for 2.5 million
timesteps over 4 seeds with a 95% confidence interval (CI). The hyperparameters are available in the
appendix.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Hardcoded Defender</title>
        <p>
          A basic hardcoded defender was designed with logic developed alongside a cyber security expert. This
hardcoded defender is not intended to be perfect, but in the absence using cyber security experts to
act as the defender remediating the threat within the IPMSRL environment, the hardcoded defender
deploys solid logic based on the NIST SP-600-61 guidance [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The hardcoded defender is therefore
able to provide us context as to what level of mean episode reward represents a “good” performance
with validated logic. The hardcoded defender takes one action per timestep, as an RL agent does in all
of the experiments explored in this paper. The defensive hardcoded agent uses the following logic to
determine which action to take next:
1. Contain infectable node that is connected to a critical node, with an alert.
2. Recover ofline critical node.
3. Contain infectable node with an alert.
4. Eradicate infectable node that is connected to critical node, with an alert.
5. Eradicate infectable node with an alert.
6. Recover infectable node that is in a contained state.
        </p>
        <p>7. Wait.</p>
        <p>The logic follows the fundamental idea of sequentially containing, eradicating and then recovering a
node, with some modifications to prioritise recovering ofline critical nodes and remediating nodes that
are ‘closer’ to critical nodes first.</p>
        <p>The hardcoded defender was able to reach a mean episode reward over 10,000 episodes of 0.988
in an easy environment, 0.883 in a medium dificulty environment and -1.895 in a hard environment
configuration. Figure 9 shows the comparison of the best performing RL defenders and the hardcoded
defender.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Baseline Results for Vanilla PPO</title>
        <p>The baseline results of a single agent acting in the IPMSRL environment with varying degrees of
dificulty are shown in Figure 2. In the easy configuration, the agent reached near optimum performance
after 1 million timesteps, winning every episode and achieving an extremely high episode reward mean
value of 0.977. For the medium configuration, the agent reached an episode reward mean of 0.104 and
in the hard configuration, the agent struggled to perform well, resulting in an episode reward mean of
-2.791. These baseline results for vanilla PPO are all below that of the episode reward mean achieved by
the hardcoded defender.</p>
        <p>
          The results in Figure 3 show that there are clear challenges for the agent to learn an optimum policy
using a standard training process with PPO [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] when the dificulty of the environment
configuration increases. For this reason, we explored methods which enabled the data eficiency and overall
performance of the agent to improve significantly.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Curriculum Learning</title>
        <p>As discussed in the background section, CL is a training method which, in this implementation, allows
the agent to initially explore simpler tasks where the agent is expected to explore more efectively,
before the task is changed to a more dificult one. This enables the incremental increase in the dificulty
of the environment during the agent’s training.</p>
        <p>
          The tasks can be changed at pre-defined points based on either standard RLlib training metrics e.g.
mean episode reward, or custom metrics which are calculated in the IPMSRL environment e.g. win rate.
This is implemented through RLlib’s TaskSettableEnv API [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>For all the CL results reported in this paper, the number of total timesteps was used to change tasks
in each training sample. The number of total timesteps was chosen to be the point at which the agent,
at the current task dificulty, had reached a ‘stable’ level of performance e.g. the agent’s mean episode
reward had plateaued. Therefore, for diferent training curricula, tasks are changed at diferent points
depending on how quickly the agent can reach a stable level of performance. During this experiment,
the tasks were changed at 850k timesteps and 1.7m timesteps. Approximately a third of the total training
was completed at each task dificulty, once the task’s training had converged. The light blue lines in
Figure 5 show the points at which the tasks were changed, and performance subsequently dropping
sharply. Figure 4 shows the curriculum used in this experiment. Figure 5 compares the results of an
agent trained on the hard environment configuration (Figure 3), with an agent trained via a curriculum
of easy, medium then hard dificulty configurations.</p>
        <p>
          In Figure 5, it is clearly shown that the agent trained via CL can reach a significantly higher level of
performance within the training time when compared to the baseline of “vanilla” training of a PPO agent
in the hard environment configuration. The episode reward mean reached was -0.569 in comparison to
the baseline of -2.791. The CL agent also significantly outperforms the hardcoded defensive agent’s
performance in a hard environment, where it scored an episode reward mean of -1.895. This behaviour
is expected and supports the related literature’s conclusions [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and the intuitive premise that
by leveraging the previous learning of transferable skills in simpler environments to more complex
environments, an agent will perform better than attempting to learn this behaviour from scratch in a
far more complex environment.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Action Masking</title>
        <p>
          The implementation of action masking encompasses two primary components: modifications within
the environment and the development of custom models that incorporate action masking. The
implementation was adapted from examples provided within RLlib [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
        <p>Action masking utilises a binary array (mask) to identify permissible actions at each step. For this
purpose, a specialised class is created, which, upon initialisation, reads a configuration file to ascertain
the applicable mask conditions. This class features a method that is called at every environment step to
evaluate which actions should be masked based on the current state of the environment. The observation
space, structured as a dictionary, facilitates the mask’s transfer to the model, by including “action mask”
and “observations” keys. The “action mask” is a 1D binary array indicating valid and invalid actions,
while “observations” provide a conventional representation of the environment’s state.</p>
        <p>
          For the custom model incorporating discrete action masking a PyTorch implementation is used, which
is compatible with RLlib and functional for DQN and Policy-Gradient style algorithms, such as PPO
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. This model initialises an internal fully connected network that processes solely the observation
component of the observation space. In the forward pass, it computes unmasked logits by feeding the
observation component through this internal network. Additionally, the action mask is transformed
into an infinite mask, setting valid actions to 0 and invalid actions to a large negative value (efectively
negative infinity), ensuring invalid actions are highly unlikely to be selected post-softmax application.
The logits derived from the internal model are then added to this infinite mask, enabling the selection
of a non-masked action. This process for action masking can be integrated with other custom models,
such as a centralised critic model, allowing for both centralised critic functionality and action masking.
        </p>
        <p>All instances of action masking reported in this paper used the masking conditions based on input
from a cyber security expert to reflect realistic cyber defence constraints and logic. These masks are
deliberately simple to avoid over-engineering and unnecessary bias. The implemented masks are:
• If there is no alert on an infectable node, mask the contain and eradicate action on that node.
• If the infectable node is contained, mask the contain action on that node.
• If the infectable node is not contained, mask the eradicate and recover action on that node.
• If an action is already in progress on an infectable node, mask all actions on that node.
• If a critical node is online, mask the recover action on that node.</p>
        <p>Figure 6 shows the baseline results of the easy, medium and hard environment configurations when
action masking is applied, compared to the baseline results without action masking. In all instances,
both a dramatic improvement in data eficiency and overall performance compared to vanilla PPO can
be seen. In the easy configuration, for all seeds, the agent reached an optimum level of performance
in less than 100k timesteps, winning every episode, resulting in an episode reward mean of 0.977. In
the medium dificulty configuration, the agent reached a good level of performance, an episode reward
mean of 0.816.</p>
        <p>In the hard environment configuration with an action mask, the agent reached a higher level of
performance than solely training on the hard environment, as shown in Figure 7, with an episode
reward mean of -0.743 in comparison to the baseline hard environment episode reward mean of -2.791.
The agent trained with an action mask achieved a slightly lower mean episode reward mean than CL,
but was able to reach that episode reward mean at a sharper rate of learning. The AM agent, similarly,
to the CL agent, significantly outperformed the hardcoded agent in the hard environment.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Action Masking and Curriculum Learning</title>
        <p>Following the conclusions of the previous sections in this paper AM and CL were applied together with
the subsequent curriculum shown in Figure 8.</p>
        <p>The curriculum outlined in Figure 8 shows that the tasks are switched much earlier than the curriculum
for the experiment without action masking presented in Figure 4. This is because, as shown in Figure 6,
the learning of policies trained with action masking plateaued far sooner than policies trained without
action masking. The curriculum is consequently adapted to match the attributes of agents trained with
action masking, changing tasks at 250k and 750k timesteps.</p>
        <p>In Figure 9 it can be observed that when the techniques of CL and AM are combined, for all seeds, an
agent trained in these conditions reached a higher mean episode reward of 0.137, in fewer timesteps,
than agents trained solely with CL, AM or without either. The combination of CL and AM was therefore
shown to produce the highest level of performance in comparison to the other training techniques
tested and the hardcoded defender.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>This paper demonstrates that the application of action masking and curriculum learning individually
improved the overall performance of training a defensive agent to remediate attacks in more complex
IPMSRL environment configurations. This includes real work dynamics such as false positive and
negative alerts and the delay that is inherent in OT systems. The benefits of applying these techniques
were even more pronounced when they were applied together.</p>
      <p>Curriculum learning alone was shown to increase episode reward mean from -2.791 to -0.569. Action
masking alone is shown to similarly improve mean episode reward. In this paper an improvement to
-0.743 was seen.</p>
      <p>Finally, both curriculum learning and action masking are applied together. This training method
resulted in the highest level of mean episode reward observed in this paper, with a mean episode reward
of 0.137.</p>
      <p>As the complexity of the environment increased, the hardcoded defender struggled to maintain a
strong level of performance. The hardcoded defender was able to outperform vanilla PPO, but the
application of AM or CL training techniques enabled the RL defensive agent to achieve a significantly
higher episode reward mean than the hardcoded defender in the hard dificulty environment.</p>
      <p>A potential reason that the hardcoded defender struggled in the hard environment, in comparison
to the best performing RL agents, is the uncertainty that a high proportion of FP alerts provides. The
hardcoded agent struggles to prioritise the remediation of true alerts as there is no mechanism in its
logic to establish whether an alert is a FP or not. It will have to randomly choose between which alert
to remediate if there are multiple alerts active. The only prioritisation that is baked into the logic of the
hardcoded defender is to focus on nodes adjacent to critical nodes. The RL agents on the other hand
may have developed policies that allow them to more eficiently decide which alerts to prioritise. This
behaviour is very dificult and time consuming to encode into a hardcoded defender’s logic, further
displaying the benefits of using an RL based approach for autonomous cyber security.</p>
      <p>An important note about the use of curriculum learning is the significance of defining appropriate
task changing criteria. This paper implemented a simple curriculum, changing tasks at a specified
number of total timesteps. There is further scope to optimise this process and potentially see higher
levels of performance.</p>
      <p>Additionally, there is a trade-of present when using action masking; the masking conditions used
during training need to be present when querying the trained policy. This is a drawback in the sense
that it adds a dependency to the agent’s deployment, and additional bias is added through the setting of
masking conditions. But the benefit of constraining certain actions which don’t meet the requirements
set out when developing the masking conditions is a tangible one. The use of action masking in this way
therefore benefits from gains in data eficiency, overall performance, and an ability to restrict actions to
meet the user-defined requirements of a given system. In future work the use of action masking could
begin to consider the safety-critical nature of OT systems. Action masking could further be used to
help to build trust in autonomous agents that aim to be applied to real systems.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>Research funded by Frazer-Nash Consultancy Ltd. on behalf of the Defence Science and Technology
Laboratory (Dstl) which is an executive agency of the UK Ministry of Defence providing world class
expertise and delivering cutting-edge science and technology for the benefit of the nation and allies.
The research supports the Autonomous Resilient Cyber Defence (ARCD) project within the Dstl Cyber
Defence Enhancement programme.</p>
      <p>The authors would also like to thank Lisa Gralewski, Marco Casassa Mont, David Foster, Clare Jubb,
Laura Caddy, Tasha Hughes and Jake Rigby for their wider contribution to the project and paper.</p>
      <p>Hyperparameters Values</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Menzies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Morarji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Foster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Casassa</given-names>
            <surname>Mont</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Turkbeyler</surname>
          </string-name>
          , L. Gralewski,
          <article-title>MultiAgent Reinforcement Learning for Maritime Operational Technology Cyber Security</article-title>
          ,
          <source>CAMLIS: Conference on Applied Machine Learning in Information Security</source>
          <volume>3652</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Narvekar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leonetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sinapov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , P. Stone,
          <article-title>Curriculum learning for reinforcement learning domains: A framework and survey</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Skinner</surname>
          </string-name>
          , Reinforcement today.,
          <source>American Psychologist</source>
          <volume>13</volume>
          (
          <year>1958</year>
          )
          <fpage>94</fpage>
          -
          <lpage>99</lpage>
          . URL: https://doi.apa. org/doi/10.1037/h0049039. doi:
          <volume>10</volume>
          .1037/h0049039.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ewalds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bartunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Georgiev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Vezhnevets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yeo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Makhzani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Küttler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Agapiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schrittwieser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gafney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schaul</surname>
          </string-name>
          , H. van Hasselt,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Calderone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Keet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Brunasso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lawrence</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekermo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Repp</surname>
          </string-name>
          , R. Tsing,
          <string-name>
            <surname>StarCraft</surname>
            <given-names>II</given-names>
          </string-name>
          :
          <article-title>A New Challenge for Reinforcement Learning</article-title>
          ,
          <year>2017</year>
          . URL: http: //arxiv.org/abs/1708.04782, arXiv:
          <fpage>1708</fpage>
          .04782 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] OpenAI, :,
          <string-name>
            <given-names>C.</given-names>
            <surname>Berner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Brockman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dębiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dennison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Farhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hashme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hesse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Józefowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Olsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pachocki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. P. d. O.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Raiman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salimans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schlatter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sidor</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wolski</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. Zhang,</surname>
          </string-name>
          <article-title>Dota 2 with Large Scale Deep Reinforcement Learning</article-title>
          ,
          <year>2019</year>
          . URL: https://arxiv.org/abs/
          <year>1912</year>
          .06680. doi:
          <volume>10</volume>
          .48550/ARXIV.
          <year>1912</year>
          .
          <volume>06680</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ontañón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Closer</given-names>
            <surname>Look at Invalid Action</surname>
          </string-name>
          <article-title>Masking in Policy Gradient Algorithms</article-title>
          ,
          <source>The International FLAIRS Conference Proceedings</source>
          <volume>35</volume>
          (
          <year>2022</year>
          ). URL: https://journals.flvc.org/FLAIRS/ article/view/130584. doi:
          <volume>10</volume>
          .32473/flairs.v35i.
          <fpage>130584</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Barto</surname>
          </string-name>
          ,
          <article-title>Reinforcement learning: An introduction</article-title>
          , MIT press,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Towers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Terry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kwiatkowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. U.</given-names>
            <surname>Balis</surname>
          </string-name>
          , G. de Cola, T. Deleu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goulão</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kallinteris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. KG</given-names>
            ,
            <surname>M. Krimmel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Perez-Vicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pierré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schulhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Tai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J. S.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. G.</given-names>
            <surname>Younis</surname>
          </string-name>
          , Gymnasium, ???? URL: https://github.com/Farama-Foundation/Gymnasium.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Corporation</surname>
          </string-name>
          , ICS Matrix |
          <source>MITRE ATT&amp;CK®</source>
          ,
          <year>2023</year>
          . URL: https://attack.mitre.org/matrices/ ics/, publication Title:
          <article-title>ICS Matrix MITRE ATT</article-title>
          &amp;
          <article-title>CK® Type: Documentation.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>N.I.S.T.</surname>
          </string-name>
          ,
          <source>NIST SP 800-61 Rev. 2 - Computer Security Incident Handling Guide</source>
          ,
          <year>2012</year>
          . URL: https: //nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.
          <fpage>800</fpage>
          -
          <lpage>61r2</lpage>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Schulman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wolski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Klimov</surname>
          </string-name>
          ,
          <source>Proximal Policy Optimization Algorithms</source>
          ,
          <year>2017</year>
          . URL: http://arxiv.org/abs/1707.06347.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Velu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Vinitsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bayen</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Wu,</surname>
          </string-name>
          <article-title>The Surprising Efectiveness of PPO in Cooperative, Multi-</article-title>
          <string-name>
            <surname>Agent Games</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: http://arxiv.org/abs/2103.
          <year>01955</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Eßer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jestel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Urbann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kerner</surname>
          </string-name>
          ,
          <article-title>Guided Reinforcement Learning: A Review and Evaluation for Eficient and Efective Real-World Robotics [Survey]</article-title>
          ,
          <source>IEEE Robotics &amp; Automation Magazine</source>
          <volume>30</volume>
          (
          <year>2023</year>
          )
          <fpage>67</fpage>
          -
          <lpage>85</lpage>
          . URL: https://ieeexplore.ieee.org/document/9926159/. doi:
          <volume>10</volume>
          .1109/ MRA.
          <year>2022</year>
          .
          <volume>3207664</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Diprasetya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Pullani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schwung</surname>
          </string-name>
          ,
          <article-title>Sim-to-Real Transfer for Robotics Using ModelFree Curriculum Reinforcement Learning</article-title>
          ,
          <source>in: 2024 IEEE International Conference on Industrial Technology (ICIT)</source>
          , IEEE,
          <year>2024</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Louradour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Collobert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weston</surname>
          </string-name>
          ,
          <article-title>Curriculum learning</article-title>
          ,
          <source>in: Proceedings of the 26th Annual International Conference on Machine Learning</source>
          , ACM, Montreal Quebec Canada,
          <year>2009</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          . URL: https://dl.acm.org/doi/10.1145/1553374.1553380. doi:
          <volume>10</volume>
          .1145/1553374.1553380.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Maddison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sifre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Van Den Driessche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schrittwieser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Antonoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Panneershelvam</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Lanctot, others, Mastering the game of Go with deep neural networks and tree search</article-title>
          ,
          <source>nature</source>
          <volume>529</volume>
          (
          <year>2016</year>
          )
          <fpage>484</fpage>
          -
          <lpage>489</lpage>
          . Publisher: Nature Publishing Group.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kanervisto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Scheller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hautamäki</surname>
          </string-name>
          , Action Space Shaping in Deep Reinforcement Learning,
          <year>2020</year>
          . URL: http://arxiv.org/abs/
          <year>2004</year>
          .00980, arXiv:
          <year>2004</year>
          .00980 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Thakur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <article-title>An investigation on cyber security threats and security models</article-title>
          ,
          <source>in: 2015 IEEE 2nd international conference on cyber security and cloud computing, IEEE</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>307</fpage>
          -
          <lpage>311</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sabatelli</surname>
          </string-name>
          ,
          <article-title>Safe and psychologically pleasant trafic signal control with reinforcement learning using action masking</article-title>
          ,
          <source>in: 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC)</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>951</fpage>
          -
          <lpage>958</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Yasutomi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ogata</surname>
          </string-name>
          ,
          <article-title>Automatic Action Space Curriculum Learning with Dynamic Per-Step Masking</article-title>
          ,
          <source>in: 2023 IEEE 19th International Conference on Automation Science and Engineering (CASE)</source>
          , IEEE, Auckland, New Zealand,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          . URL: https://ieeexplore.ieee.org/document/ 10260397/. doi:
          <volume>10</volume>
          .1109/CASE56687.
          <year>2023</year>
          .
          <volume>10260397</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>R. RLLib</surname>
          </string-name>
          ,
          <article-title>Advanced python api's - curriculum learning</article-title>
          ,
          <year>2023</year>
          . URL: https://docs.ray.io/en/releases-2. 4.0/rllib/rllib-advanced
          <article-title>-api.html?highlight=curriculum%20leanrign#curriculum-learning, publication Title: Advanced Python API's - Curriculum Learning Type: Documentation.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>E.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Moritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nishihara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Stoica</surname>
          </string-name>
          ,
          <article-title>Rllib: Abstractions for distributed reinforcement learning</article-title>
          ,
          <year>2018</year>
          . URL: https://arxiv.org/abs/1712. 09381. arXiv:
          <volume>1712</volume>
          .
          <fpage>09381</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>