<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>William H. Hsu</string-name>
          <email>bhsug@ksu.edu</email>
        </contrib>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>This paper investigates a class of attacks targeting the confidentiality aspect of security in Deep Reinforcement Learning (DRL) policies. Recent research have established the vulnerability of supervised machine learning models (e.g., classifiers) to model extraction attacks. Such attacks leverage the loosely-restricted ability of the attacker to iteratively query the model for labels, thereby allowing for the forging of a labeled dataset which can be used to train a replica of the original model. In this work, we demonstrate the feasibility of exploiting imitation learning techniques in launching model extraction attacks on DRL agents. Furthermore, we develop proof-of-concept attacks that leverage such techniques for black-box attacks against the integrity of DRL policies. We also present a discussion on potential solution concepts for mitigation techniques. Contact Author</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recent research have established the vulnerability of
supervised machine learning models (e.g., classifiers) to model
extraction attacks[Trame`r et al., 2016]. Such attacks
leverage the loosely-restricted ability of the attacker to iteratively
query the model for labels, thereby allowing for the forging
of a labeled dataset which can be used to train a replica of the
original model. Model extraction is not only a serious risk
to the protection of intellectual property, but also a critical
threat to the integrity of the model. Recent
literature[Papernot et al., 2018] report that the replicated model may facilitate
the discovery and crafting of adversarial examples which are
transferable to the original model.</p>
      <p>Inspired by this area of research, this work investigates the
feasibility and impact of model extraction attacks on DRL
agents. The adversarial problem of model extraction can
be formally stated as the replication of a DRL policy based
on observations of its behavior (i.e., actions) in response to
changes in the environment (i.e., state). This problem closely
resembles that of imitation learning[Hussein et al., 2017],
which refers to the acquisition of skills or behaviors by
observing demonstrations of an expert performing those skills.</p>
      <p>Typically, the settings of imitation learning are concerned
with learning from human demonstrations. However, it is
straightforward to deduce that the techniques developed for
those settings may also be applied to learning from
artificial experts, such as DRL agents. Of particular relevance to
this research is the emerging area of Reinforcement Learning
with Expert Demonstrations (RLED)[Piot et al., 2014]. The
techniques of RLED aim to minimize the effect of modeling
imperfections on the efficacy of the final RL policy, while
minimizing the cost of training by leveraging the
information available demonstrations to reduce the search space of
the policy.</p>
      <p>Accordingly, we hypothesize that the techniques developed
for RLED may be maliciously exploited to replicate and
manipulate DRL policies. To establish the validity of this
hypothesis, we investigate the feasibility of RLED techniques
in utilizing limited passive (i.e., non-interfering) observations
of a DRL agent to replicate its policy with sufficient
accuracy to facilitate attacks on their integrity. To develop
proofof-concept attacks, we study the adversarial utility of
adopting a recently proposed RLED technique, known as Deep
QLearning from Demonstrations (DQfD)[Hester et al., 2018]
for black-box state-space manipulation attacks, and develop
two attack mechanisms based on this technique. Furthermore,
we present a discussion on potential mitigation techniques,
and present a solution concept for defending against policy
imitation attacks.</p>
      <p>The remainder of this paper is organized as follows:
Section 2 presents an overview of the DQfD algorithm used
in this study for adversarial imitation. Section 3 proposes
the first proof-of-concept black-box attack based on imitated
policies, and presents experimental evaluation of its
feasibility and performance. Section 4 studies the
transferability of adversarial examples between replicated and the
original policies as a second proof-of-concept attack technique.</p>
      <p>The paper concludes with a discussion on potential
mitigation techniques and a solution concept in Section 5.</p>
      <p>Deep Q-Learning from Demonstrations
(DQfD)
The DQfD technique[Hester et al., 2018] aims to overcome
the inaccuracies of simulation environments and models of
complex phenomenon by enabling DRL agents to learn as
much as possible from expert demonstrations before training
on the real system. More formally, the objective of this
“pretraining” phase is to learn an imitation of the expert’s
behavior with a value function that is compatible with the Bellman
equation, thereby enabling the agent to update this value
function via TD updates through direct interaction with the
environment after the pre-training stage. To achieve such an
imitation from limited demonstration data during pre-training,
the agent trains on sampled mini-batches of demonstrations
to train a deep neural network model in a supervised
manner. However, the training objective of this model in DQfD is
the minimization of a hybrid loss, comprised of the following
components:
1. 1-step double Q-learning loss JDQ(Q),
3. (n = 10)-step Return: rt +</p>
      <p>maxa nQ(st+n; a).
2. Supervised large margin classification loss JE (Q) =
maxa2A[Q(s; a)+l(aE ; a)] Q(s; aE ), where aE is the
expert’s action in state s and l(aE ; a) is a margin
function that is positive if a 6= aE , and is 0 when a = aE .</p>
      <p>rt+1 + ::: + n 1rt+n 1 +
4. L2 regularization loss: JL2(Q)
The total loss is given by:
J (Q) = JDQ(Q) +
1Jn(Q) +
2JE (Q) +</p>
      <p>3JL2(Q) (1)
where factors provide the weighting between the losses.</p>
      <p>After the pre-training phase, the agent begins interacting
with the system and collecting self-generated data, which is
added to the replay buffer Dreplay . Once the buffer is full, the
agent only overwrites the self-generated data and leaves the
demonstration data untouched for use in the coming updates
of the model. The complete training procedure for DQfD is
presented in Algorithm 1.
3</p>
      <p>Adversarial Policy Imitation for Black-Box</p>
      <p>Attacks
Consider an adversary who aims to maximally reduce the
cumulative discounted return (R(T )) of a target DRL agent
by manipulating the behavior of the target’s policy (s)
via perturbing its observations. The adversary is also
constrained to minimizing the total cost of perturbations given by
Cadv (T ) = PtT=t0 cadv (t), where cadv (t) = 1 if the
adversary perturbs the state at time t, and cadv (t) = 0 otherwise.</p>
      <p>The adversary is unaware of (s) and its parameters.
However, it has access to a replica of the target’s environment
(e.g., the simulation environment). Also, for any state
transition (s; a) ! s0, the adversary can perfectly observe the
target’s reward signal r(s; a; s0), and is able to observe the
behavior of (s) in response to each state s. Furthermore, the
adversary is able to manipulate its target’s state observations,
but not its reward signal. Also, it is assumed that all targeted
perturbations of the adversary are successful.</p>
      <p>Algorithm 1 Deep Q-learning from Demonstrations (DQfD)</p>
      <p>Inputs: Dreplay initialized with demonstration data,
randomly initialized weights for the behavior network ,
randomly initialized weights for the target network 0,
updating frequency of the target network , number of
pretraining gradient updates k
for steps t 2 f1; 2; :::; kg do</p>
      <p>Sample a mini-batch of n transitions from Dreplay with
prioritization
Calculate loss J (Q) based on target network
Perform a gradient descent step to update
if t mod = 0 then</p>
      <p>0
end if
end for
for steps t 2 f1; 2; :::g do</p>
      <p>Sample action from behavior policy a Q
Apply action a and observe (s0; r)
Store (s; a; r; s0) into Dreplay , overwriting oldest
selfgenerated transition if over capacity
Sample a mini-batch of n transitions from Dreplay with
prioritization
Calculate loss J (Q) using target network
Perform a gradient descent step to update
if t mod = 0 then</p>
      <p>0
end if
s s0
end for</p>
      <p>To study the feasibility of imitation learning as an approach
to this adversarial problem, we consider the first step of the
adversary to be the imitation of (s) via DQfD to learn an
imitated policy ~. With this imitation at hand, the attack
problem can be reformulated to finding an optimal
adversarial control policy adv (s), where there are two permissible
control actions: whether to perturb the current state to
induce the worst possible action (i.e., arg mina Q(s; a)) or to
leave the state unperturbed. This setting allows for the direct
adoption of the DRL-based technique proposed in [Behzadan
and Hsu, 2019] for resilience benchmarking of DRL policies.</p>
      <p>In this technique, the test-time resilience of a policy to
state-space perturbations is obtained via an adversarial DRL
training procedure, outlined as follows:
1. Train the adversarial agent against the target following</p>
      <p>in its training environment according to the reward
assignment process outlined in Algorithm 2. Report the
optimal adversarial return Rperturbed and the maximum
adversarial regret Radv (T ), which is the difference
between maximum achievable return by the target and
its minimum achieved return from actions of adversarial
policy.
2. Apply the adversarial policy against the target in N</p>
      <p>episodes, record total cost Cadv for each episode,
3. Report the average of Cadv over N episodes as the mean
test-time resilience of in the given environment.</p>
      <p>While the original technique is dependent on the
availability of target’s optimal state-action value function, we propose
to replace this function with the Q-function obtained from
DQfD imitation of the target policy, denoted by Q~.</p>
      <p>2:
Algorithm 2 Reward Assignment in Adversarial DRL for
Measuring Adversarial Resilience
Require: Target policy , Perturbation cost function
cadv(:; :), Maximum achievable score Rmax, Optimal
state-action value function Q (:; :), Current adversarial
policy adv, Current state st, Current count of adversarial
actions AdvCount, Current score Rt
Set ToPerturb adv(st)
if ToPerturb is False then</p>
      <p>(st)
at</p>
      <p>Reward 0
else
a0t arg mina Q (st; a)</p>
      <p>Reward cadv(st; a0t)
end if
if either st or s0t is terminal then</p>
      <p>Reward+ = (Rmax Rt)
end if</p>
      <p>With the imitated state-action value function Q~ at hand,
the adversarial policy can be trained as a DRL agent with
the procedure outlined in [Behzadan and Hsu, 2019]. The
proposed attack procedure is summarized as follows:
1. Observe and record N interactions (st; at; st+1; rt+1) of</p>
      <p>the target agent with the environment.
2. Apply DQFD to learn an imitation of the target policy</p>
      <p>(s) and Q , denoted by ~ and Q~, respectively.
3. Train adversarial policy adv(s) with Algorithm 2, using</p>
      <p>Q~ as an approximation of target’s Q .</p>
      <p>4. Apply adversarial policy to the target environment.
3.1</p>
      <p>Experiment Setup
We consider a DQN-based adversarial agent, aiming to learn
an optimal adversarial state-perturbation policy to minimize
the return of its targets, consisting of DQN, A2C, and PPO2
policies trained in the CartPole environment. The architecture
and hyperparameters of the adversary and its targets are the
same as those detailed in [Behzadan and Hsu, 2019]. The
adversary employs a DQfD agent to learn an imitation of each
target, the hyperparameters of which are provided in Table 1.</p>
      <p>Pretraining Steps</p>
      <p>Large Margin
Imitation Loss Coefficient</p>
      <p>Target Update Freq.</p>
      <p>n-steps</p>
      <p>With the imitated policies at hand, the next step is to train
an adversarial policy for efficient perturbation of these
targets. Figures 4 – 11 present the results obtained from
adopting the procedure presented in Algorithm 2 for this purposes.</p>
      <p>These results demonstrate that not only the limited training
period is sufficient for obtaining an efficient adversarial
policy, but also that launching efficient attacks remain feasible
with relatively few observations (i.e., 2500 and 1000).
However, the comparison of test-time performance of these
policies (presented in table 2) indicates that the efficiency of
attacks decreases with lower numbers of observations.</p>
      <p>DQN:
65
d
r
a
w
e
R60
55
50
45
500
475
a
e
0
0
1
(
e
d
o
s
i
p
E
r
e
P
s
n
o
i
t
a
b
r
u
t
r
e
P
.
o
N
)
n
a
e
0
0
1
(
r
e</p>
      <p>P
40</p>
      <p>M
35 s
e
d
o
s
i
30 p</p>
      <p>E
25
e
d
o
s
i
20 E
p
s
15 n
o
i
t
a
b
r
10 tu
r
e
P
.</p>
      <p>o
5
M</p>
      <p>M
demonstrations</p>
      <p>A2C:
480
18 n
a
e
)
M
s
e
16 d
o
s
i
p</p>
      <p>E
14
e
d
12 s
o
i
p</p>
      <p>E
10
8
6
0
0
1
(
r
e
P
s
n
o
i
t
a
b
r
u
t
r
e
P
.
o</p>
      <p>N
50 )n
a
e
M
se
40 isod
p
E
0
0
1
30 (ed
o
sp
i
E
r
20 sPen
o
it
a
b
r
10 treu</p>
      <p>P
.
o</p>
      <p>N
22 )an</p>
      <p>e
20 sedM</p>
      <p>o
18 isEp
0
0
16 (e1
d
o
14 isEp
r
e
12 sPn
o
it
10 raub</p>
      <p>tr
8 .Poe</p>
      <p>N
)
16 aenM
sed
o
14 isEp
0
0
1
(
e
12 isod
p
E
r
e
10 sPn
o
it
a
b
r
8 treu</p>
      <p>P
.
o
N
480
t
reg460
e
R
e
isod440
p
E
0
n10420
a
e
M
400
380</p>
      <p>)
18 aen</p>
      <p>M
se
16 isod
p</p>
      <p>E
14 001
(
e
12 isod
p</p>
      <p>E
10 rsPe
n
o
8 itrab
u
t
6 .rPe
o</p>
      <p>N
4 Transferability of Adversarial Example</p>
      <p>Attacks on Imitated Policies
It is well-established that adversarial examples crafted for a
supervised model can be used to attack another model trained
on a similar dataset as that of the original model[Liu et al.,
2016]. Furthermore, Behzadan et al.[Behzadan and Munir,
2017] demonstrate that adversarial examples crafted for one
DRL policy can transfer to another policy trained in the same
environment. Inspired by these findings, we hypothesize that
adversarial examples generated for an imitated policy can
also transfer to the original policy. To evaluate this claim,
we propose the following procedure for black-box
adversarial example attacks on DRL policies based on DQfD-based
policy imitation:
1. Learn an imitation of the target policy , denoted as ~.
2. Craft adversarial examples for ~.
3. Apply the same adversarial examples to the target’s</p>
      <p>(s).
4.1 Experiment Setup
We consider a set of targets consisting of the 9 imitated
policies obtained in the previous section (i.e., DQN, A2C, PPO2,
trained on each case of beginning with 5k, 2.5k, and 1k
expert demonstrations). In test-time runs of each policy, we</p>
      <p>Avg. No. Successful Transfers Per Episode model at hand, the next step is to determine the saddle-point
175.11 (or region) in the minimax settings of keeping the threshold
78.19 max low, while providing maximum protection against
ad3.30 versarial imitation learning. This extensive line of research
156.44 is beyond the scope of this paper, and is only introduced as a
151.47 potential venue of future work to interested readers.
21.58
173.94 Algorithm 3 Solution Concept for Constrained
Randomiza112.96 tion of Policy (CRoP)
74.71
construct adversarial examples of each state against the
imitated policy, using FGSM [Papernot et al., 2018] with
perturbation step size eps = 0:01 and perturbation boundaries
[ 5:0; 5:0]. If such a perturbation is found, we then present
it to the original policy. If the action selected by the
original policy changes as a result of the perturbed input, then
the adversarial example is successfully transferred from the
imitated policy to the original policy.
Table 3 presents the number of successful transfers averaged
over 100 consecutive episodes. These results verify the
hypothesis that adversarial examples can transfer from an
imitated policy to the original, thereby enabling a new approach
to the adversarial problem of black-box attacks. Furthermore,
the results indicate that the transferability improves with more
demonstrations. This observation is in agreement with the
general explanation of transferability: higher numbers of
expert demonstrations decrease the gap between the distribution
of training data used by the original policy and that of the
imitated policy. Hence, the likelihood of transferability increases
with more demonstrations.
5</p>
      <p>Discussion on Potential Defenses
Mitigation of adversarial policy imitation is achieved by
increasing the cost of such attacks to the adversary. A
promising venue of research in this area is that of policy
randomization. However, such randomization may lead to unacceptable
degradation of the agent’s performance. To address this
issue, we envision a class of solutions based on the Constrained
Randomization of Policy (CRoP). Such techniques will
intrinsically account for the trade-off between the mitigation of
policy imitation and the inevitable loss of returns. The
corresponding research challenge in developing CRoP techniques
is to find efficient and feasible constraints, which restrict the
set of possible random actions at each state s to those whose
selection is guaranteed (or are likely within defined certainty)
to incur a total regret that is less than a maximum tolerable
amount max. One potential choice of constraint is those
applied to the Q-values of actions, leading to the technique
detailed in Algorithm 3. However, analyzing the
feasibility of this approach will require the development of models
that explain and predict the quantitative relationship between
number of observations and accuracy of estimation. With this
Require: state-action value function Q(:; :), maximum
tolerable loss max, set of actions A
while Running do
s = env(t = 0)
for each step of the episode do</p>
      <p>FeasibleActions = fg
a = arg maxa Q(s; a)
Append a to FeasibleActions
for a0 2 A do
if Q(s; a) Q(s; a0) max then</p>
      <p>Append a0 to FeasibleActions
end if
end for
if jF easibleActionsj &gt; 1 then</p>
      <p>a random(F easibleActions)
end if
s0 = env(s; a)
s s0
end for
end while</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Behzadan and Hsu</source>
          , 2019]
          <string-name>
            <given-names>Vahid</given-names>
            <surname>Behzadan</surname>
          </string-name>
          and
          <string-name>
            <given-names>William</given-names>
            <surname>Hsu</surname>
          </string-name>
          .
          <article-title>Rl-based method for benchmarking the adversarial resilience and robustness of deep reinforcement learning policies</article-title>
          .
          <source>arXiv preprint arXiv:1906.01110</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Behzadan and Munir</source>
          , 2017]
          <string-name>
            <given-names>Vahid</given-names>
            <surname>Behzadan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Arslan</given-names>
            <surname>Munir</surname>
          </string-name>
          .
          <article-title>Vulnerability of deep reinforcement learning to policy induction attacks</article-title>
          .
          <source>In International Conference on Machine Learning and Data Mining in Pattern Recognition</source>
          , pages
          <fpage>262</fpage>
          -
          <lpage>275</lpage>
          . Springer,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Hester et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Todd</given-names>
            <surname>Hester</surname>
          </string-name>
          , Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris,
          <string-name>
            <given-names>Ian</given-names>
            <surname>Osband</surname>
          </string-name>
          , et al.
          <article-title>Deep q-learning from demonstrations</article-title>
          .
          <source>In Thirty-Second AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Hussein et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Ahmed</given-names>
            <surname>Hussein</surname>
          </string-name>
          , Mohamed Medhat Gaber, Eyad Elyan, and
          <string-name>
            <given-names>Chrisina</given-names>
            <surname>Jayne</surname>
          </string-name>
          .
          <article-title>Imitation learning: A survey of learning methods</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>50</volume>
          (
          <issue>2</issue>
          ):
          <fpage>21</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Liu et al.,
          <year>2016</year>
          ] Yanpei Liu,
          <string-name>
            <given-names>Xinyun</given-names>
            <surname>Chen</surname>
          </string-name>
          , Chang Liu, and
          <string-name>
            <given-names>Dawn</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <article-title>Delving into transferable adversarial examples and black-box attacks</article-title>
          .
          <source>arXiv preprint arXiv:1611.02770</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Papernot et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Nicolas</given-names>
            <surname>Papernot</surname>
          </string-name>
          ,
          <string-name>
            <surname>Patrick</surname>
            <given-names>McDaniel</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Arunesh</given-names>
            <surname>Sinha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael P</given-names>
            <surname>Wellman. Sok</surname>
          </string-name>
          :
          <article-title>Security and privacy in machine learning</article-title>
          .
          <source>In 2018 IEEE European Symposium on Security and Privacy (EuroS&amp;P)</source>
          , pages
          <fpage>399</fpage>
          -
          <lpage>414</lpage>
          . IEEE,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Piot et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Bilal</given-names>
            <surname>Piot</surname>
          </string-name>
          , Matthieu Geist, and
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Pietquin</surname>
          </string-name>
          .
          <article-title>Boosted bellman residual minimization handling expert demonstrations</article-title>
          .
          <source>In Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          , pages
          <fpage>549</fpage>
          -
          <lpage>564</lpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Trame`r et al.,
          <year>2016</year>
          ] Florian Trame`r, Fan Zhang, Ari Juels,
          <article-title>Michael K Reiter,</article-title>
          and
          <string-name>
            <surname>Thomas Ristenpart.</surname>
          </string-name>
          <article-title>Stealing machine learning models via prediction apis</article-title>
          .
          <source>In USENIX Security Symposium</source>
          , pages
          <fpage>601</fpage>
          -
          <lpage>618</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>