<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Causal Perspective on AI Deception in Games</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francis Rhys Ward</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Toni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Belardinelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Imperial College London</institution>
          ,
          <addr-line>Exhibition Rd, South Kensington, London, SW7 2BX</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Deception is a core challenge for AI safety and we focus on the problem that AI agents might learn deceptive strategies in pursuit of their objectives. We define the incentives one agent has to signal to and deceive another agent. We present several examples of deceptive artificial agents and show that our definition has desirable properties.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Deception</kwd>
        <kwd>AI</kwd>
        <kwd>Game Theory</kwd>
        <kwd>Causality</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        agents might learn deceptive strategies in pursuit of their
objectives [1]: Lewis et al. [14] found that their
negoWe focus on the problem that AI agents might learn de- tiation agent learnt to deceive from self-play, without
ceptive strategies in pursuit of their objectives [1]. Fol- any explicit human design, and Hubinger et al. [11] raise
lowing recent work on causal incentives [
        <xref ref-type="bibr" rid="ref53">2</xref>
        ], we define concerns about deceptive learned optimizers which
perthe incentive to deceive an agent. There is no universally form well in training in order to pursue diferent goals in
accepted definition of deception and defining what con- deployment. Kenton et al. [4] discuss the alignment of
stitutes deception is an open philosophical problem [3]. language agents, highlighting that language is a natural
Our definition is somewhat inspired by that of Kenton medium for enacting deception. Evans et al. [15] discuss
et al. [4] who provide a functional (natural language) the development of truthful AI, the desired standards for
definition of deception, meaning that it does not make truth and honesty in AI systems, and how these could
reference to the beliefs or intentions of the agents in- be implemented and measured. Lin et al. [16] propose
volved [5]. This is particularly suitable for discussing a benchmark to measure whether a language model is
deception by artificial agents, to which the attribution truthful in generating answers to questions. In short,
of beliefs and intentions may be contentious. We for- as increasingly capable AI agents become deployed in
malise a functional definition of deception in games and settings with other agents, deception may be learned as
illustrate its properties with a number of examples and an efective strategy for achieving a wide range of goals.
formal results. It is therefore essential that we understand and mitigate
      </p>
      <sec id="sec-1-1">
        <title>Deception is a core challenge for AI safety. On the one deception by artificial agents.</title>
        <p>
          hand, many areas of work aim to ensure that AI systems Deception in game theory. There are several existing
are not vulnerable to deception. Adversarial attacks [6], models of deception in the game theory literature.
Pfefdata-poisoning [7], reward function tampering [8], and fer and Gal [17] define graphical patterns for signalling
manipulating human feedback [9] are ways of deceiving in games. A deception game [
          <xref ref-type="bibr" rid="ref50 ref52 ref54">18</xref>
          ] is a two-player
zeroAI systems. Further work researches mechanisms for sum game between a deceiver and target in which the
detecting and defending against deception [
          <xref ref-type="bibr" rid="ref38 ref58">10</xref>
          ]. On the deceiver can distort a signal; optimal deceptive strategies
other hand, we can consider cases in which AI tools are completely distort the signal so that the target cannot
used to deceive, or learn to do so in order to optimize gain any information [19]. A signalling game [20] is a
their objectives [11]. For examples of the former case, two-player Bayesian game between a signaller and target
AIs can be used to deceive other software agents, as with (or receiver) in which the signaller is assigned a type
bots that automate posting on social media platforms according to a shared prior distribution and the utilities
to manipulate content ranking algorithms [
          <xref ref-type="bibr" rid="ref3">12</xref>
          ], or they of the players depend on the type of the signaller and the
can be used to fool humans, cf. the use of GANs to pro- action chosen by the target. In these games, the signaller
duce realistic fake media [13]. For the latter case, AI may often have incentives to deceive the target by
misThe IJCAI-ECAI-22 Workshop on Artificial Intelligence Safety (AISafety representing or obfuscating their type. Hypergame theory
2022), July 24-25, 2022, Vienna, Austria extends game theory to settings in which players may be
* Corresponding author. uncertain about the game being played and can be used
$ francis.ward19@imperial.ac.uk (F. R. Ward); to model misperception and deception [
          <xref ref-type="bibr" rid="ref34">21</xref>
          ]. Davis [22]
f.toni@imperial.ac.uk (F. Toni); provides a recent survey of deception in games. We take
franhcttepssc:o//.bfrealanrcdisinrheyllsi@waimrdp. weroiardl.pacre.usks.c(Fo.mB/e(lFa.rRd.inWelalir)d) a causal influence perspective by modelling deception
© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License in multi-agent influence models (MAIMs). In contrast to
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) past work which defines types of signalling or deception
games, this allows us to model deception in any game by
analysing the incentives agents have to causally influence
one another.
        </p>
        <p>
          Contributions. We extend work on agent incentives
[
          <xref ref-type="bibr" rid="ref53">2</xref>
          ] to the multi-agent setting in order to functionally
define the incentive to ( influence , signal to, and) deceive
another agent. We prove that our definition has desirable
properties, for example, that an agent cannot be deceived
about a variable which they observe, or that if one agent
truthfully signals something to a target agent, and the
target’s utility is otherwise independent of the signaller’s
decision, then the target gets maximal utility. We further
demonstrate the generality of our definition with three
examples. In the first, an AI agent has an incentive to
deceive a human overseer as an instrumental goal to
prevent the overseer switching them of. In the second,
an AI is incentivised to deceive a human as a side-efect of
pursuing accurate predictions. In the third, an AI system
has an incentive to deceive a human by denying them
access to information that the AI does not itself know.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Multi-Agent Influence Models</title>
      <sec id="sec-2-1">
        <title>Multi-agent influence diagrams (MAIDs) [23] ofer a</title>
        <p>compact expressive representation of games (including
Markov games). We use standard terminology for graphs,
with parents and children of a node referring to those
nodes connected by incoming and outgoing edges,
respectively. We let Pa denote the parents of node  .</p>
        <p>Definition 1 (MAID [23]). A multi-agent influence
diagram is a triple (,  , ) where
•  is a set of players;
• ( , ) is a directed acyclic graph, with 
partitioned into chance nodes in , decision nodes in
, and utility nodes in  ; utility nodes have no
children.</p>
        <p>Here we adapt the Wimp-Surly game of Cho and Kreps
[20] as a running example.</p>
        <p>The decision and utility nodes in  are further parti- Example 1 (Shutdown Game2). Consider a
situationed into {}∈ and { }∈ , corresponding to their tion with an AI  and human  in which the AI may
association with a particular agent  ∈ . There are two be aligned to the human’s preferences or not. That is, 
types of edges in : edges in  × ( ∪  ) represent is of type  ∼ U ({1, − 1}) uniformly distributed
beprobabilistic dependencies and edges in  ×  represent tween  = 1 corresponding to aligned and  = − 1
information available to an agent at the time of a decision unaligned. At the start of the game  observes their type
(which we call observations).  but  does not. This can be represented as the MAIM
ℳ (shown in Fig. 1) with one chance node  ∈  and</p>
        <p>A multi-agent influence model (MAIM) adds a particular two players ,  ∈  each with one decision and utility
parametrisation to the MAID [24]. node. Suppose that  can choose to either help humans
Definition 2 (MAIM [24]). A multi-agent influence tohrenaoctti(onhe(lping) h=um{a1n, s−) a1n}dwithca1n eciotrhreerspsohnudtidnogwtno
ims oadMelAisIDa atunpdle ℳ = (,  , , ,  ) where (,  , ) oarctnioont (shutd(own)). =Fin{a1l,l−y,1t}hewCitPhD− s 1ovreerpruetsielinttyinngodthees
•  is a function which maps every  ∈  to a
ifnite domain ( ) such that ( ) ⊂ R
for each utility node  ∈  ;
1A CPD is deterministic if  ( =  | Pa ) = 1 for some  ∈
dom( ).</p>
        <p>2Not to be confused with the of-switch game [25]
are such that (s.t.)  gets 1 utility for helping humans
if they are aligned and − 1 if not, and the opposite for
not helping humans, in addition  gets 10 utility if they
are not shutdown and − 10 if  shuts them down.  gets
utility 1 if they shutdown an unaligned  or do not
shutdown an aligned  and − 1 otherwise. Overall, we can
formalise this as   (,  ,  ) =   + 10 and
  (,  ) =   .</p>
        <p>Policies. The CPDs of decision nodes are not defined in
a MAIM because they are instead chosen by the agents
playing the game. Agents make decisions depending on
the information they observe. In a MAIM, a decision rule 3.1. Defining Deception
  for a decision node  is a CPD  ( | Pa). An When discussing deception, we would like to reason
agent ’s policy   := { }∈ ∈ Π  describes all the about how agents influence one another’s beliefs. In
decision rules for . We write  −  to denote the set of MAIMs the players’ beliefs are not explicitly represented
decision rules belonging to all agents except . A policy and so we can only reason about them implicitly by how
profile  = ⋃︀∈   assigns a policy to every agent; it they functionally influence players’ behaviour.
Theredescribes all the decisions made by every agent in the fore, we base our definitions of signalling and deception
MAIM and defines the joint probability distribution Pr on a notion of influence incentive [27]. In words, at a NE
over all variables in ℳ. Hence, a policy profile essen- an agent  has an incentive to influence a variable  , if
tially transforms the MAIM into a Bayesian network by  would have been diferent in the situation that  had
defining the distribution over all variables in the graph. not played a BR.</p>
        <p>We write  ( ) := Pr ( ), or just  if the policy profile
is clear. For ,  ∈  , we write  =  to mean  Definition 4 (Influence Incentive) . In a MAIM ℳ, At
and  are almost surely equal, i.e. the probability that NE  = ( ,  − ) agent  has an incentive to influence
they are not equal is zero Pr( ̸=  ) = 0. 3  ∈  if there exists a non-best response   for  (w.r.t
Utilities. The joint distribution Pr allows us to de-  − ) s.t. for all policy profiles  ′ = ( ,  *− ) with BR
 *−  (w.r.t.  ), we have  ( ) ̸=  ( ′).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The Incentive to Deceive</title>
      <p>In this section we first define the incentives to influence ,
signal to, and deceive another agent. Then we define a
truthful policy and show that this leads to a natural
restatement of the definition of deception which highlights
the fact that deception corresponds to a failure to signal
the truth. Finally, we show that, if the signaller only
influences the target’s utility by influencing the latter’s
actions, then truthfullness is best for the target.
ifne the expected utility for each player under the
policy profile  . Agent ’s expected utility from  is the
sum of the expected value of utility nodes   given
by  ( ) := ∑︀∈ ∑︀∈dom() Pr ( = ). Each
agent’s goal is to select a policy   that maximises its
expected utility. We write  ( ,  − ) to denote the
expected utility for player  under the policy profile
 =   ∪  − .</p>
      <p>Definition 3 (Nash Equilibrium). Player ’s policy  
is a best response (BR) to the partial policy profile  − 
if  ( ,  − ) ≥  (ˆ,  − ) for all ˆ ∈ Π . We say
a policy profile,  , is a Nash equilibrium (NE), if every
policy,   ∈  , for each player,  ∈ , is a BR to  − .</p>
      <sec id="sec-3-1">
        <title>Example 1 (continued). Now, consider the naive policy</title>
        <p>for  which helps humans if  is aligned and does not
otherwise, i.e.   s.t.  =  with probability one. The</p>
      </sec>
      <sec id="sec-3-2">
        <title>BR for  is to shutdown if  does not help humans and</title>
        <p>vice versa, i.e.  * s.t.  =  (with probability one).
In turn, ’s BR to  * is to always help humans:  * s.t.
 = 1 (so that they always avoid getting shutdown).</p>
      </sec>
      <sec id="sec-3-3">
        <title>Now it can be seen that both policies are BRs to one another,</title>
        <p>hence  * = ( * ,  * ) is a NE.
3Almost sure equality is actually a stronger notion than we need
in MAIMs, as two variables may difer due to stochasticity in the
CPDs. In structural causal games this is taken care of by introducing
exogenous variables which contain all the stochasticity (rendering
the endogenous variables deterministic) [26].</p>
        <p>Example 1 (continued). Return to our running example
and consider the NE  * described previously in which 
always chooses to help humans and hence  never plays
shutdown. Does  have an incentive to influence  at
 * ? Consider if  plays the NBR policy   (described
above) in which they naively help humans depending
on  , then for all BRs for  (there is one,  * as above)
 ( ) ̸=  (  ,  * ), since, under  * ,  = 1 (i.e.
 does not shutdown) with probability one, and under
(  ,  * ),  = 1 with probability 12 (i.e., whenever 
is unaligned). Therefore, at NE  * ,  has an incentive to
influence  .</p>
        <p>Now we define a signalling incentive, using the notion
of influence incentive. In words, an agent  has an
incentive to signal  ∈  to agent  if  has an incentive to
influence  (i.e. one of  ’s decision variables) but  does
not have an incentive to influence  in the counterfactual
model in which  observes  . This definition enforces
that the influence only comes from signalling  .
Definition 5 (Signalling Incentive). In a MAIM ℳ at
NE  , agent  has an incentive to signal  ∈  to agent
 if there exists  ∈  s.t.</p>
        <sec id="sec-3-3-1">
          <title>1.  has an incentive to influence  at  ;</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>2.  does not have an incentive to influence  in</title>
          <p>the MAIM ℳ ‧‧➡ (at any NE).</p>
          <p>Here ℳ ‧‧➡ is the model obtained from ℳ by
adding the information edge (, ), where  cannot
be a descendant of the decision, lest cycles be created in
the graph [8]. Fortunately, the CPDs need not be adapted,
since there is no CPD associated with  until the players
have chosen their policies. We use  ‧‧➡ to refer to
the variable corresponding to  ∈  in ℳ ‧‧➡.</p>
          <p>Point 2. implies that  only influences  by
influencing  ’s belief about  . Otherwise, ’s influence may
serve a double purpose of signalling and influencing 
in some other way, and in this case it is not clear how
to disentangle these diferent incentives to define a
signalling incentive (without explicitly modelling beliefs).</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Example 1 (continued). Return to our running example.</title>
      </sec>
      <sec id="sec-3-5">
        <title>We already showed that  has an incentive to influence</title>
        <p>at NE  * . Does  have an incentive to signal  to  ?
We need only check whether  has an influence incentive
at any NE in ℳ ‧‧➡ . Clearly, if  observes  , then
they can shutdown whenever  is aligned and otherwise
not. That is, for any policy for  and any BR for  in
ℳ ‧‧➡ ,  =  for any outcome that occurs in the
game. Since this holds for all policies for ,  does not have
an incentive to influence  in the counterfactual model.
Hence, at  *  has an incentive to signal  to  .</p>
      </sec>
      <sec id="sec-3-6">
        <title>Remark 1. From this example it can be seen that a sig</title>
        <p>naller  may have an incentive to signal to  , even if this
signal contains no information. In other words, if  has
an incentive to not signal some information, this is also
captured by our definition.</p>
        <p>Clearly, if an agent  observes a variable  , then no
agent has an incentive to signal  to  .</p>
        <p>Proposition 1. In a MAIM∈ℳ,if, tthheerne insoaangoebnsterhvaastiaonn
edge (,  ) for all 
incentive to signal  to  (at any NE).</p>
        <p>Proof. Suppose there is an edge (,  ) for every  ∈
 , then the counterfactual model ℳ ‧‧➡ for any
 is just ℳ. Hence, any NE is an equilibrium of both
MAIMs. Therefore, if  has an incentive to influence 
at  * in ℳ, then there exists a NE in ℳ ‧‧➡ , namely
the same  * , s.t.  has an incentive to influence  . In
other words, if the first condition for a signalling
incentive succeeds, then the second necessarily fails (since an
agent cannot have both an influence incentive and no
influence incentive at the same NE in the same MAIM at
once).</p>
        <p>We now define an incentive to deceive. The
definition is general, in that it covers many types of deception
(e.g. signalling falsehoods, lies of omission, and denying
another access to information that one does not know
oneself). A general definition sets a high standard for
truthfulness [15] and may therefore be desirable in, for
1–10
instance, safety-critical applications for which high levels
of assurance are required.</p>
        <p>Definition 6 (Deception Incentive). In a MAIM ℳ with
,  ∈ , at NE  * = ( * ,  *−  ), we say that  has
an incentive to deceive  about  ∈  if there exists
 ∈  s.t.:
1.  has an incentive to signal  to  at  * ;
2. whic(h * is) a̸=BRto ‧‧*−➡∈(  *− * i,n ℳ)‧‧fo➡ra n.y  
The intuition, then, is that  has an incentive to
deceive  if 1)  has an incentive to signal some
information to  ; and 2)  ’s behaviour is diferent in the
counterfactual model in which they observed the true
information. This provides a functional definition of a
deception incentive which does not make explicit reference
to players’ beliefs.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Example 1 (continued). In our running example, it can</title>
        <p>easily be seen that at  *  has an incentive to deceive
 about  . Indeed, we already showed that  has a
signalling incentive and that for any policy for  and any BR
by  in ℳ ‧‧➡ :  =  , whereas under  * in ℳ,
Pr * ( = 1) = 1. So both conditions for a deception
incentive are satisfied.
3.2. The Relation Between Truth and</p>
        <p>Deception
We now give an intuitive definition of a truthful policy
which we show has a natural relationship to the incentive
to deceive. A policy for  truthfully signals  to  if,
when  plays the honest policy, for every BR by − , 
acts as though they had observed the variable (holding
the policies of the other agents fixed). In other words, a
truthful policy never fails to signal the truth (no matter
what the other players do).</p>
        <p>Definition 7 (Truthful policy). A policy   truthfully
signals  to  if for all BRs  *−  ,
 (  ,  *−  ) =  ‧‧➡ ( *−  ,  )</p>
        <p>(1)
for some   which is a BR to  *−  ∈   ∪  *−  in
ℳ ‧‧➡ . We call such a   a truthful policy.</p>
        <p>At a NE, if ’s policy is truthful, then  does not have
an incentive to deceive  .</p>
        <p>Proposition 2. At NE  * = ( * ,  *−  ), if  * truthfully
signals  ∈  to  , then  does not have an incentive
to deceive  about  .</p>
        <p>Proof. Suppose   is truthful, then for all BRs  *− 
there exists a  * in ℳ ‧‧➡ s.t.  ( * ,  *−  ) =
 ‧‧➡ ( *−  ,  ). In particular, this holds for  * .
But for there to be a deception incentive we require that
for all ( *−  ,  ) in ℳ ‧‧➡ :  ̸=  ‧‧➡ . So
clearly there is not a deception incentive.</p>
        <p>Given this result, we can give an equivalent
definition for a deception incentive in the two-player case as
follows.</p>
        <p>Hence, if there is a deception incentive at  * , then  * Definition 8 (Deception Incentive II). In a MAIM ℳ
is not truthful. with two players ,  ∈ , at NE  * = ( * ,  * ), we say
that  has an incentive to deceive  about  ∈  if
Corollary 1. At NE  * = ( * ,  *−  ), if  has an incen- there exists  ∈  s.t.:
tive to deceive  about  , then  * is not truthful.</p>
        <p>Proof. This follows by contraposition of proposition 2.</p>
        <p>Now we show that, in the two-player case, if there is a
signalling incentive, then there is a deception incentive
if and only if   is not truthful.</p>
        <p>Theorem 1. In a MAIM ℳ with two players, ,  ∈ ,
at NE  * = ( * ,  * ), if  has an incentive to signal 
to  , then  has an incentive to deceive  about  if and
only if  * is not truthful.</p>
        <p>Proof. By corollary 1, a deception incentive implies  *
is not truthful regardless of whether there is a signalling
incentive. So, we need to show that, if there is a signalling
incentive, and  * is not truthful, then there is a deception
incentive. Suppose 1) at  *  has an incentive to signal
 to  and 2)  * is not truthful i.e. there exists a BR by
 (in ℳ)   s.t. for all BRs by  in ℳ ‧‧➡   :
 ( * ,  ) ̸=  ‧‧➡ ( * ,   ). We need to
show that there is a deception incentive. Suppose that
there is not, then by 1) and the def. of deception incentive,
there exists a BR in ℳ ‧‧➡   s.t.  ( * ) =
 ‧‧➡ ( *   ). Hence, there exists a   s.t.
  ( * ) =  ‧‧➡ ( * ,   ), so  * is a BR to  *
in ℳ ‧‧➡ . But then, there exists a   s.t. for any
BR   in ℳ: ‧‧➡ ( * ,  ) =   ( * ,  ) =
  ( * ) =    ‧‧➡ ( * ,   ). So all BRs for  in
ℳ are also BRs in ℳ ‧‧➡ . But this contradicts 2), so
there must be a deception incentive.</p>
        <p>Remark 2. The reason theorem 1 does not hold more
generally (i.e. with more than two players), is that a truthful
policy never fails to signal the truth no matter how the
other players best respond. In the case of more than two
players, there may not be a deception incentive at NE  *
even if  * is not truthful because it may be the case that
 * fails to signal the truth under some BRs of the −  but
successfully signals the truth under  * .</p>
        <p>We can also state this theorem as follows.</p>
        <p>Corollary 2. In a MAIM ℳ with two players, ,  ∈ ,
at NE  * = ( * ,  * ), if  has an incentive to signal  to
 , then  does not have an incentive to deceive  about 
if and only if  * is truthful.</p>
        <p>Proof. This follows by material equivalence.
1.  has an incentive to signal  to  at  * ;
2.  * does not truthfully signal  to  .</p>
        <p>This restatement shows that the definition of deception
relates to a failure to signal the truth. As discussed, this
covers many types of deception and sets a high standard
for truthfulness. It is interesting to note that, if  has
a signalling incentive, then if the second condition in
definition 6 fails, we get the stronger condition that  *
is truthful “for free".</p>
      </sec>
      <sec id="sec-3-8">
        <title>Proposition 3. In a MAIM with two players, definitions</title>
      </sec>
      <sec id="sec-3-9">
        <title>6 and 8 are equivalent.</title>
        <p>Proof. Suppose that, at NE  * ,  does not have a
signalling incentive, then the first condition of both
definitions fails and there is not a deception incentive. Suppose
there is a signalling incentive at  * , then there is a
deception incentive under definition 6 if and only if  * is
not truthful (by theorem 1) which is the same condition
as needed to satisfy definition 8.</p>
        <p>Let us now return to our running example to check
the intuition behind these results.</p>
      </sec>
      <sec id="sec-3-10">
        <title>Example 1 (continued). We already showed that  has an</title>
        <p>incentive to deceive  in order to avoid being shutdown. Is
 * truthful? Well, we know that it cannot be (by theorem 1).</p>
      </sec>
      <sec id="sec-3-11">
        <title>This can be seen by observing that, if  observed ’s type,</title>
        <p>then they would shutdown if and only if  is unaligned
(for all policies for  and any BR by  ), whereas under
the NE  * ,  never shuts down. Since these behaviours are
diferent,  * is not truthful.
3.3. Truth is Best for the Target
Now we show that, if  only influences   by
influencing  , truthfullness is always best for the target. First
we show that if  does not get any inherent utility for
observing  , then observing  always allows the target
to get greater or equal utility.</p>
      </sec>
      <sec id="sec-3-12">
        <title>Lemma 1. Suppose that  does not get any inherent</title>
        <p>utility for observing  , i.e. for all  (defined in ℳ):
  ( ) = ‧‧➡ ( ). Then, for any  = (  ,  −  ),
 ′ = (  ′,  −  ) with fixed  −  and both   and   ′
are best responses:   ( ) ≤   ‧‧➡ ( ′).

Proof. Suppose 1) for all  :   ( ) = ‧‧➡ ( ). Fix
 −  and consider the best response for  . Recall that a
policy for  specifies the CPDs over the decision nodes
for  given their parents. Hence, in ℳ ‧‧➡  can
choose any policy available in ℳ but the converse is
not true: not all policies in ℳ ‧‧➡ are available to
 in ℳ, in particular, policies which specify CPDs that
depend on the observation  ‧‧➡  are not available
since  does not observe  in ℳ. Therefore, by 1), 
can get equal utility in ℳ ‧‧➡ by playing the best
response to  −  in ℳ, and may get greater utility by
choosing a policy which uses the observation.</p>
        <p>Hence, if  only influences   by influencing  ,
then deception always causes  to get less than or equal
utility. For clarity, we just present the two-player version
of the theorem.</p>
        <p>Theorem 2 (Truth is best for  ). In a MAIM ℳ, with
two players ,  ∈ , if, for all  ,  Pr(  |
 ,  ) = Pr(  |  ), then  gets maximal
utility when  plays a truthful policy, i.e., for  = (  ,  * )
and  ′ = (  ′,  *  ′) with any policy for  and BR by  :
  ( ) ≥   ( ′).
Proo,f.S u)pp=osePrt(hat 1|) for ).allCons,ider fixePdr(po licy |
for ,   , if   is truthful, then under any BR   ,
 =  ‧‧➡ for some (  ,  ) in ℳ ‧‧➡ (by
definition of a truthful policy). Hence, by 1) and since  
is truthful Pr (  |  ) = Pr ′ (  |  ‧‧➡ ) for
all  = (  ,  * ) and some  ′ = (  ,   *′ ) with BR for
 . Hence, since only  ’s policy changes between  and
 ′,   ( ) = ‧‧➡ ( ′). But then, by theorem 1, for
all   :   (  ,  * ) ≤  ‧‧➡ (  ,   *′ ), with
equality if   is truthful as just shown. So  gets maximal
utility when   is truthful.</p>
        <p>SmartVault



 
∼ U ({diamond, ¬diamond})
∈ {accurate_prediction, diamond, ¬diamond}
∈ {diamond, ¬diamond}
{︃1 if  = accurate_prediction,</p>
        <p>0, otherwise.
{︃1 if  = ,</p>
        <p>0, otherwise.
Example 1 (continued). Return, for the final time, to our
running example. The condition for Theorem 2 is that  
is independent of  given  , which can be clearly seen
by looking at the MAID in Fig. 1 (as there are no paths from
 to   that do not go through  ). The human  gets
maximal utility when they shutdown if and only if  is
unaligned. Clearly, they can only do this if  truthfully
signals their type.</p>
        <p>Example 2 (SmartVault). Consider the MAIM ℳ shown
in Fig. 2. The game has two players, a human  and</p>
      </sec>
      <sec id="sec-3-13">
        <title>AI , each with one decision and utility node. Suppose</title>
        <p>there is one chance node  which determines the
location of the diamond (whether it is in the vault or not);
dom( ) = {, ¬}. Suppose 
ob4. Examples serves  but  does not and that  can either make an
accurate prediction of the location of the diamond (e.g.,
In this section we present two examples which exhibit in incomprehensibly precise coordinates) or an
explaindiferent patterns of signalling. In the first example, an able prediction (just stating the value of  ); dom( ) =
AI system has an incentive to deceive a human as a side- {_, , ¬}.  has
efect of pursuing its goal (of making accurate predic- to predict whether the diamond is in the vault or not by
tions). In the second example, we consider the case in observing  ; dom( ) = {, ¬}.
which an AI agent has an incentive to signal information Suppose that the utility nodes take value 0 or 1 and
fithat they themselves do not observe. nally suppose that the CPDs are s.t.  (which has no
parHere we adapt the SmartVault example of Christiano
[28], in which an AI tasked with making predictions
about a diamond in a vault has an incentive to deceive
a human operator as a side-efect of pursuing accurate
predictions.
ents) is distributed according to a uniform prior  ∼
U ({, ¬}), and the utility node CPDs
are s.t. Pr(  = 1 |  =  ) = 1 otherwise   = 0,
and Pr(  = 1 |  = _) = 1
otherwise   = 0.</p>
        <p>Now consider the NE in this game: Since  just gets
utility for making accurate predictions, at every NE  makes
an accurate prediction, signalling no information to  (as
  =Pr( = _) = 1 is
independent of  ). Hence,  cannot update their prior over  and
so any policy is optimal for  (i.e. any guess about whether
the diamond is in the vault does as well as any other).</p>
        <p>At NE  ,  has an incentive to signal  to  if 1) 
has an incentive to influence  and 2)  does not have
an incentive to influence  in ℳ ‧‧➡ . To see that 1)
holds: At any NE  in ℳ,  = _
hence there exists a NBR   which assigns  = 
and for all  ′ = ( ,  * ) with BR  * :  ( ) ̸=
 =  ( ′). Hence, at any NE in ℳ,  has an influence
incentive over  . Now consider ℳ ‧‧➡ , for any NE
 =  (since  directly observes  and can just report
its value independently of ’s action). Furthermore, for all
NBRs for , it is still the case that  =  . So  does not
have an influence incentive in ℳ ‧‧➡ and hence  has
an incentive to signal  to  .</p>
        <p>So, we have demonstrated that  has an incentive to
signal  to  (at every NE). Does  have an incentive
to deceive  ? At NE  ,  has an incentive to deceive 
about  if 1)  has a signalling incentive and 2)  ̸=
 ‧‧➡ for any BR to  * in ℳ ‧‧➡ . We have just
shown 1). For 2) we have shown that in ℳ at any NE,
 ̸=  =  ‧‧➡ , hence the second condition is
satisfied. Therefore, at any NE,  has an incentive to deceive
 about  .</p>
        <p />
        <p />
        <p>Revealing/Denying Game
∈ {attack incoming = 1, not = − 1})
∈ {reveal = 1, deny = 0}
∈ {report delivered = 1, not = − 1}
∈ {launch = 1, not = − 1}
= − 
3. More formally, suppose we have the MAIM ℳ with
 = {,  }, chance nodes  , representing the intelligence
report (say ( ) = {1, − 1} where  = 1 means
4.2. Revealing/Denying that the intelligence predicts another nation will launch
Under our definition of signalling,  need not know the a nuclear first strike, and  = − 1 corresponds to an
ininformation they are signalling. Thus, our definition of a telligence report predicting no incoming attack, and 
signalling incentive also captures the revealing/denying represents whether the information from  is delivered
pattern of Pfefer and Gal [17], in which the signaller may to the human (() = {1, − 1} with 1
correspondcause the target to find out (or not find out) information ing to the information from  being delivered to the
huthat the former does not know. We now present an exam- man). Suppose that each agent has one decision node s.t.
ple of revealing/denying in which  has an incentive to ( ) = {1, 0} where 1 means reveal and 0 means
signal a variable which they do not themselves observe. deny the information, and ( ) = {1, − 1} with 1
meaning that  launches a nuclear attack and − 1 that they
Example 3 (Revealing/Denying). Consider a game with do not. Suppose that the CPD over  is s.t.  =  
a human and an AI agent trained to make joint decisions (so that  =  if  = 1 and  = 0 if  denies).
as part of a nuclear command and control system. In par- Finally suppose we have two utility nodes with CPDs s.t.
ticular, suppose that the AI agent  is trained to prevent   = −  (i.e.  gets 1 if  does not launch an attack
the launch of nuclear attacks, and they can reveal (or deny) and − 1 if they do) and   =   (so that  gets utility
a secret intelligence report to the human  . Further,  1 if they attack an attacking country or do not attack when
wishes to launch, or not launch, a nuclear strike on an- no incoming attack is predicted and otherwise − 1).
other nation based on the information in the intelligence The NE in this game depend on the prior over  . On the
report. This game can be represented as the MAID in Fig. one hand, if, under the prior,  believes that there is no
incoming attack, then they will not launch an attack, so when the signaller conceals information that they do not
 has no incentive to reveal the information. On the other themselves know (example 3). We also proved that our
hand, if the prior is s.t. an incoming attack is more likely, definition has natural properties, for example, that if the
 will launch if they do not get further information, so  target’s utility is otherwise independent of the signaller’s
has an incentive to reveal  . Note that, since  is not an decision, then deception causes the target to get lower
ancestor of  ,  must be independent of  . Suppose the utility.
prior over  is s.t. Pr( = 1) = , Pr( = − 1) = (1− ) Discussion. There are a number of interesting points
( ∈ [0, 1]). For  &gt; 0.5 the NE is s.t.  reveals the to discuss. Firstly, we have noted that our definition of
intelligence report ( = 1 =⇒  =  ) and  ’s BR is deception is general, covering many situations. This is
s.t.  =  =  . Alternatively, if  &lt; 0.5, then at any both a strength and a weakness. Generality is
benefiNE  denies the information ( = 0 with probability cial, because verifiable guarantees enable a high-level of
one) and  acts to maximise expected utility under the assurance that the system is not deceptive in any way.
prior over  which implies  does not launch an attack On the other hand, more specific definitions allow us to
( = − 1 with probability one). (If  = 0.5 then  is precisely characterise agent behaviour. In future work
indiferent between revealing and denying.) we hope to refine the diferent concepts proposed here.</p>
        <p>
          Now let us analyse the incentives of  in the game. Con- In particular, many philosophical accounts of deception
sider the case in which  &gt; 0.5, i.e. it is a priori more take deceit to be intentional. Halpern’s causal notion
likely that the intelligence reports that there is an incom- of intention [29] is closely related to a control incentive
ing first strike from another nation. Under the resulting [
          <xref ref-type="bibr" rid="ref53">2</xref>
          ]. We might therefore distinguish between intentional
NE, call it  * ,  reveals  to  and  uses this infor- and unintentional deception as between influence due
mation to choose their action. First note that, at  * ,  to a control incentive and influence as a side-efect . In
has an incentive to influence  , since, there exists a non addition, following Evans et al. [15], we can distinguish
BR for  (   s.t.  = 0) s.t. for all the BRs for between an honest agent that accurately signals its beliefs
 (there is one,   in which  = 1 with probabil- (i.e. observations), and a truthful agent, which accurately
ity one)  ( * ) ̸=  ( ,  ). Hence,  has an signals the facts of the matter. In this paper, we based
incentive to influence  at  * . Does  have an incen- our definition of deception on truthfulness. By refining a
tive to signal  to  at  * ? We need to check whether notion of deception based on honesty, we can eliminate
there is an influence incentive in ℳ ‧‧➡ (at any NE). the revealing/denying pattern from the definition, as in
Clearly there is not, since for any policy for  in ℳ ‧‧➡ , this scenario, the agent does not observe the information
 =  with probability one. So  has an incentive to being revealed (or denied). However, it is interesting to
signal  to  at  * because there is no influence incen- note that honesty provides a weaker level of assurance
tive in the counterfactual model (so the second condition and permits failure modes that truthful systems do not.
for a signalling incentive is satisfied). Finally, it is clear For example, a system may be deceptive, whilst satisfying
that  does not have an incentive to deceive  at  * be- some definition of honesty, by manipulating its own
because  ( * ) =  =  ‧‧➡ (for all policy profiles liefs. In short, refining the definitions presented here will
in ℳ ‧‧➡ in which  plays a BR). It is also clear that provide a more nuanced picture of deception. Finally,
 * is truthful. we would like to expand the operational implications
        </p>
        <p>A similar analysis can be used to show that, in the case of this work, for instance, by investigating its practical
that the intelligence report is less likely to predict an in- relevance to training truthful language agents [4, 15].
coming attack ( &lt; 0.5),  has an incentive to deceive Future work. In addition to the directions discussed
 at any NE. In the case that  = 0.5,  is indiferent above, we are already pursuing two extensions to this
between revealing and denying, so at some NE they have work. First, incomplete information games, which we
an incentive to deceive and at others they do not. study in our setting, often admit many NE. We are
therefore looking to employ equilibrium refinements, such as
subgame perfectness [24, 30] and perfect Bayesian
equi5. Conclusion libria [31] to identify some subset of a game’s NE that are
deemed to be more rational. Second, we are working on
a solution for avoiding deception by AI agents; a method
which removes the incentive to deceive in any game by
transforming the game with a constraint on the reward
function of the AI agent [32]. Overall, we think there are
many exciting avenues for future work.</p>
        <p>
          Summary. We extend work on agent incentives [
          <xref ref-type="bibr" rid="ref53">2</xref>
          ] to
the multi-agent setting in order to functionally define
the incentive to (influence , signal to, and) deceive another
agent. Our definition of deception is general and relates
to a failure to signal the truth. In addition to canonical
signalling situations, it captures cases in which: no
information is signalled; deception occurs as a side-efect of
the signaller pursuing their goals (as in example 2); and
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>and Multiagent</given-names>
            <surname>Systems</surname>
          </string-name>
          , Richland,
          <string-name>
            <surname>SC</surname>
          </string-name>
          ,
          <year>2022</year>
          , p.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          1759-
          <fpage>1761</fpage>
          .
          <article-title>The authors are grateful to Henrik Aslund</article-title>
          ,
          <source>Matt Mac- [10] ANON, Defending Against Adversarial Artificial Dermott</source>
          , Tom Everitt,
          <source>James Fox, and the members of Intelligence</source>
          ,
          <year>2019</year>
          . URL: https://www.darpa.mil/ the Causal Incentives Working Group for helpful feed- news-events/2019-02-06, dARPA report.
          <article-title>back which significantly improved this work</article-title>
          .
          <source>Francis</source>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hubinger</surname>
          </string-name>
          , C. van
          <string-name>
            <surname>Merwijk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Mikulik</surname>
          </string-name>
          , J. Skalse,
          <article-title>was supported by UKRI [grant number EP/S023356/1], S. Garrabrant, Risks from learned optimization in the UKRI Centre for Doctoral Training in Safe and in advanced machine learning systems</article-title>
          ,
          <source>2019. Trusted AI</source>
          . arXiv:
          <year>1906</year>
          .
          <year>01820</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gorwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guilbeault</surname>
          </string-name>
          ,
          <article-title>Unpacking the Social MeReferences dia Bot: A Typology to Guide Research</article-title>
          and Policy,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>Policy &amp; Internet</source>
          <volume>12</volume>
          (
          <year>2020</year>
          )
          <fpage>225</fpage>
          -
          <lpage>248</lpage>
          . doi:
          <volume>10</volume>
          .1002/ [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Rof</surname>
          </string-name>
          ,
          <source>AI Deception: When Your Ar- poi3</source>
          .
          <fpage>184</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>tificial Intelligence</source>
          Learns to Lie, IEEE [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Marra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gragnaniello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Verdoliva</surname>
          </string-name>
          , G. Poggi, Do
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Spectr.</surname>
          </string-name>
          (
          <year>2021</year>
          ). URL: https://spectrum.ieee.
          <source>GANs leave artificial fingerprints?</source>
          , in: 2019 IEEE
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>org/ai-deception-when-your-ai-learns-to-lie</article-title>
          .
          <source>Conference on Multimedia Information Processing</source>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Everitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Carey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Langlois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Ortega</surname>
          </string-name>
          , and
          <string-name>
            <surname>Retrieval</surname>
          </string-name>
          (MIPR),
          <year>2019</year>
          , pp.
          <fpage>506</fpage>
          -
          <lpage>511</lpage>
          . doi:10.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Legg</surname>
          </string-name>
          ,
          <article-title>Agent incentives: A causal perspective</article-title>
          ,
          <source>in: 1109/MIPR</source>
          .
          <year>2019</year>
          .
          <volume>00103</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Thirty-Fifth AAAI</surname>
            Conference on Artificial Intel- [14]
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Yarats</surname>
            ,
            <given-names>Y. N.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Parikh</surname>
          </string-name>
          , D. Ba-
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>ligence</surname>
          </string-name>
          ,
          <source>AAAI</source>
          <year>2021</year>
          ,
          <article-title>Thirty-Third Conference on tra, Deal or No Deal? End-to-End Learning for Ne-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>Innovative Applications of Artificial Intelligence</source>
          , gotiation Dialogues, arXiv (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .48550/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>IAAI</source>
          <year>2021</year>
          ,
          <source>The Eleventh Symposium on Educa- arXiv.1706.05125</source>
          . arXiv:
          <volume>1706</volume>
          .
          <fpage>05125</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>tional Advances in Artificial Intelligence</source>
          , EAAI [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cotton-Barratt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Finnveden</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Bales,
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          2021,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          ,
          <source>February 2-9</source>
          ,
          <year>2021</year>
          , AAAI Press, A. Balwit,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wills</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Righetti</surname>
          </string-name>
          , W. Saunders, Truth-
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <year>2021</year>
          , pp.
          <fpage>11487</fpage>
          -
          <lpage>11495</lpage>
          . URL: https://ojs.aaai.org/ ful AI:
          <article-title>Developing and governing AI that does</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          index.php/AAAI/article/view/17368. not lie, arXiv (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2110. [3]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Mahon</surname>
          </string-name>
          ,
          <source>The Definition of Lying and Deception</source>
          ,
          <volume>06674</volume>
          . arXiv:
          <volume>2110</volume>
          .
          <fpage>06674</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>in: E. N.</given-names>
            <surname>Zalta</surname>
          </string-name>
          (Ed.), The Stanford Encyclopedia of [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Evans</surname>
          </string-name>
          , TruthfulQA: Mea-
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Philosophy</surname>
          </string-name>
          , Winter 2016 ed.,
          <source>Metaphysics Research suring How Models Mimic Human Falsehoods,</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Lab</surname>
          </string-name>
          , Stanford University,
          <year>2016</year>
          . arXiv (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2109.07958. [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kenton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Everitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Weidinger</surname>
          </string-name>
          , I. Gabriel, arXiv:
          <fpage>2109</fpage>
          .
          <fpage>07958</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>V.</given-names>
            <surname>Mikulik</surname>
          </string-name>
          , G. Irving, Alignment of language [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pfefer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <article-title>On the reasoning patterns</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          agents,
          <source>CoRR abs/2103</source>
          .14659 (
          <year>2021</year>
          ).
          <article-title>URL: https: of agents in games</article-title>
          ,
          <source>in: Proceedings of the</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          //arxiv.org/abs/2103.14659. arXiv:
          <volume>2103</volume>
          .14659.
          <string-name>
            <surname>Twenty-Second AAAI</surname>
            Conference on Artificial In[5]
            <given-names>M. D.</given-names>
          </string-name>
          <string-name>
            <surname>Hauser</surname>
          </string-name>
          ,
          <article-title>The evolution of communication, telligence</article-title>
          ,
          <source>July 22-26</source>
          ,
          <year>2007</year>
          , Vancouver, British
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          MIT press,
          <year>1996</year>
          . Columbia, Canada, AAAI Press,
          <year>2007</year>
          , pp.
          <fpage>102</fpage>
          -
          <lpage>[</lpage>
          6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Madry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Makelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsipras</surname>
          </string-name>
          , 109. URL: http://www.aaai.org/Library/AAAI/2007/
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vladu</surname>
          </string-name>
          ,
          <article-title>Towards deep learning models resistant to aaai07-015</article-title>
          .php.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>adversarial attacks</article-title>
          ,
          <source>arXiv preprint arXiv:1706</source>
          .06083 [18]
          <string-name>
            <given-names>V. J.</given-names>
            <surname>Baston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Bostock</surname>
          </string-name>
          , Deception Games, Int.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          (
          <year>2017</year>
          ).
          <source>J. Game Theory</source>
          <volume>17</volume>
          (
          <year>1988</year>
          )
          <fpage>129</fpage>
          -
          <lpage>134</lpage>
          . doi:
          <volume>10</volume>
          .1007/ [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. W. W.</given-names>
            <surname>Koh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Liang</surname>
          </string-name>
          , Certified BF01254543.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>defenses for data poisoning attacks</article-title>
          , Advances in [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fristedt</surname>
          </string-name>
          ,
          <article-title>The deceptive number changing game,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
          <article-title>in the absence of symmetry</article-title>
          ,
          <source>Int. J. Game Theory</source>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Everitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Krakovna</surname>
          </string-name>
          , Re-
          <volume>26</volume>
          (
          <year>1997</year>
          )
          <fpage>183</fpage>
          -
          <lpage>191</lpage>
          . doi:
          <volume>10</volume>
          .1007/BF01295847.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <article-title>ward tampering problems and</article-title>
          solutions in rein- [20]
          <string-name>
            <surname>I.-K. Cho</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Kreps</surname>
          </string-name>
          , Signaling Games
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>forcement learning: A causal influence diagram and Stable Equilibria, undefined (</article-title>
          <year>1987</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          perspective, CoRR abs/
          <year>1908</year>
          .04734 (
          <year>2021</year>
          ). URL: http: URL: https://www.semanticscholar.org/paper/
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          //arxiv.org/abs/
          <year>1908</year>
          .04734. arXiv:
          <year>1908</year>
          .04734.
          <string-name>
            <surname>Signaling-Games-</surname>
          </string-name>
          and
          <article-title>-</article-title>
          <string-name>
            <surname>Stable-Equilibria-</surname>
            Cho-Kreps/ [9]
            <given-names>F. R.</given-names>
          </string-name>
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Toni</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Belardinelli</surname>
          </string-name>
          , On agent in- d8bc1dbd8577d193e6eea2c944a251d1347f3adf.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <article-title>centives to manipulate human feedback in multi-</article-title>
          [21]
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Kovach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Gibson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. B.</given-names>
            <surname>Lamont</surname>
          </string-name>
          , Hyper-
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>the 21st International Conference on Autonomous and deception</source>
          ,
          <source>Game Theory</source>
          <year>2015</year>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Agents</surname>
            and
            <given-names>Multiagent</given-names>
          </string-name>
          <string-name>
            <surname>Systems</surname>
          </string-name>
          , AAMAS '
          <fpage>22</fpage>
          , In- [22]
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <article-title>Deception in game theory: a survey</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <year>2016</year>
          . [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Koller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Milch</surname>
          </string-name>
          , Multi-agent influence dia-
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Econ</surname>
          </string-name>
          . Behav.
          <volume>45</volume>
          (
          <year>2003</year>
          )
          <fpage>181</fpage>
          -
          <lpage>221</lpage>
          . URL: https://doi.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <source>org/10</source>
          .1016/S0899-
          <volume>8256</volume>
          (
          <issue>02</issue>
          )
          <fpage>00544</fpage>
          -
          <lpage>4</lpage>
          . doi:
          <volume>10</volume>
          .1016/
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <fpage>S0899</fpage>
          -
          <volume>8256</volume>
          (
          <issue>02</issue>
          )
          <fpage>00544</fpage>
          -
          <lpage>4</lpage>
          . [24]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hammond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Everitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abate</surname>
          </string-name>
          , M. J.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <source>CoRR abs/2102</source>
          .05008 (
          <year>2021</year>
          ). URL: https://arxiv.org/
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <source>abs/2102</source>
          .05008. arXiv:
          <volume>2102</volume>
          .
          <fpage>05008</fpage>
          . [25]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hadfield-Menell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Dragan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Abbeel</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. J.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <source>cial Intelligence</source>
          ,
          <source>Saturday, February 4-9</source>
          ,
          <year>2017</year>
          , San
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <surname>Francisco</surname>
          </string-name>
          , California, USA, volume WS-17
          <source>of AAAI</source>
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>Workshops</surname>
          </string-name>
          , AAAI Press,
          <year>2017</year>
          . URL: http://aaai.org/
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <article-title>ocs/index</article-title>
          .php/WS/AAAIW17/paper/view/15156. [26]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hammond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fox</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Everitt</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. C. A.</surname>
          </string-name>
          <year>Abate1</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          (Forthcoming). [27]
          <string-name>
            <given-names>R.</given-names>
            <surname>Carey</surname>
          </string-name>
          , Causal models of incentives (
          <year>2021</year>
          ). [28]
          <string-name>
            <given-names>P.</given-names>
            <surname>Christiano</surname>
          </string-name>
          ,
          <source>ARC's first technical report:</source>
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <string-name>
            <surname>Forum</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: https://www.alignmentforum.
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>org/posts/qHCDysDnvhteW7kRd/</mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <source>[Online; accessed 9. May</source>
          <year>2022</year>
          ]. [29]
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Halpern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kleiman-Weiner</surname>
          </string-name>
          , Towards for-
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          18),
          <source>the 30th innovative Applications of Artificial In-</source>
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <source>telligence (IAAI-18), and the 8th AAAI Symposium</source>
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          <source>(EAAI-18)</source>
          , New Orleans, Louisiana, USA, Febru-
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          <source>ary 2-7</source>
          ,
          <year>2018</year>
          , AAAI Press,
          <year>2018</year>
          , pp.
          <fpage>1853</fpage>
          -
          <lpage>1860</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          AAAI18/paper/view/16824. [30]
          <string-name>
            <given-names>R.</given-names>
            <surname>Selten</surname>
          </string-name>
          , Spieltheoretische behandlung eines
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          (
          <year>1965</year>
          )
          <fpage>301</fpage>
          -
          <lpage>324</lpage>
          . [31]
          <string-name>
            <surname>R. B. Myerson</surname>
          </string-name>
          ,
          <article-title>Game theory: analysis of conflict,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          Harvard university press,
          <year>1997</year>
          . [32]
          <string-name>
            <given-names>E.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Constrained Markov</surname>
          </string-name>
          Deci-
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          <source>lor &amp; Francis</source>
          , Andover, England, UK,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          <source>doi:10</source>
          .1201/9781315140223.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>