<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Promoting Training of Multi-Agent Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petro.O.Kravets@lpnu.ua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasyl.V.Lytvyn@lpnu.ua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victoria.A.Vysotska@lpnu.ua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yevhen.V.Burov@lpnu.ua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The problem of incentive training of multi-agent systems in the game formulation for collective decision making under uncertainty is considered. Methods of incentive training do not require a mathematical model of the environment and enable decision making directly in the training process. Markov model of stochastic game is constructed and the criteria for its solution are formulated. An iterative Q-method for solving a stochastic game based on the numerical identification of a characteristic function of a dynamic system in space of state-action is described. Players' current gains are determined by the method of randomization of payment Q-matrix elements. Mixed player strategies are calculated using the Boltzmann method. Pure strategies are determined on the basis of discrete random distributions given by mixed player strategies. The algorithm for stochastic game solving is developed and results of computer implementation of game Q-method are analyzed.</p>
      </abstract>
      <kwd-group>
        <kwd>- Multi-Agent System</kwd>
        <kwd>Stochastic Game</kwd>
        <kwd>Promotional Training</kwd>
        <kwd>Qmethod</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The functioning of most modern information systems (IS) is based on rigidly
programmed algorithms. Unforeseen environmental influences in such systems may
impair the stability of operating modes, which can lead to various types of emergency
situations. To prevent critical states, distributed IS software must consist of
interoperable standalone modules, be intelligent, flexible, and capable of independently
monitoring environmental changes and making timely and appropriate decisions.
Otherwise, such systems should be built on the principles of an agent-oriented methodology
[
        <xref ref-type="bibr" rid="ref1 ref10 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1 – 10</xref>
        ]. An IS agent is a standalone software module with elements of artificial
intelligence, capable of making decisions on its own, interacting with the environment,
other agents, and people as they accomplish the task. IS agents interact within the
computer network. A population of computer network agents who solve a common
problem is called a multi-agent system (MAS).
      </p>
    </sec>
    <sec id="sec-2">
      <title>The operation of the MAS is usually carried out in the context of a priori uncer</title>
      <p>tainty about the state of the decision-making environment and the actions of other</p>
    </sec>
    <sec id="sec-3">
      <title>Copyright © 2020 for this paper by its authors. Use permitted under Creative</title>
      <p>
        Commons License Attribution 4.0 International (CC BY 4.0).
agents. In this regard, agent’s behavior strategies must be adaptive at the expense of
agents' ability to learn [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Among the methods of learning under uncertainty,
incentive-based methods have gained practical appeal [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ] because they do not require
a mathematical model of the environment and provide decision-making power
directly in the learning process. The mechanisms of reflexive behavior of living
organisms with developed nervous system are the basis of the stimulating training. An
effective method of incentive learning is Markov Q-learning [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which performs
numerical identification of the characteristic function of a dynamic system in
statespace. As a characteristic function is usually the function of the total expected reward
agent.
      </p>
      <p>
        Compared to single agent systems, the structure, operation, and research of
multiagent Q-learning methods are much more complicated [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Due to the collective
interaction of agents, the stationary environment is transformed into a non-stationary
class. The change in the state of the environment and the value of the benefits of each
agent depend on the actions of the other agents. Generally, in an MAS, an agent
cannot achieve a maximum gain equal to that of a single agent system. The optimal
payoffs of agents must be balanced and meet the criteria of benefit, fairness, balance.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Thus, instead of the criterion of scalar maximization of the benefits of a single-agent</title>
      <p>
        system, the criteria of vector maximization of MAS winnings are introduced, for
example, Nash equilibrium, Pareto optimality or other [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Provided the use of methods of Q-learning of MAS, an iterative construction of a
system of characteristic Q-functions in the state-action space takes place, and the
increment of the elements of these functions is carried out in the direction of
achieving their collective equilibrium. To build the MAS, it is necessary to carry out
preliminary studies on the basis of adequate mathematical models that will allow to study
the dynamics of the system under uncertainty, to build strategies for the behavior of
agents that provide optimal technical and economic parameters of the system
functioning. Given the peculiarities of the subject area, namely, multi-agency, uncertainty
of decision-making environment, antagonism or competitiveness of goals,
communicativeness, coordination of actions, adaptability of agent behavior strategies, we use
the mathematical apparatus of stochastic game theory [
        <xref ref-type="bibr" rid="ref17 ref18 ref19 ref20 ref21 ref22 ref23 ref24 ref25 ref26 ref27 ref28 ref29">17 – 29</xref>
        ]. Solving a stochastic
game is to find the strategies of agents that maximize their winnings so as to ensure a
certain collective balance of interests for all players. The search for optimal strategies
for players in uncertainty will be performed on the basis of the promotional training
method.
      </p>
    </sec>
    <sec id="sec-5">
      <title>The purpose of the work is to construct an iterative method of incentive learning to solve the stochastic MAS game in uncertainty. To achieve this goal, it is necessary to develop a model of multi-agent stochastic game, to determine the criteria of collective equilibrium, the method and algorithm for solving the game problem.</title>
      <p>2</p>
      <p>The Mathematical Model of Stochastic Game</p>
    </sec>
    <sec id="sec-6">
      <title>The stochastic game is determined by the tuple:</title>
      <p>(S, p, Ai , ri |i  I ) , I  {1, 2,..., L} ,
where S  {s1,..., sM } is set of all states of the environment, p : S  A (S ) is
system state change function defined in the space of probability distributions (S ) on
the plural S , Ai  ai (1),..., ai ( Ni ) is multiple actions or pure strategies i -agent,</p>
      <sec id="sec-6-1">
        <title>A  iI Ai is set of combined agent actions, ri : S  A  R is reward function i -agent,</title>
        <sec id="sec-6-1-1">
          <title>I is multiple agents, L is number of agents, M is number of states, Ni is number</title>
          <p>of strategies i -agent.</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>In the general case of multiple actions Ai  Ai (s) i  I and combined actions</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>A  A(s) may depend on the state of the environment s  S .</title>
    </sec>
    <sec id="sec-8">
      <title>We adopt a Markov model [30, 31] of the dynamics of states of a system in which the probability of change of states p depends only on the current state of the environment and the current actions of the agents:</title>
      <p>pst1  s |(s , b ),  0,1, 2,..., t pst1  s | st , bt  ,
where bt  A is combined action at time t .</p>
      <p>
        At each point in time, the environment is in one of the states s  S and agents
choose actions independently ai  Ai . After the implementation of the combined
option a  (a1,..., aL )  A agents get random winnings ri (otherwise is incentives or
reinforcements), and the environment changes its state according to the probability
distribution p(s, a) with values per segment [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] :
 p(s | s, a)  1 .
      </p>
      <p>sS</p>
      <sec id="sec-8-1">
        <title>The agent implements actions based on a mixed strategy  i : S  Ai , which deter</title>
        <p>mines the likelihood of action ai  Ai in every state of the environment s  S .</p>
      </sec>
      <sec id="sec-8-2">
        <title>Distribution  i  i takes the value on a unit simplex [11]</title>
        <p> 
i     (s, ai )  1, (s, ai )  0 .</p>
        <p> aiAi 
If  i (s, ai ) {0, 1} , then the agent determines the choices of solutions. Let the total
payoff of each agent be determined by the function of discounted total payoffs:

i   trti ,
t0
(1)
where   (0, 1] is discount option.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>The goal i -agent is to maximize function (1) by formulating an effective strategy</title>
      <p> i :
2) Pareto optimality: V i (s, *) V i (s, ) .
3</p>
      <p>Learning of Stochastic Game
(2)
(3)
(4)
(5)</p>
    </sec>
    <sec id="sec-10">
      <title>Expression (4) defines a tabular function of the values of the action options a in the states s .</title>
    </sec>
    <sec id="sec-11">
      <title>Similarly to (3) we obtain:</title>
      <p>Vi (s)  E i | s0  s  max, i  I ,
 i
where   ( 1,..., L ) ; E is symbol of mathematical expectation.</p>
      <p>Stochastic game resolution is about defining agent behavior strategies  i*
( i  I ), which ensure fulfilment of one of the conditions of collective optimality, for
example:
1) Nash equilibrium: V i (s, 1*, 2*,..., L* ) V i (s, 1*, 2*,..., i*1, i , i*1,..., L* ) ;</p>
      <p>Q (s, a)  E R s0  s, a0  a .</p>
      <p>Q (s, a)  r(s, a)   p(s | s, a)V (s) .</p>
      <p>
        sS
Calculation Vi (s) can be performed in a recursive form known in the literature as
the Bellman equation [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12 – 14</xref>
        ]. Given (1), we obtain after simple transformations:

V (s | st  s)  E(rt )    k E(rtk 1 ) E(rt )  V (st1 ) 
      </p>
      <p>k 0
 r(s, (s))   p(s | s, (s))V (s)</p>
      <p>sS
where s is probable future states of the system.</p>
      <sec id="sec-11-1">
        <title>The agent's goal is to find a strategy  * that maximizes function (2) for all states</title>
        <p>of the environment:
 s  S
V  *(s)  V  (s) .</p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>Since the choice of options is made by chance, it is for the purpose of comparing the</title>
      <p>effectiveness of actions when the system is in a state s  S , current gains are useful to
obtain from (2). For this purpose is specially built Q is average payoff feature that
determines the cost of the action – the total payoff of the agent in the state s chose
action a :</p>
    </sec>
    <sec id="sec-13">
      <title>Adherence to the Bellman principle of optimality (5) ensures the optimal gain of the</title>
      <p>agent from the current state achieved s  S at all future times. Applying this principle
to all states ensures a global optimal solution.</p>
      <sec id="sec-13-1">
        <title>For optimal function selection strategies  * for each state s  S we will get:</title>
        <p>(6)
(7)
 
V * (s)  max r(s, a)   p(s | s, a)V * (s) .
aA  sS </p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>From (6) one can obtain the optimal function of choosing strategies</title>
      <p> *(s)  arg max Q * (s, a) .
aA</p>
    </sec>
    <sec id="sec-15">
      <title>Optimization (7) can be performed by dynamic programming methods [30].</title>
    </sec>
    <sec id="sec-16">
      <title>By analogy to single agent training, we define the payoff matrix i -players with</title>
      <p>current and future winnings in the direction of movement to the optimal collective
state in space S  A :</p>
      <p>Q*i (s, a1,..., aL )  ri (s, a1,..., aL )   p(s | s, a1,..., aL ) V i (s, 1*,..., L* ) ,
sS
where Q*i (s, a1,..., aL ) is total discounted gain i -player provided the players select
the action (a1,..., aL ) in the state s according to the optimal strategy of the game
 *  ( 1*,..., L* ) .</p>
    </sec>
    <sec id="sec-17">
      <title>In the conditions of a priori uncertainty of the transition probabilities between the</title>
      <p>
        states of the system p(s, a1,..., aL ) and winnings features ri (s, a1,..., aL ) an iterative
method is used to calculate the elements of the payoff matrix Q -learning [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]:
Qti1(s, a1,..., aL )  (1  t )Qti (s, a1,..., aL )  t [rti Vti (st1)]
(8)
where  t  (0, 1) is training option; Vti (st1) is the operator of the cost of the system
state in the direction of the optimal collective solution.
      </p>
      <p>The type of the operator Vti (st1) is determined by the condition of collective
equilibrium, for example:</p>
      <sec id="sec-17-1">
        <title>Vti (st1)  MM Qt (st1) is maximin equilibrium;</title>
      </sec>
      <sec id="sec-17-2">
        <title>Vti (st1)  NE Qti (st1) is Nash equilibrium;</title>
      </sec>
      <sec id="sec-17-3">
        <title>Vti (st1)  BRQti (st1) is best agent response;</title>
      </sec>
      <sec id="sec-17-4">
        <title>Vti (st1)  CE Qti (st1) is correlated equilibrium;</title>
      </sec>
      <sec id="sec-17-5">
        <title>Vti (st1)  PE Qti (st1) is Pareto optimality.</title>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>The above list may be supplemented by other already known and new equilibrium</title>
      <p>states that will determine the target aspect of the functioning of a distributed dynamic
system. Method (8) can be applied to decipher a single agent game with nature as a
partial case N-agent stochastic game if I  {i},| I | 1, A  Ai , when
Vti (st1)  max Qti (s, b) , where s  st1 .</p>
      <p>bAi</p>
    </sec>
    <sec id="sec-19">
      <title>The Maximin Equilibrium (MM) takes place in the game of two agents with zero sum of their payoff functions:</title>
      <p>V 1(st1)  max min  1(s, a1)Qt (s, a1, a2 )  V 2 (st1) .</p>
      <p>11 a2A2 a1A1</p>
    </sec>
    <sec id="sec-20">
      <title>Nash Equilibrium (NE) is determined by the independent distribution of strategies by</title>
      <p>players who choose their own strategies independently of the choice of other agents.</p>
      <sec id="sec-20-1">
        <title>In a Nash equilibrium situation in mixed strategies  NE (s)   1NE (s),..., LNE (s) it is</title>
        <p>
          not profitable for each agent to deviate from its own optimal strategy  iNE (s) , if other
agents stick to the equilibrium point [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]:
        </p>
        <p>L L
Qti (s, a) iNE (ai ) NjE (s, a j )  Qti (s, a)~i (ai ) NjE (s, a j ) ,
aA ji aA ji
(9)
where a  (a1,..., aL ) ;  iNE , ~i  i .</p>
      </sec>
    </sec>
    <sec id="sec-21">
      <title>Method (8) ensures that condition (9) is satisfied when the current value of the sys</title>
      <p>tem state value operator is determined at point  NE Nash equilibrium:
L
NEQti (st1)  Qti (s, a) NjE (s, a j ) .</p>
      <p>aA j1</p>
    </sec>
    <sec id="sec-22">
      <title>The set of NE equilibrium points in mixed strategies is a convex compact and can be calculated by linear programming methods (for bi-matrix games) or by solving a system of polylinear equations that determine the complementary rigidity condition:</title>
    </sec>
    <sec id="sec-23">
      <title>Unlike the Nash equilibrium, the Best Response (BR) method generates an optimal agent strategy in response to the actions of all other agents. The corresponding system state value operator in method (8) has the form [32]:</title>
      <p> L 
BRQti (st1)  mai x aAQti (s, a) j (s, a j )  .

j1 </p>
    </sec>
    <sec id="sec-24">
      <title>Correlated Equilibrium (CE) generalizes the Nash equilibrium by allowing players'</title>
      <p>strategies to depend. For this purpose, there is an arbitrator in the collective
decisionmaking system, which according to the generalized distribution   ( A)</p>
      <p>L
(  (a)  1 , A  i1 Ai ) recommends that players choose to take actions that form a
aA
combined option a  (a1,..., aL ) . Player with a number i only receives component
information combined option a  A . This signal is perceived i -player as an optional
offer to take action ai . Each player secretly and independently chooses at a moment's
time t action option ai , possibly different from the proposed variant, and receives a
current payoff ri (st , at ) , which is a function of the current state of the system st and
the combined option at  A . The environment then moves to a new state st1
according to the probability distribution p(st1 | st , at ) and the process is repeated at
time t  1 .</p>
    </sec>
    <sec id="sec-25">
      <title>Correlated equilibrium is determined by the united distribution of player strategies</title>
      <p>
          ( A) , when each agent does not have the motivation to deviate unilaterally [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]:
 CE (ai | ai )Qi [s, (ai , ai )] 
      </p>
      <p> CE (ai | ai )Qi [s, (ai , a~i )],
aiAi
 CE (ai , ai ) ,  CE (ai | ai )   CE (ai , ai )  CE (ai ) ,  CE (ai )  0 .
points are correlated equilibrium
points. If
i  I
 CE (ai | ai )   CE (ai | a~i ) , ai , a~i  Ai , ai  Ai ,  CE (ai ), CE (a~i )  0 , then
correlated equilibrium is also Nash equilibrium. To solve the game by method (8), the
cost operator Vti (st1) the state of the system is determined by the point  CE
correlated equilibrium: CE(Qti (st1))   CE (a)Qti (s, a) . The set of CE equilibrium
aA
points is non-empty, convex and compact and can be effectively calculated using
linear programming methods. In the case of maximizing total player winnings, the
task of linear programming is to find  CE can be formulated as follows:</p>
      <p>L
 (a)Qi (s, a)  max ,
aA i1 </p>
      <p> (ai , ai )Qi [s, (ai , ai )]  Qi [s, (ai , a~i ) 0 ,
i  I , ai  Ai , a~i  Ai , s  S ,  (a)  0 a  A ,  (a)  1 .
aA
Based on the distribution  (a) a  A players' own strategies are determined
 i i  I . There are various options for switching from  to  i depending on the
type of strategies and players' level of awareness. For example, pure strategies
determine the maximum value of the operator CE(Qi (s)) :
ai  arg max CE(Qi (s)) .</p>
      <p>
 (ai )  1 , it can be assumed that mixed strategies
aiAi
Taking into account that
 i (ai )   (ai ) </p>
      <p> (ai , ai ) , or:
aiAi
 i (ai ) </p>
      <p> (ai )Qi (s, ai , ai ) CE(Qi (s)) .</p>
      <p>aiAi</p>
    </sec>
    <sec id="sec-26">
      <title>Pareto Equilibrium (PE) optimality occurs in the Common-Interest Markov Game,</title>
      <p>
        where the payoff matrices are the same for all players Qti (s, a)  Qtj (s, a) i, j  I ,
s  S, a  A [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. A game with different payoff matrices can be turned into a
game of shared interests by a convolution
      </p>
      <p>L L
PEQt (st1)  k Qtk (s, a) PjE (s, a j ) ,</p>
      <p>k 1 aA j1
where  j  0 ( j  1..L) .</p>
    </sec>
    <sec id="sec-27">
      <title>The search for the game's PE solution is done independently by agent strategies,</title>
      <p>similar to the search for a NE solution. A multi-agent game is optimal for Pareto if
there is no common player strategy that improves the winnings of all players:
Qti (s, PE )  Qti (s, ) . Pareto-optimal mixed strategies  PE (s)   1PE (s),..., LPE (s)
can be obtained by maximizing the convolution of concave (up) payoff functions:
L L
k Qtk (s, a) j (s, a j )  max .</p>
      <p>k 1 aA j1 </p>
      <sec id="sec-27-1">
        <title>To calculate optimal collective strategies  *(s)   1*(s),..., L* (s) (NE, CE, PE)</title>
        <p>
          agent with number i need to know Q -functions of all agents:
Q(s)  Qt1(s),..., QtL (s). In the absence of such information, each agent should
evaluate the value Q - functions in the learning process. For this i -agent monitors the
current gains of other agents and modifies their estimates Q -functions according to
(8). To ensure that method (8) converges to one of the points of collective
equilibrium, it is necessary to impose a limit on the rate of change of its adjustable
parameters. The general limitations are as follows [
          <xref ref-type="bibr" rid="ref11 ref14 ref34">11, 14, 34</xref>
          ]:

 t  ,
t0

 t2   ,
t0
(10)
where  t  t  (  0) is monotonically decreasing positive sequences of real values.
4
        </p>
        <p>Stochastic Game Solving Algorithm</p>
      </sec>
    </sec>
    <sec id="sec-28">
      <title>Step 1. Set the start time t  0 ; the initial values of the payoff matrices</title>
      <p>Qti (s, a1,..., aL )   s  S , ai  Ai , i  I , where 0    1 is small
positive value; the value of the gain discount parameter   (0, 1] ; the initial state
of the system s0 .</p>
      <p>Step 2. Perform a random selection of agent actions a  (a1,..., aL ) based on
strategies   ( 1,..., L ) . The value of the strategies can be calculated from the
current estimates of the payoff matrices i  I :
 i (s, ai ) 
Qti (s, ai , ai )
Qti (s, a) , ai  Ai .
aA
Step 3. Get current agent payouts rt  (rt1,..., rtL ) .</p>
      <sec id="sec-28-1">
        <title>Step 4. Determine the new state of the system st1  st (a1,..., aL ) .</title>
        <p>Step 5. Calculate function Vti (st1) according to the specific condition of collective
equilibrium (NE, CE, PE).</p>
        <sec id="sec-28-1-1">
          <title>Step 6. Modify the payoff matrix Qt1  Qti1(st , at ) | i  1..L according to (8).</title>
          <p>Step 7. If Qti1  Qti   i  1..L , then ask t : t  1 and go to step 2.</p>
        </sec>
        <sec id="sec-28-1-2">
          <title>Step 8. Print the calculated values of the payoff matrices Q(s)  Q1(s),...,Q L (s) and</title>
          <p>strategies  (s)  ( 1(s),..., L (s)) s  S . End of algorithm.
5</p>
          <p>The Results of Computer Simulation</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-29">
      <title>Let's solve the stochastic game of two agents with two pure strategies in a two-state environment. The matrices of the average payoffs of such a game are given in Table 1.</title>
    </sec>
    <sec id="sec-30">
      <title>States</title>
      <p>s1
s2</p>
      <p>Strategies</p>
      <p>
        
 1(s1, a1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ])
 1(s1, a1[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ])
      </p>
      <p>
        
Under uncertainty, the elements of the average payoff matrix vi (s, a)sS a priori
aA
unknown and available for observation in the form of random current values
ri (s, a)  Normal(vi (s, a), d i (s, a)) ,
distributed by normal law with mathematical expectation vi (s, a) and dispersion
d i (s, a) . Normally distributed random variables are obtained by summing twelve
evenly distributed random numbers  [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] :
(11)
(12)
(13)
ri (s, a)  vi (s, a) 
      </p>
      <p> 12 
d i (s, a)  j1 j  6 ,
where d i (s, a)  d  0 s  S, a  A . If at time t the system was in a state of
disrepair s  S , then after implementing pure strategies a  (a1,..., aL ) , where</p>
      <p>L
a  A   Ai , agents receive current winnings rti (s, a) , calculated according to (11).</p>
      <p>i1</p>
    </sec>
    <sec id="sec-31">
      <title>After receiving current winnings, each agent lists the corresponding item Q</title>
      <p>matrix according to the algorithm modified for uncertainty conditions BR :
Qti1(s, a1,..., aL )  (1  t )Qti (s, a1,..., aL )  t (rti   max Qti (s, a1,..., aL )) .
ai</p>
    </sec>
    <sec id="sec-32">
      <title>Based on Q -matrices the current values of mixed strategies are calculated using the</title>
      <p>Boltzmann method:
 i (ai (k ) | s)  eQi*(s,ai (k )) /T
Ni eQi*(s,ai ( j)) /T , k  1..Ni ,
j1
where Qi*(s, ai (k ))  max ri (ai , ai (k )) , ai  Ai , Ai  L A j , T  0 is
temperaai jj1i
ture coefficient.</p>
      <sec id="sec-32-1">
        <title>Vector elements of mixed strategies  i define a discrete distribution by which the</title>
        <p>
          values of random pure strategies are determined i -agent at the next time:
  k  
ai (s)   Ai (s, k ) k  arg min  i (s, ai ( j))   , k  1.. Ni  , s  S , i  I , (14)
  k j1  
where  [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] is random variable with uniform distribution.
        </p>
      </sec>
    </sec>
    <sec id="sec-33">
      <title>The change of states of the dynamic system is determined by the discrete distribu</title>
      <p>tion p(s | s, a)  p s  S, a  A :</p>
      <p>  k  
s  S (k ) k  arg min  p( j)   , k  1.. M  .
  k j1  
(15)</p>
    </sec>
    <sec id="sec-34">
      <title>Let the states change s  S system is implemented with equal probabilities</title>
      <p>p(s | s, a)  (| S |k1, k  1..M ) s  S, a  A , that is p(s | s, a)  (0.5; 0.5) for
| S | 2 . Agents' average payoffs are calculated taking into account the transition
probabilities of the medium from one state to another: V i   p(s)V i (s) , i  1..L ,
sS</p>
      <p>L
where V i (s)   vi (s, a) j (s, a) is average agent gain in the state s  S .</p>
      <p>aA j1</p>
    </sec>
    <sec id="sec-35">
      <title>The trajectories of changing agent strategies within a unit simplex and the appear</title>
      <p>ance of average payoff functions V i (s) , which correspond to the data of the table 1,
is shown in Fig. 1 and Fig. 2.</p>
      <sec id="sec-35-1">
        <title>The BR method (12 – 15) ensures that stochastic play is solved at the vertices of a</title>
        <p>unit simplex with the maximum value of the mean gain function. The percentage of
options for achieving optimal game resolution depends on the absolute difference
between the two largest consecutive values of the average payoff features.</p>
      </sec>
    </sec>
    <sec id="sec-36">
      <title>The convergence of the method is estimated by the error of fulfillment of the com</title>
      <p>
        plementary slackness condition [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ], weighted by mixed strategies:
  L1  i ~i 2 , where ~i  diag( i ) Vi /V i ; diag( i ) is diagonal square
iI
matrix of order Ni , formed from vector elements  i ; V i  V i [ j] j  1..Ni  is
vecNi
tor median payoff function for fixed net strategies i -player; V i  V i [ j] i [ j] is
j1
average payoff function i -player;  is Euclidean vector norm. The complementary
slackness condition characterizes the gameplay in Nash mixed strategies. The
weighted condition additionally takes into account the game's solutions in pure
strategies. Average payoff function graphs  and norms for deviating mixed strategies
from their target values  filed in Fig. 3.
strategies from their target indicates the convergence of the game Q -method.
      </p>
    </sec>
    <sec id="sec-37">
      <title>The value of the temperature coefficient T has a significant impact on the conver</title>
      <p>gence of the game method. The rate of convergence is determined by the rapid decline
of the function graph  , which can be estimated by the value of the acute angle of
linear approximation of the function graph  with the time axis. With the growth T
rate of convergence of game Q -method decreases.
6</p>
      <p>Conclusions
The promotional training method (8) considered in the deterministic version requires
the knowledge of each agent Q -functions of all other agents. These functions are
used by agents to identify strategies that provide method dynamics toward the points
of collective equilibrium. Value Q -functions can be obtained by exchanging
information between agents. If integrated information about Q -functions are not available to
the agent, then he must determine their value independently in the learning process,
observing the current benefits of other agents and performing evaluations Q
functions according to (8). If such observations are not possible, the agents may
perform reflective assessments Q -functions of other agents.</p>
    </sec>
    <sec id="sec-38">
      <title>Another method of constructing incentive training algorithms for agents under uncertainty is to apply the stochastic approximation method to the corresponding collective equilibrium condition.</title>
    </sec>
    <sec id="sec-39">
      <title>The practical use of game-based promotional training methods requires their prior</title>
      <p>analysis to determine the conditions for convergence to a state of collective
equilibrium. Such studies are based on the evaluation of sequences of random variables,
which characterize the current deviations of players' strategies from their optimal
values.</p>
      <p>The rate of convergence of the game method of Q-learning is determined by the
parameters  t and T. Parameter  t must satisfy the general conditions of stochastic
approximation (10). The value of the parameter T depends on the absolute values of
the elements of Q-matrices. It is experimentally established that for the given matrices
of average payoffs the convergence of the game Q-method is provided at T  (0, 0.2]
in the parameter value range  t  t  ,   (0, 1] . The highest convergence rate of the
game method of promotional learning is achieved at T  102 .</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Weiss</surname>
          </string-name>
          , G.:
          <article-title>Multiagent Systems</article-title>
          .
          <source>Second Edition</source>
          . The MIT Press.
          <article-title>(</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chong</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <article-title>: Multi-agent Systems Support for Community-Based Learning</article-title>
          .
          <source>Interacting with Computers</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <fpage>33</fpage>
          -
          <lpage>55</lpage>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kravets</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The Methodology of Multi-Agent Systems: a Modern State and Future Trends</article-title>
          .
          <source>In: Proceedings of the International Conference on computer science and information technologies - CSIT</source>
          '
          <year>2006</year>
          ,
          <fpage>125</fpage>
          -
          <lpage>127</lpage>
          . (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dignum</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradshaw</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silverman</surname>
            ,
            <given-names>B.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doesburg</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Agent for Games and Simulations: Trends in Techniques, Concepts</article-title>
          and
          <source>Design</source>
          . Springer. (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kravets</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The Control Agent with Fuzzy Logic</article-title>
          .
          <source>In: Perspective Technologies and Methods in MEMS Design, MEMSTECH</source>
          ,
          <fpage>40</fpage>
          -
          <lpage>41</lpage>
          . (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Scerri</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vincent</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mailler</surname>
          </string-name>
          , R.T.:
          <article-title>Coordination of Large-Scale Multiagent Systems</article-title>
          . Springer. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Byrski</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kisiel-Dorohinicki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <string-name>
            <given-names>Evolutionary</given-names>
            <surname>Multi-Agent</surname>
          </string-name>
          <string-name>
            <surname>Systems</surname>
          </string-name>
          : From Inspirations to Applications. Springer. (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Radley</surname>
          </string-name>
          , N.:
          <string-name>
            <surname>Multi-Agent Systems</surname>
          </string-name>
          - Modeling, Control, Programming,
          <source>Simulations and Applications</source>
          . Scitus Academics LLC. (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>J.-X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Iterative Learning Control for Multi-Agent Systems Coordination</article-title>
          . Wiley-IEEE Press.
          <article-title>(</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Cooperative Coordination and Formation Control for Multi-Agent Systems</article-title>
          . Springer. (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nazin</surname>
            ,
            <given-names>A. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poznyak</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          :
          <article-title>Adaptive Choice of Variants: Recurrence Algorithms (in russian)</article-title>
          . Moscow: Science. (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kaelbling</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Littman</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          :
          <article-title>Reinforcement learning: A survey</article-title>
          .
          <source>In: Journal of Artificial Intelligence Research</source>
          ,
          <volume>4</volume>
          ,
          <fpage>237</fpage>
          -
          <lpage>285</lpage>
          . (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sutton</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barto</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>Reinforcement Learning: An Introduction</article-title>
          . MIT Press.
          <article-title>(</article-title>
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Watkins</surname>
            ,
            <given-names>C.J.C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dayan</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Q-Learning</article-title>
          .
          <source>In: Machine Learning</source>
          , Kluwer Academic Publishers, Boston,
          <volume>8</volume>
          ,
          <fpage>279</fpage>
          -
          <lpage>292</lpage>
          . (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ummels</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Stochastic Multiplayer Games: Theory and Algorithms</article-title>
          . Amsterdam University Press. (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Ungureanu</surname>
          </string-name>
          , V.:
          <string-name>
            <surname>Pareto-Nash-Stackelberg Game</surname>
          </string-name>
          and
          <source>Control Theory: Intelligent Paradigms and Applications</source>
          . Springer. (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Dynamic Computation Offloading for Mobile Cloud Computing: A Stochastic Game-Theoretic Approach</article-title>
          .
          <source>In: IEEE Transaction on Mobile Computing</source>
          ,
          <volume>18</volume>
          (
          <issue>4</issue>
          ),
          <fpage>771</fpage>
          -
          <lpage>786</lpage>
          . (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.-S.:</given-names>
          </string-name>
          <article-title>Stochastic Game Strategies and their Applications</article-title>
          . CRC Press.
          <article-title>(</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kravets</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasichnyk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kunanets</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veretennikova</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <article-title>Game Method of Event Synchronization in Multiagent Systems</article-title>
          .
          <source>In: Advances in Intelligent Systems and Computing (AISC)</source>
          ,
          <source>ICCSEEA 2019 - Proceedings</source>
          ,
          <volume>938</volume>
          ,
          <fpage>378</fpage>
          -
          <lpage>387</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Kravets</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lytvyn</surname>
            ,
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vysotska</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Gaming Method of Ontology Clusterization</article-title>
          . In: Webology,
          <volume>16</volume>
          (
          <issue>1</issue>
          ),
          <fpage>55</fpage>
          -
          <lpage>76</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>You</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A Stochastic Game Theoretic Framework for Decentralized Optimization of Multi-Stakeholder Supply Chain Under Uncertainty</article-title>
          . In: Compute &amp; Chemical Engineering,
          <volume>122</volume>
          ,
          <fpage>31</fpage>
          -
          <lpage>46</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Lalropuia</surname>
            ,
            <given-names>K. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Modeling Cyber-Physical Attacks Based on Stochastic Game and Markov Processes</article-title>
          . In: Reliability Engineering &amp; System Safety,
          <volume>181</volume>
          ,
          <fpage>28</fpage>
          -
          <lpage>37</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Lozovanu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Pure and Mixed Stationary Nash Equilibria for Average Stochastic Positional Games</article-title>
          .
          <source>In: Frontiers of Dynamic Games</source>
          ,
          <fpage>131</fpage>
          -
          <lpage>155</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Garrec</surname>
            ,
            <given-names>T. Communicating</given-names>
          </string-name>
          <string-name>
            <surname>Zero-Sum Product Stochastic</surname>
          </string-name>
          <article-title>Games</article-title>
          .
          <source>In: Journal of Mathematical Analysis and Applications</source>
          ,
          <volume>477</volume>
          (
          <issue>1</issue>
          ),
          <fpage>60</fpage>
          -
          <lpage>84</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Kloosterman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Cooperation in Stochastic Games: a Prisoner's Dilemma Experiment</article-title>
          .
          <source>In: Experimental Economics</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>S. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beling</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>A. Multi-Agent Inverse Reinforcement Learning for Certain General-Sum Stochastic Games</article-title>
          .
          <source>In: Journal of Artificial Intelligence Research</source>
          ,
          <volume>66</volume>
          ,
          <fpage>473</fpage>
          -
          <lpage>502</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Saldi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basar</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raginsky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Approximate Nash Equilibria in Partially Observed Stochastic Games with Mean-Field Interactions</article-title>
          .
          <source>In: Mathematics of Operations Research</source>
          ,
          <volume>44</volume>
          (
          <issue>3</issue>
          ),
          <fpage>1006</fpage>
          -
          <lpage>1033</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Yamamoto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Stochastic Games with Hidden States</article-title>
          .
          <source>In: Theoretical Economics</source>
          ,
          <volume>14</volume>
          ,
          <fpage>1115</fpage>
          -
          <lpage>1167</lpage>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Fudenberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levine</surname>
            ,
            <given-names>D.K.</given-names>
          </string-name>
          :
          <article-title>The Theory of Learning in Games</article-title>
          . Cambridge. (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Puterman</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          :
          <article-title>Markov Decision Processes: Discrete Stochastic Dynamic Programming</article-title>
          . John Wiley &amp; Sons, New York. (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wellman</surname>
            ,
            <given-names>M. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nash</surname>
            <given-names>Q</given-names>
          </string-name>
          <article-title>-learning for general-sum stochastic games</article-title>
          .
          <source>In: Machine Learning Research</source>
          ,
          <volume>4</volume>
          ,
          <fpage>1039</fpage>
          -
          <lpage>1069</lpage>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Weinberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenschein</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          :
          <article-title>Best-Response Multiagent Learning in Non-Stationary Environments</article-title>
          . In: AAMAS'
          <fpage>04</fpage>
          , New York, USA. (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Greenwald</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Correlated Q-learning</article-title>
          .
          <source>In: Proceedings of the Twentieth International Conference on Machine Learning</source>
          ,
          <fpage>242</fpage>
          -
          <lpage>249</lpage>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Kushner</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>G. G.</given-names>
          </string-name>
          :
          <article-title>Stochastic Approximation and Recursive Algorithms</article-title>
          and Applications. Springer Science &amp; Business Media. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Neogy</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bapat</surname>
            ,
            <given-names>R. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <source>Mathematical Programming and Game Theory</source>
          . Springer. (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>