<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Stochastic Dichotomous Game-Theoretic Model of Technology Efficiency</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladimir Tsyganov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Control Sciences</institution>
          ,
          <addr-line>Profsoyuznaya str. 65, 117997 Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Theoretical, algorithmic and methodological aspects of stochastic modeling and technology efficiency control based on game-theoretic approach and machine learning are considered. The problem of assigning one of two ranks to a control object with the stochastic potential of technology is posed and solved for the case when probabilistic characteristics are known. Otherwise, to determine the optimal parameter of the ranking rule it is proposed to use the procedure of machine learning. The case of asymmetric awareness of the manager deciding on the ranking and the staff responsible for the effectiveness of the technology is considered. Far-sighted staff selects indicators of technology efficiency in such a way as to maximize own objective function, which depends on current and future results of ranking. There is a game between staff and manager which can lead to a decrease in the effectiveness of the technology and distortion of the estimates of the ranking parameters. This makes machine learning ineffective. To solve these problems, stochastic game model and ranking learning mechanism are proposed. The results of this mechanism functioning are estimates of ranking parameters, standards and ranks that determine staff stimuli. Sufficient conditions for the synthesis of ranking learning mechanism have been found, allowing to reveal the potential of technology effectiveness and to determine the optimal parameters of the ranking rule. These conditions are illustrated by the example of machine learning of ranking the technology electricity effectiveness in the process of implementing the program to increase the energy efficiency of the Russian Railways holding.</p>
      </abstract>
      <kwd-group>
        <kwd>stochastic modeling</kwd>
        <kwd>game theory</kwd>
        <kwd>control</kwd>
        <kwd>learning</kwd>
        <kwd>ranking</kwd>
        <kwd>stimulation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>machine</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>The important aspect of production efficiency in the face of changes is the disclosure
of internal reserves and technological resources by activating the human factor. This
determines the relevance of the game-theoretic approach to stochastic modeling of the
dynamics of organizational and technical systems. This approach allows take into
account conflicts of interest and the resulting activity of people in the production
process.</p>
      <p>The problem of reconciling these interests under conditions of uncertainty has
traditionally been considered in hierarchical games, or inverse Stackelberg games first
introduced in [1]. In Russian/Soviet literature, hierarchical games were previously
considered by Yu. Germeier [2].</p>
      <p>At the heart of modern research of organizational systems control, taking into
account the activity of their elements is a game-theoretic approach. Among the
scientific directions, it should be noted first of all, the theory of mechanisms
(mechanism design), which has incorporated the control theory, the theory of
contracts and the theory of feasibility [3]. It should also be noted theory of active
systems [4], as well as works on the study of individual models using the
gametheoretic approach for example [5].</p>
      <p>On the other hand, at present, the theoretical, algorithmic and methodological
aspects of stochastic game-theoretic modeling of organizational and technical systems
with elements of artificial intelligence are acquiring more importance. For example,
within the framework of the theory of active systems, mathematical methods are
developed for smart organization mechanism design [4].</p>
      <p>The most important element of artificial intelligence is learning. Accordingly,
game-theoretic approach support mechanism design with learning. For example, a
behavioral model for mechanism design is based on individual evolutionary learning
[6].</p>
      <p>In recent years, machine learning has attracted a lot of interest from both scientists
and managers of organizational and technical systems. But the development of
machine learning at the beginning of the 21st century has led to a certain divergence
with control theory [7]. Control-related views have only been published for narrower
areas of iterative learning by Bristow et al. [8] and of reinforcing learning by Recht
[9]. What is a challenge for control theorists is that there is very little rigorous
mathematical proof in the new machine learning “tsunamis” [7]. Today this
divergence is the subject of discussion in the control community on “Control and
learning – is there really a divide?” Among the issues currently being discussed is:
How can we use control theory to improve machine learning algorithms? [10].</p>
      <p>To answer these questions in relation to control of organizational and technical
systems, mechanisms have been designed in recent years using machine learning
algorithms. So, in [11] on the basis of a game-theoretic approach the mechanism for
learning digital control of a large-scale industrial system was designed. In [12], a
business control mechanism based on a learning algorithm with a tutor proposed. In
[13], a control mechanism of production was designed based on a self-learning of
identification. This paper discusses a game-theoretic approach to designing a
mechanism that uses a self-learning algorithm for ranking and stimulation in
organizational and technical system.</p>
    </sec>
    <sec id="sec-3">
      <title>2 Stochastic Dichotomy</title>
      <p>Many tasks of governing body (briefly – Center) come down to assigning one of two
ranks to the control object (briefly – dichotomous ranking). The better the rank, the
higher the incentive for the controlled object. For this, a decision rule is needed. With
enough complete a priori information, Center can use the rules of the theory of
statistical decisions.</p>
      <sec id="sec-3-1">
        <title>2.1 Minimization of Losses in Dichotomy</title>
        <p>Denote by t the time period, t = 0,1,... The sequence t = 0,1,... is a chronologically
sorted sequence of time intervals. These intervals are constant in length, not
overlapping and cover the whole timeline. Unit of time is used here correspond to the
application problem of organization control is considered (for example month, see
Section 4).</p>
        <p>
          Let zt be a random variable characterizing the efficiency potential of the object in
period t, zt ∈ ∆, where ∆ = [λ, σ] is the finite subset of R1 : ∆ ⊂ R1, λ &lt; 0, σ &gt; 0.
Variables zt corresponding different periods t = 0,1,... , are independent identically
distributed random variables. Dichotomous ranking involves zt assigning one of the
two areas that make up the set ∆ . Incorrect ranking leads to losses.
2
Denote {∆1, ∆2} some partition of the set ∆ into 2 areas, U ∆ k = ∆ . When
k =1
ranking, i.e. assigning a situation zt to one of these areas, Center makes a decision
associated with some losses. The challenge is to define a partition {∆1, ∆2} that
minimizes the average losses associated with ranking. We introduce for each, so far
unknown area ∆ k , k = 1,2, the dichotomous ranking loss function:
– L1( c,z ) – losses in case of assignment z the rank 1, while z ∈ ∆2 ;
– L2( c,z ) – losses in case of assignment z the rank 2, whereas z∈∆1,
where c is an unknown parameter. Then the affiliation z of one or another area is
determined by the sign of the decision rule
µ12(c,z) =L1(c,z)_L2(c,z): z∈∆1 for µ12 (c, z) &lt; 0, and z ∈ ∆2 for μ12 (c, z) ≥ 0. (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
Assume that Center knows the distribution density q( z ) of a random variable z.
Then the problem is solved by determining the parameter с of the decision rule (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
that minimizes the average losses of ranking:
        </p>
        <p>
          2
J (c ) = Σ ∫ Lk (c, z )q ( z )dz a min
k =1 ∆ k c
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2 Learning of Stochastic Dichotomy</title>
        <p>
          Knowing q( z ) we can find the parameter as a solution to problem (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ). However, in a
stochastic setting a priori information is often not enough. Suppose q( z ) it is
unknown to Center. Therefore, the direct determination of the parameter c by solving
optimization problem (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) is impossible. Then we can then try to tune decision rule
parameter с using observations zt , t = 0,1,... , to minimize losses J ( c ).
        </p>
        <p>
          Write the condition for the minimum average losses of ranking (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) in the form:
M z {
2 )
∑ Bk ( c,z )∂Lk ( c,z ) / ∂c }= 0 ,
k =1
)  1 if
Bk (c, z ) = 
 0
        </p>
        <p>z ∈ Δ k
otherwise
,
where M z is the mathematical expectation operator. For simplicity of calculations,
consider below the linear loss functions:</p>
        <p>L1(c, z) = z −vc,</p>
        <p>L2(c, z) =d(c−z),
where</p>
        <p>– v is the parameter of loss elasticity when z is assigned rank 1, while
z ∈ ∆2 , 0&lt;v&lt;1 ;</p>
        <p>– d is the parameter of loss elasticity when assigning rank 2 to z, whereas z∈∆1,
d&gt;0.</p>
        <p>
          Substituting (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) in (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ), we obtain the decision rule in the form:
z ∈ ∆1 if z &lt; ( d + ν )c /( d +1 ) , and z ∈ ∆2 if z ≥ (d + ν)c /(d + 1)
(
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
        </p>
        <p>
          Solve equation (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) using the method of stochastic approximation [14]. Denote ct
the optimal estimate of unknown parameter c in period t, obtained by this method.
Then, according to (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) we obtain a recursive equation for such estimate:
ct+1 = It (ct , zt ) → c* = argmin J (c)
t c
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
(9)
2 )
ct+1 = ct - γt ∑ Bk (ct , zt )∂Lk (ct , zt ) / ∂ct , c0 = b0 , z0 = g0 , t = 0,1,...
k =1
(7)
There b0 is the initial value of unknown parameter estimate, g0 is the initial value of
the efficiency potential, γt is the adaptation coefficient in period t,
∞
γt &gt; 0, ∑ γt &lt; ∞ [14]. Substituting (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ), (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) into (7), and taking into account (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ), we
t =0
obtain the algorithm for optimal estimate the parameter of the decision rule
ct+1 = It (ct , zt ) = {
ct + γtv if zt &lt; (d +v)ct /(d +1)
ct - γtd if zt ≥ (d +v)ct /(d +1)
, c0 = b0 , z0 = g0 , t = 0,1,... (8)
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Game-Theoretic Approach to Machine Learning of Dichotomy</title>
      <p>The elements of production are people, as well as the technologies and control
processes by which they carry out their activities. Therefore, when researching and
developing organizational and technological systems in the face of uncertainty, it is
necessary to take into account the socio-psychological aspects of production activities
including undesirable activity of staff.</p>
      <sec id="sec-4-1">
        <title>3.1 Stochastic Technological Active System</title>
        <p>Let us consider a two-level stochastic technological active system, at the upper level
of which Center is located, and at the lower level is forward-thinking staff (FTS) that
implements the technological process. Let us characterize game-theoretic approach to
the study of the interaction of Center and FTS to increase efficiency of the
technology, using the results of Section 2.</p>
        <p>Suppose that random value of technological potential zt becomes known FTS
before choosing indicator yt of technology efficiency in period t. Moreover, value zt
is unknown to Center. Based on the condition that the efficiency indicator cannot
exceed the potential (i.e. yt ≤ zt ), FTS chooses yt , t = 1, 2,..., in such a way as to
maximize own target function, depending on current and future ranks assigned by
Center.</p>
        <p>Center, on the other hand, observes only indicator yt that does not necessarily
coincide with potential zt (since yt ≤ zt ). Therefore, Center is forced to form the rank
of FTS under conditions of uncertainty caused not only by stochastic potential of the
technology, but also by the undesirable activity of FTS. For this, Center uses the
learning algorithm (8), substituting the observed indicator yt in it, instead of the
unknown zt . In this case, the estimate at +1 of the parameter of the decision rule is
obtained using algorithm similar to (8):</p>
        <p>at + γtv if yt &lt; (d + v)at /(d +1)
at+1 = It (at , yt ) = 
at - γtd if yt ≥ (d + v)at /(d +1)
, a0 = b0 , y0 = g0 , t = 0,1,... (10)</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2 Goals of Center</title>
        <p>Center is interested in unlocking the potential of a technology, as well as in increasing
its predictability. Constant disclosure of technology potential is achieved when
yt = zt , t = 0,1,... Unpredictability of this potential is due not only to the random
factors, but also to undesirable activity of FTS, due to which yt &lt; zt .</p>
        <p>Since in general case yt ≠ zt , then at ≠ ct , t = 1, 2,... Therefore, estimate at
calculated using recurrence algorithm (10) does not converge to optimal estimate c*
determined according to (9). This makes machine learning algorithm (10) ineffective.
Substantially, the reason is that Center is not able to take into account random factors
that become known to FTS in the process of production. This not only reduces the
effectiveness ( yt &lt; zt ), but also makes such machine learning inefficient. To improve
the efficiency of machine learning, it is necessary to ensure convergence of estimate
at calculated by recurrent algorithm (10) to optimal estimate c * determined
according to (9):
at +1 = It (at , yt ) → c* = arg min J (c)
t c
(11)</p>
        <p>Thus, Center’s goals are to unleash potential of technologies ( yt = zt , t = 1, 2,... ),
as well as to increase efficiency of machine learning by implementing (11). Thus,
expected payoff of Center is maximal if yt = zt , t = 1, 2,... To achieve this, Center
establishes the necessary order and mechanism for functioning of technological active
system.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3 Dichotomous Ranking Learning Mechanism</title>
        <p>Consider the following order of system’s functioning in period t, t = 0,1... In period
t = 0 Center knowing initial values a0 and y0 calculates estimate a1 for period 1 by
means of (10). Then Center reports a1 to FTS.</p>
        <p>In period t, t = 1, 2,..., FTS knows not only at , but also zt , zt ∈ ∆. Based on this,
FTS chooses indicator yt* that is preferable for itself, yt* ≤ zt . Then Center, based on
observation of yt* and known estimate at , determines rank FTS r t=R (a t , y *t). Also
Center calculates estimate at+1 for next period t+1 by means of (10). Then Center
reports at+1 to FTS, t = 1, 2,...</p>
        <p>We will call learning procedure a one-way infinite sequence of functions It (at , yt )
defined according to (10) with t = 0,1,..., and denote it</p>
        <p>I = { It (at , yt ), t = 0,1,... } (12)</p>
        <p>Similarly, we will call R = { R(at , yt ), t = 0,1,... } ranking procedure. Then learning
procedure I and ranking procedure R are combined into a dichotomous ranking
learning mechanism Σ = {I , R} in two-level technological active system shown in
fig.1.</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.4 Target of Forward-Thinking Staff</title>
        <p>Knowing at , zt , and Σ , FTS chooses yt in such a way as to increase own target
function Vt , which depends on current and future ranks r τ=R(aτ, yτ), τ= t ,t + T :
t+T
Vt (Σ) = ∑ρτ-t R(aτ , yτ )
τ=t
(13)
where ρ is the discount rate, 0 &lt; ρ &lt; 1, T is the number of periods taken into account
by FTS.</p>
        <p>at
zt</p>
        <p>Dichotomous ranking learning mechanism Σ = {I , R}</p>
        <p>Machine learning procedure I : at +1 = It (at , yt )</p>
        <p>a t
Ranking procedure R: rt = R(a t , y t )</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Rank rt</title>
      <sec id="sec-5-1">
        <title>Forward-thinking staff</title>
      </sec>
      <sec id="sec-5-2">
        <title>TECHNOLOGY</title>
        <p>yt</p>
        <p>We will assume that FTS know that 0 ≤ yτ ≤ zτ, and zτ ∈ ∆ in each future
period τ, τ = t +1,t + T. In order to make a choice in conditions of such uncertainty,
FTS is guided by expected payoff equals to guaranteed value of (13):
wt (at , yt , zt , Σ) =</p>
        <p>min max
zτ∈ ∆, τ =t+1,t+T 0≤ yτ ≤zτ , τ=t +1,t +T</p>
        <p>Vt (Σ) , t = 1, 2,...</p>
        <p>(14)</p>
        <p>Then the set of possible choices of FTS has the form:
W t (at, z t,Σ) ={ yt* | 0 ≤ y *t≤ zt, wt (at, y *t, zt,Σ) ≥ wt (at, yt, zt,Σ), 0≤ yt ≤ zt }, t = 1, 2,... (15)</p>
        <p>Suppose that the set of FTS possible choices (15) includes zt , i.e. zt ∈Wt (at, zt,Σ).
In this case we say that the benevolence hypothesis of FTS with respect to Center is
valid if FTS chooses yt*= zt, t = 1, 2,...</p>
      </sec>
      <sec id="sec-5-3">
        <title>3.5 Synthesis of Dichotomous Ranking Learning Mechanism</title>
        <p>Let us turn to current practice of production management. First, ranking and stimulus
procedures are usually designed in such a way stimuli grow as production indicators
increase compared with their current scores (plans, standards). Usually, stimulation is
carried out in case these scores are exceeded [4]. Therefore, the higher a score the
more difficult it is to get a stimulus.</p>
        <p>Secondly, forecasting procedure in a large corporation is often organized in such a
way that a score (like plan) in each subsequent period increases by a certain
percentage of result achieved today (the so-called “planning from the achieved level”)
[4]. Then future score (plan, standard) will be the higher, the higher today's indicators
are achieved. Therefore, staff may not be interested in exceeding the score (after all,
the higher the score in the future the more difficult it will be to get a stimulus).</p>
        <p>Thus, problem arises of the lack of interest of forward-thinking staff in unlocking
the potential of technology. In this case yt &lt; zt , and it is impossible to determine с*
with the aid of machine learning algorithm (10). Therefore, expected payoff of Center
is not maximal if yt*&lt; zt, t = 0,1,...</p>
        <p>Consider the non-cooperative game of Center and FTS in two-level technological
active system shown in fig.1. The first move in period t is made by Center, setting
mechanism Σ = {I , R} . The second move is made by FTS, choosing indicator yt*.
Then these moves are repeated in next period t+1, t = 1, 2,...</p>
        <p>Statement. To unlock the potential ( yt*= zt , t = 1, 2,... ) and improve the efficiency of
machine learning to obtain (11), it is enough to set a mechanism Σ = {I , R} with
procedure I that satisfies (12), and</p>
        <p>R(at , yt ) = Θ( yt - xt ),</p>
        <p>xt = at (d + v) /(d +1) ,
 2 if yt ≥ xt ,
Θ( yt - xt ) = 
 1 if yt &lt; xt
(16)
(17)
Proof. The expected payoff of FTS (14) depends according to (13) on both current
and future ranks rτ = R(aτ, yτ ) , τ = t ,t + T . By condition (16), with an increase in
indicator yt , current rank FTS rt = R(at , yt ) increases (does not decrease). In
addition, by the hypothesis of Statement, Center uses learning procedure (12). Hence,
according to (10) estimates аτ decrease (do not increase) with an increase in
yt , τ = t +1,t + T . Therefore, according to (16) future ranks FTS rτ = R(aτ, yτ )
increase (do not decrease) with an increase of yt , τ = t +1,t + T .</p>
        <p>According to (13), Vt (Σ) increases monotonically in rτ = R(aτ, yτ ) , τ = t ,t + T .
But rt monotonously increases (does not decrease) by yt . Therefore, with an increase
in indicator yt , expected payoff (14) also increases (does not decrease). Since
yt ≤ zt , maximum of wt (at , yt , zt , Σ) is reached at yt = zt . Therefore, according to
(15), zt ∈Wt (at, zt,Σ). Hence, by virtue of the benevolence hypothesis, yt*= zt ,
t = 0,1,... But then (8) and (10) coincide. Therefore, (11) follows from (9), Q.E.D.</p>
        <p>Following the common game-theoretic notation [15], let comment relation between
this Statement and Nash equilibrium. In our case, Nash equilibrium is a solution of
a non-cooperative game in which both Center and FTS know equilibrium strategies of
the other player, and no player has anything to gain by changing only own strategy.
Center strategy Σ determines actions based on what it has seen happen so far in the
game. Expected payoff of Center is maximal if yt*= zt , t = 1, 2,... FTS makes the
strategy choice yt —its own action based on Σ and zt . Expected payoff of FTS is
(14). According to Statement, no player can increase expected payoff by changing
strategy while the other player keep own strategy unchanged. Therefore the set of
strategy choices {Σ, yt* | t =1,2,...} constitutes Nash equilibrium. This unleashes
technology potential and increases efficiency of machine learning</p>
        <p>Consider a simple interpretation of Statement. Suppose that a performance
stimulation is such that the higher the rank the higher the stimulus for staff. Center
observes a value yt* characterizing actual effectiveness in period t, yt* ≤ zt , where zt
is unknown random maximal efficiency. Center learns tuning the decision rule
parameter with the aid of (10). Further, in accordance with adopted decision rule,
Center ranks FTS according to actual effectiveness of technology. Namely, with
yt* ≥ xt FTS refers to successful ( Θ = 2 ) and is encouraged. If yt* &lt; xt then FTS
refers to dysfunctional ( Θ = 1 ) and is punished.</p>
        <p>Any of these decisions is associated with a certain losses for Center. In first case,
losses L1 increase with a decrease in efficiency yt (undeserved encouragement or
bonus of staff). In second case, these losses L2 increase with increasing efficiency
and unfair punishment of staff. The standard xt = ( d + v )at /( d +1 ) corresponds to
the lower limit of satisfactory work of staff.</p>
        <p>Note that according to (10), the higher is technology efficiency indicator ( y*t ) the
lower is estimate for the next period (at+1). But, according to (17), this estimate plays
the role of the threshold value of indicator yt+1 at which FTS receives a stimulus in
period t+1. Therefore, FTS becomes easier to get a stimulus in period t+1 even with a
smaller value of random potential zt.</p>
        <p>In other words, with an increase in indicator yt*, staff receives not only a higher
stimulus. With growth y*, estimate for the next period at+1 decreases. Therefore
t
threshold value for stimulation in the future decreases. This further interests the staff
in unlocking the potential of the technology, i.e. in choosing yt* = zt . Thus, in
accordance with (8) – (10) learning procedure (12) ensures convergence of estimate
at to optimal value c* (11). This makes machine learning algorithm (10) more
efficient.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4 Example: Using 2 dichotomous ranking learning mechanisms for ranking electricity efficiency in four ranks</title>
      <p>By Statement, dichotomous ranking learning mechanism promotes revelation of
potential effectiveness of technology and increase efficiency of machine learning.
Consider the application of this mechanism to ranking the efficiency of energy-saving
technology under the program of increasing the energy efficiency of the Russian
Railways holding.</p>
      <sec id="sec-6-1">
        <title>4.1 Ranking Learning Mechanism</title>
        <p>Improving energy efficiency of production technology can be achieved through
modernization and optimization of technological processes, as well as operating
modes of heating and lighting systems of an enterprise. The decision on ranking the
efficiency of energy-saving technology is made by the employee of the energy
commission of the Russian Railways, responsible for achieving indicators of the
program for improving energy efficiency in region. This employee acts as Center,
observing actual effectiveness of energy-saving technology at a subordinate enterprise
providing wagon-repairing [16]. Consider the application of a mechanism Σ = {I , R}
that satisfies conditions of Statement to improve efficiency of electricity use at such
enterprise (briefly, FTS).</p>
        <p>The monthly electricity efficiency of wagon-repairing αt is calculated as number
nt of wagons repaired in month t, divided by appropriate number et of
megawattmonth of consumed electricity: αt =nt /et. There t is the number of month in a year,
t = 1,12. Indicator yt of monthly electricity efficiency is calculated as deviation of αt
from the norm (plan) of electricity efficiency βt divided by this norm: yt = (αt -βt )/βt
t = 1,12. Thus, the value of indicator yt depends on variables nt, et, βt. In turn, these
variables depend on many random factors: volume of orders for repairs, weather,
season, etc. Therefore value yt has a stochastic character.</p>
        <p>Based on indicator yt, a monthly ranking of technology effectiveness is determined.
For this, Center uses decision rule with customizable estimates based on algorithm
(10) and conditions of Statement. In order to make the results more vivid, Center
defines four ranks of electricity efficiency. There rank 4 corresponds to excellent
assessment of electricity efficiency, rank 3 – to good assessment, rank 2 – to
satisfactory assessment, and rank 1 – to poor assessment.</p>
        <p>To determine these four ranks of electricity efficiency, Center uses the following
procedure for estimate parameters of decision rule and procedure for ranking, based
on the approach developed above.</p>
        <p>1. In case indicator yt is negative ( yt &lt; 0 ) then rank rt of electricity efficiency in
period t can be 1 or 2. At the beginning of the year (in period t = 0 ), Center and FTS
know d, v, initial values of indicator y0 , estimate a0 , and standard
x0 = a0 (d + v) /(d +1). After that, within a year the estimate of decision rule parameter
is adjusted according to algorithm (10), by formula</p>
        <p>
aτ+1 =  aτ + γ τv if yτ &lt; xτ ,
aτ - γ τd if yτ ≥ xτ
xτ = aτ (d + v) /(d +1) , τ = 0,12,
(18)
Then according to Statement, the rank of electricity efficiency is</p>
        <p>
           2 if xt ≤ yt &lt; 0
rt = 
 1 if yt &lt; xt
,
t = 0,12
(19)
2. In case indicator yt is non-negative ( yt ≥ 0 ) then the rank of electricity
efficiency may be 3 or 4. In this case, procedures for parameter estimates of decision
rule and ranking are based on the approach developed above. Namely, similar to (
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
denote: at+ – the estimate of parameter of decision rule for ranking of electricity
efficiency to 3 or 4 in period t; yt − va+ – the loss function for erroneous assignment
of rank 3 (instead of rank 4), 0&lt;v+&lt;1; d + (a+ − yt ) – the loss function for erroneous
assignment of rank 4, d+&gt;0; xt+ = at+ (d + + v+ ) /(d + +1); t = 0,12.
        </p>
        <p>At the beginning of the year (in period t = 0 ) Center and FTS know d +, v+,
initial values of indicator y0 , estimate a0 , and standard x0+ = a0+ (d + + v+ ) /(d + +1).</p>
        <p>+
After that, within a year the estimate of decision rule parameter is adjusted similar to
(18) by formula:
aτ++1 =  aτ+ + γ τv + if yτ &lt; xτ+ , xτ+ = aτ+ (d + + v+ ) /(d + +1) , τ = 0,12
aτ+ - γ τd + if yτ ≥ xτ+</p>
        <p>Then in accordance with Statement, the rank of electricity efficiency is calculated
using a formula similar to (19):
rt+ =  4 if xt+ ≤ yt
3 if 0 ≤ yt &lt; xt
+ ,
t = 0,12
(20)
(21)
Note that (20) and (21) are similar to (18) and (19), if at , xt , rt replaced by
at+ , xt+, rt+ , respectively.</p>
      </sec>
      <sec id="sec-6-2">
        <title>4.2 Parameters of Calculations</title>
        <p>Estimates at and at+ in period t were calculated, respectively, using formulas (18)
and (20). Based on these estimates, values of xt and xt+ were determined according
to (18) and (20). Values xt and xt+ make sense, respectively, of the lower and the
upper standard of electricity efficiency in period t, t = 0,12.</p>
        <p>The values of parameters used in calculations of at and at+ were based on hard
and soft knowledge. Information theory of identification [14] shows that in optimal
algorithm γt = 1/(t +1), t = 0,1,... The other needed values were assigned by expert –
invited specialist or tutor [12]. Then these values were adjusted by the employee
himself or a manager of a higher level.</p>
        <p>In this expert environment, it was generally accepted that the loss in case of
mistaken assignment of a higher rank was lower than the loss in case of erroneous
assignment of a lower rank. Based on this, parameters of loss elasticity in case of an
erroneous assignment of a higher rank were taken equal to 0.15 (i.e. d = d + = 0,15).
In case of an erroneous assignment of a lower rank, parameters of loss elasticity were
taken equal to 0.10 (i.e. v = v+ = 0,10).</p>
        <p>In addition, experts a priori considered the deviation of indicator yt from zero level
by more than 5% noticeable. The result of such a deviation should be an assignment
to a different rank, and appropriately stimulated. Therefore, the initial values of
standards of electricity efficiency were taken equal x0 = −0,05, x0+ = 0,05. Hence,
according to (19) if y0 &lt; x0 = −0,05 then FTS attributed to a rank 1. Also according
to (21) if y0 ≥ x0+ = 0,05, then FTS attributed to a rank 4.</p>
      </sec>
      <sec id="sec-6-3">
        <title>4.3 Estimations and Standards Calculations</title>
        <p>Note if x0 = −0,05, x0+ = 0,05, then according to (18) and (20), a0 = −4,6,
a0+ = 4,6. The results of estimations and standards calculations for above parameters
are shown in fig.2. There are graphs of indicator yt, as well as lower and upper
standard xt and xt+ during the year, t = 0,12.</p>
        <p>According to (20), if indicator yt is less than the upper standard xt+ then this
standard is increased. For example, according to fig.2 the upper standard xt+ increases
in February and March. Also from September until the end of the year, the standard
xt+ was raised because yt was less than the standard (here the growth of the lower
standard xt+ is less noticeable on fig. 2 due to the smallness of the γt). On the other
hand, according to (21), the standard xt+ is a threshold of excellent rank in the future
period τ, τ&gt;t. Therefore, even in November, the FTS did not receive an excellent rank,
although indicator yt exceeded 0.1. Note that before September this would have been
enough to get an excellent rank.</p>
        <p>Also from fig.2 we see growth of lower standard xt during the second month of
the year (in February). The reason is the low FTS indicator in January. This is
explained by the fact that according to (18), the smaller is indicator yt the higher is
the lower standard xτ using as a threshold for satisfactory rank in future period τ, τ&gt;t.
In other words, by being poor FTS worsened own ability to become satisfactory rank.
The situation is similar in the ninth month (here the growth of the lower standard xt
is also less noticeable on fig. 2 due to the smallness of the γt).</p>
        <p>However, it was enough for FTS not to get poor rank in the following months
(including August) as lower standard xt began to decline. In essence, from February
to August FTS worked for own authority.</p>
      </sec>
      <sec id="sec-6-4">
        <title>4.4 Ranks and Stimuli</title>
        <p>Further, Center determined rank Rt in period t, t = 0,12. If current indicator of
electricity efficiency was negative (yt&lt;0) then to assign rank 1 or 2, (19) was used.
Otherwise (at yt≥0) to assign rank 3 or 4, (21) was used. Thus, the procedure for
ranking electricity efficiency in period t, combining (19) and (21), had the form:



rˆt = 



4
3
if
if
2 if
where rˆt is the rank of electricity efficiency in month t. In fig.3 shows a graph of rank
rˆt calculated by (22), t = 0,12.</p>
        <p>Substantially, standard xt+ is the lower limit of electricity efficiency yt ,
corresponding to excellent work of staff. Standard xt is the lower limit of electricity
efficiency yt corresponding to satisfactory performance. When electricity efficiency
was below xt in February and October, the intervention of Center was required.</p>
        <p>Let comment on the relation between fig. 2 and fig. 3. Note that in general terms
dynamics of rˆt resembles dynamics of yt. However fig. 3 gives rougher assessments.
This was necessary for a qualitative analysis (for example in a report to the higher
management when there was no time to go into details). Whereas fig. 2 provides more
accurate quantitative estimates of the ratio of indicators and standards useful for
indepth analysis. For example fig. 3 shows that in November (t =11) the rank was 3.
However, fig. 2 shows that this month indicator yt was very close to upper standard,
that is, could get rank 4. This shows more progress in increasing electricity efficiency
at the end of the year.
for the next period xt+1 and xt++1. Thus, with an increase in indicator yt the staff
receives not only higher stimulus st . Also threshold value for future stimulation xt+1
+
and xt +1 will decrease. This further piques the interest of staff in maximum disclosure
of electricity efficiency potential and makes machine learning more efficient.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5 Conclusions</title>
      <p>An important aspect of development and application of game theory and stochastic
modeling in a wide range of organizational and technical systems is study of
possibility of use of internal reserves and resources of the technologies used. For this,
it is necessary to take into account human factor, interests of elements of the system.
Often management does not know random interference and other factors that become
known to staff in the process of production. This reduces effectiveness of the
technology and makes the machine learning inefficient.</p>
      <p>The task of the theory was the synthesis of game-theoretic mechanisms of
coordinated control in stochastic conditions, in which desire of elements to achieve
own interests leads to an increase in technology effectiveness. The method for solving
this problem by synthesizing the optimal learning mechanism for dichotomous
ranking was proposed. This method includes construction of procedures for stochastic
approximation of decision rule, ranking and stimulation. Thanks to this mechanism,
Nash equilibrium arises which makes machine learning more efficient.</p>
      <p>This approach was illustrated by the example of machine learning to rank electrical
efficiency of wagon-repairing aimed at achieving targets established by the program
of increasing the energy efficiency of the Russian Railways. Further research will
focus on the development of theoretical, algorithmic and methodological aspects of
game theory and stochastic modeling with a focus on their application in a wide range
of active organizational and technical systems.
7. Fradkov, A.: Early history of machine learning. In: 21st IFAC World Congress, 3439.</p>
      <p>Berlin (2020).
8. Bristow, D., Tharayil, M., Alleyne, A.: A survey of iterative learning control. IEEE</p>
      <p>
        Control Systems Magazine 26(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ), 96-114 (2006).
9. Recht, B.: A tour of reinforcement learning: the view from continuous control. In:
arXiv:1806.09460v2 [math.OC] 10 November 2018, https://arxiv.org/pdf/1806.09460.pdf,
last accessed 2020/08/06.
10. Recht, B.: Reflections on the learning-to-control renaissance. In: 21st IFAC World
      </p>
      <p>Congress, 4707. Berlin (2020).
11. Tsyganov, V.: Learning mechanisms in digital control of large-scale industrial systems. In:</p>
      <p>Global Smart Industry Conference, pp. 1–5. IEEE, Chelyabinsk, Russia (2018).
12. Tsyganov, V.: Tutoring mechanisms of business management. In: 21st International</p>
      <p>Conference on Business Informatics, vol. 2, pp. 60-67. IEEE, Moscow (2019).
13. Tsyganov, V.: Designing adaptive information models for production management.</p>
      <p>Procedia CIRP 84, 1088-1093 (2019).
14. Tsypkin, Ya.: Fundamentals of the information theory of identification. Nauka, Мoscow
(1984).
15. Osborne, M., Rubinstein, A.: Course in game theory. MIT, Cambridge MA (1994).
16. Tsyganov, V.: Decision making and learning in wagon-repairing. In: 12th Conference on
Management of Large-Scale System Development. pp. 1–5. IEEE, Moscow (2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>Y.-C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muralidharan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Information structure, Stackelberg games, and incentive controllability</article-title>
          .
          <source>IEEE Trans. Automat. Control</source>
          <volume>26</volume>
          (
          <issue>2</issue>
          ),
          <fpage>454</fpage>
          -
          <lpage>460</lpage>
          (
          <year>1981</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Germeier</surname>
          </string-name>
          , Yu.:
          <article-title>Igry s neprotivopolozhnymi interesami (Games with non-opposing interests</article-title>
          ), Nauka, Moscow (
          <year>1976</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Mechanism theory</article-title>
          . In: Derigs U. (ed.)
          <article-title>Optimization and Operations Research, Encyclopedia of Life Support Systems</article-title>
          . EOLSS Publishers, Oxford (
          <year>2003</year>
          ), https://web.stanford.edu/~jacksonm/mechtheo.pdf,
          <source>last accessed</source>
          <year>2020</year>
          /08/06.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Burkov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gubko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondratiev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korgin</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novikov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Mechanism design and management. Mathematical methods for smart organizations</article-title>
          . NOVA Publishers, New York (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Van</given-names>
            <surname>Essen</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>A note on the stability of Chen's Lindahl mechanism</article-title>
          .
          <source>Social Choice and Welfare</source>
          <volume>38</volume>
          (
          <issue>2</issue>
          ),
          <fpage>365</fpage>
          -
          <lpage>370</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Arifovic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ledyard</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A behavioral model for mechanism design: individual evolutionary learning</article-title>
          .
          <source>Journal of Economic Behavior and Organization</source>
          <volume>78</volume>
          ,
          <fpage>375</fpage>
          -
          <lpage>395</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>