<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Electronic training polygon for artificial intelligence systems</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>A.A. Karandeev</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Karandeev Alexander A., phd student and Junior Scientific Associate of the Keldysh Institute of Applied Mathematics, Russian Academy of Sciences</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Keldysh Institute of Applied Mathematics, Russian Academy of Sciences</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Plekhanov Russian University of Economics</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the framework of the concept of neoconflictology, the possibilities and methods of mathematical modelling of conflict dynamics under uncertainty are considered. Complex contradictory situations, when it is necessary to quickly respond to changes in the situation and perform some actions in conditions of uncertainty, are often found in different spheres of activity. In this case, the uncertainty may be due to incomplete knowledge of the situation, the inability to quickly understand and evaluate possible options for its development, the influence on it from other participants in the process, etc. The result of further developments depends on how the decision will be to the situation and what actions will follow on its basis. Usually, the effectiveness of behavior in difficult situations is determined by the level of special training and experience of the decision maker. The trend of using artificial intelligence systems to support decisionmaking, including complex conflict situations with a high level of uncertainty, is now more frequent. Focusing on the use of such systems requires the creation and improvement of specialized tools and technologies, thanks to which the artificial intelligence systems themselves can form an information base, which is a necessary factor for developing rational solutions in a complex environment. The creation of a virtual polygons is one of the ways for producing such information base.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction</p>
      <p>In an antagonistic conflict, the effectiveness of
achieving goals depends on the adopted strategy of
behaviour, the speed of decision-making in a changing
environment and the level of their optimality, or in other
words, the degree of their compliance with the current
situation.</p>
      <p>
        The best way to analyze situations of this kind that
demonstrate emerging phenomena or generate unforeseen
patterns is to model and simulate them [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
problematic complexity of solving these issues is directly
related to the level of uncertainty, which is often due to a
lack of time to obtain the information necessary for
decision-making in a conflict situation and many other
factors.
      </p>
      <p>Given the modern development of computer methods
and tools, it is numerical modelling that assumes the role
of the tool by which you can explore a variety of conflicts
to identify the role and significance of certain behavioural
strategies, determine the list of effective actions in various
situations. One of the methods that greatly simplify the
modelling process is the creation of an electronic polygon.
Despite the complexity of developing such kind of
software systems, they make it possible to visually
simulate various situations and see what these or those
solutions lead to.</p>
      <p>
        The main participants in a computer polygon are
intelligent agents [
        <xref ref-type="bibr" rid="ref2 ref3">2-3</xref>
        ] who are trying to achieve their
goals. They can have both common goals and completely
opposite. It is worth considering that for each of the
participants in the conflict, their own space of probable
states can be formed, due to its capabilities and ideas about
the "world in which it operates." Moreover, the state of the
agent can be changed not only due to the actions taken by
him, but also under the influence of other agents who are
most often the adversaries. The agent’s task is to construct
an optimal trajectory for transition from its initial state to
a given target state.
      </p>
      <p>Usually, the set of actions that an agent can perform is
limited, but each of these actions must be taken into
account not only affects its own state, but also the state of
other agents in their phase space. The purpose of these
repeated experiments is to train the agent for rational
action in the given conditions. To achieve this effect, a
mathematical model is designed.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Basic provisions of the model approach</title>
      <p>
        The developed research technology is based on
numerical modelling of possible ways to actualize an
antagonistic conflict based on the following
considerations. In the simulation model, each of the
subjects of the conflict can be represented as an
“intelligent” agent. The theory of agents or multi-agent
systems [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a computer theory which seeks to apprehend
the coordination of competing independent process. An
agent is thus a computer process [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which can be
considered as autonomous since it is capable of adapting
when its environment changes. The basic characteristics of
such an agent are the formation of an internal model of the
surrounding world and the presentation of its place in it, a
system of rules for creating and changing this model,
decision-making methods for taking actions to respond
adequately in response to a changing situation. Intelligent
agents have targeted behaviour to achieve a certain goal in
the most optimal way.
      </p>
      <p>
        This process comes down to setting optimization
problems and finding their solutions. Based on the
described representations, each of the subjects of the
confrontation can be represented in the form of a complex
system that has certain resources and has its own goals. It
is worth considering that many types of conflicts, such as
social and political, cannot be strictly defined [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Because
of this, it may be necessary to discard some non-significant
parameters.
      </p>
      <p>The subject's capabilities depend on his position in the
phase state space at each current point in time, which is
characterized by a certain set of particular characteristics.
Otherwise, a change in its state over time depends on the
position of the subject in the phase space of his states,
which also include the physical coordinates of his location.
Moreover, each subject has its own phase space, only
partial intersection with the opponent’s space is possible.
The transition of the subject from the initial state to the
target is carried out through a set of intermediate states, the
trajectory of motion in phase space is a result of their
change.</p>
      <p>Conventionally, the reflection of the model of the
intelligent agent in the form of a graph, according to
[710], can be displayed in the form of the circuit shown in
Fig. 1. The initial position is conditionally shown in a blue
circle, and the target in yellow. Choosing a path from one
state to another involves a certain sequence of actions
(steps), which is determined by the subject's ideas about
the cost of resources for each of them and an assessment
of their proximity to the target state. The diagram shows
that each individual action is associated with a certain
resource of the real world, the availability of which (in the
representation of the object) is a factor determining the
possibility of its implementation.</p>
      <p>In other words, any change in the parameter that
describes some of the characteristics of the subject’s
position in the phase space can be associated with a change
in some scalar quantity — the transition price, which is
understood as the conditional cost of the action. Each of
the parameters has its own “price scale”, which
dynamically depends on the current state of the subject and
is determined by the value of this parameter to solve the
problem.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Adapting rules of conduct</title>
      <p>In the phase space or in the space of states we define
some arbitrary point, which will be considered as the
initial position of the subject. We also set the second point
denoting the position of the target. The state of the subject
changes as a result of some elementary actions in the state
space, while the set of possible actions is limited and is
described by a set of rules of behavior, each of which is
displayed by a vector in the phase space. The sequential
step-by-step execution of actions forms a certain trajectory
of the subject in the state space. In a two-dimensional
interpretation, the task is shown schematically in Fig. 2.</p>
      <p>The shortest direction or trajectory from the source to
the target state is shown by a dashed line. It is obvious that
the result of the implementation of one or another rule of
behavior, which is displayed in the phase space by some
action vector, is unlikely to coincide with the shortest
direction. Clearly, a rational trajectory of movement
should be as close to this line as possible.</p>
      <p>But in conditions of inaccurate information about the
properties of each of the vectors, it is difficult to construct.
For the two-dimensional problem shown in figure 2, such
an operation can be done “manually”, based on visual
estimates. One of the possible results may be, for example,
the path options shown in Fig. 3.
In the presented schematic example, based on a visual
assessment of the set of action vectors, such trajectories
can be constructed “manually”. There are at least two
rational options for the trajectories of the transition from
the initial state to the target, and it is enough to use a
combination of only two of all possible actions. When
comparing the obtained trajectories, it is seen that that the
upper one is slightly more preferable in terms of the degree
of approximation to the target position, but its
implementation requires one more step. To determine the
best of them, additional evaluation criteria are needed, for
example, a comparison of criteria for accuracy and
resource costs. Based on the above examples, it can be
argued that in a multidimensional phase space, the
construction of the optimal path is possible only with full
knowledge of the situation: where, by what means, with
what intermediate results certain actions arising from the
rules of behaviour of agents can be realized.</p>
      <p>The following scheme was proposed and tested.
Multidimensional phase space and an unordered set of
permissible actions are set. At the same time, the number
of actions and the dimension of space are determined by
the conditions of a specific task. It is assumed that the basis
for decision-making by the agent in the current situation to
choose a specific action is some a priori assessment of the
value of certain actions. Initially, it can be set by experts
or randomly generated.</p>
      <p>Further refinement of the assessment of the actions that
the agent must carry out is the result of a wide series of
numerical experiments, in each of which the trajectory of
motion of the agent states in the phase space from the
source to the target is constructed. In the process of
performing calculations, a certain rating is assigned to
each perfect action. The rating of each action is determined
by the degree of approximation of the new state to the
target state obtained as a result of this action from this
position. With a positive effect (approaching the goal),
incentive is introduced, with a negative effect (moving
away from the goal) - a penalty. At the beginning of the
calculation, all ratings are set close to zero.</p>
      <p>The next move step is randomly selected from the list
of all possible actions, but taking into account their rating.
The probability of choosing an action increases in
proportion to the rating. As a result of the calculations, the
initially uniform distribution of the rating of actions is
gradually deformed towards growing ratings. At each
iteration, the current state in which the subject was located
is remembered. Thus, a model of stereotypical situations
is formed in which the effects of the same actions can
manifest themselves in different ways. Each iteration stops
after reaching the target state or after completing a given
number of steps. At each subsequent iteration, the ratings
accumulated in previous experiments are saved and used
when choosing actions. Therefore, the knowledge base is
accumulated and the conditional agent is “trained” in the
rational choice of actions depending on the position in the
state space in which it is located.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>Testing of the methodology was carried out on an
arbitrary example in a space of 12 measurements with a
total number of rules of behavior equal to 50 (Fig. 4).</p>
      <p>The fig. 4 shows 33 of 50 actions in a 12-dimensional
state space. The positions of the initial and target states
were randomly selected as vectors in the phase space. In
addition, in the form of increments of vectors, action
vectors were arbitrarily specified.</p>
      <p>Let us dwell in more detail on the data given in the
table. As already mentioned above, the agent’s training
technology is demonstrated using a hypothetical example,
in which a certain zone of the phase space is
conventionally defined as a range of changes in the
coordinate values of the parameters, and the initial and
target position of the agent are determined inside this zone.
Variants of the agent’s possible actions are randomly
generated by coordinatewise specifying a certain set of
increment vectors. When generating incremental vectors,
the only condition is to limit the length of each of these
vectors to a value significantly less than the distance
between the initial and target position of the agent in the
phase space.</p>
      <p>The data given in the table reflect the obtained picture
of the learning outcomes at some intermediate stage of the
experiments. For clarity, a visual analysis introduced the
color of the cells. In this case, color means the directivity
of the effect, and the intensity of the color is its immediate
meaning. In other words, the cells highlighted in red
indicate that the choice of this action will lead to the loss
of a parameter or resource. Moreover, the more intense the
color, the greater the loss. Positive effects highlighted in
green.</p>
      <p>Fig. 5 shows the result of a computational procedure
for finding a rational trajectory.</p>
      <p>The process of qualitatively improving the procedure
for choosing rational actions from a given set is visible in
the fig. 5. The x-axis shows the number of experiments, in
each of which a motion pattern limited in the number of
steps is plotted in the phase space, and the y-axis is the
distance from the target state in the adopted metric. It can
be seen that the agent training process leads to the
achievement of some non-improved level of performance,
which, however, allows to get close enough to the target
state.
5. Conclusions</p>
      <p>Proposed technological solution for modeling the
behavior of intelligent agents provides a fairly broad basis
for studying the features of constructing trajectories of
achieving goals in phase space and the adaptive behavior
of agents in a conflict situation.</p>
      <p>Due to the specifics of the problem, approaches and
methods for modeling the phase space of states of a
conflict environment are proposed, which allow
determining strategies for rational trajectories of goal
achievement. The technology for modeling agent behavior
is based on a detailed description of their behavior based
on a scheme of intelligent transitions between states and
behavior models.</p>
      <p>Such a detailed description of complex behavior
models makes it possible to uniformly display in the model
of a conflict environment the essential aspects of the
behavior of real-world prototype objects. The proposed
scheme for ensuring the adaptive behavior of agents in a
conflict environment is an integral part of the
methodological and algorithmic support of an electronic
training ground for studying the conflict interaction of
complex systems.</p>
      <p>According to the authors, this technology has shown its
effectiveness in test tasks and can be recommended for use
as part of an electronic training polygon to test its
capabilities in developing strategies for agents' behavior
under conditions of uncertainty during their interaction.</p>
      <p>A similar technology can be applied in automated
process control support systems, where an operational
assessment of the situation is necessary based on the
available incomplete or inaccurate data with the
development of recommendations for rational actions. In
particular, when creating combat control systems,
ensuring the safety of facilities, responding in emergency
situations, as well as in control systems for unmanned
vehicles.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Klügl</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oechslein</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Puppe</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dornhaus</surname>
            <given-names>A.</given-names>
          </string-name>
          , MultiAgent Modelling in Comparison to Standard Modelling,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Nechaev</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osipov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetverushkin</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baluta</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <article-title>Ontological synthesis of management decisions in conditions of antagonistic conflicts</article-title>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Baluta</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nechaev</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osipov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetverushkin</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <article-title>Conceptual base of a supercomputer platform of application-oriented simulation, prediction and expertizes of conflict interaction</article-title>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bousquet</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barreteau</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mullon</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            <given-names>J.,</given-names>
          </string-name>
          <article-title>An environmental modelling approach. The use of multi-agent simulations</article-title>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Ferber</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <article-title>La kénétique: dés systems multi-agents à une science</article-title>
          de l'interaction,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Petukhov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Мalhanov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandalov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petukhov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <article-title>Mathematical Modeling of the Ethno-social Conflicts by Non-linear Dynamics</article-title>
          .
          <source>Proceedings of the 7th International Conference on Simulation and Modeling Methodologies, Technologies and Applications</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kh. Khakimova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.V.</given-names>
            <surname>Zolotarev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Berberova</surname>
          </string-name>
          .
          <article-title>Visualization of bibliometric networks of scientific publications on the study of the human factor in the operation of nuclear power plants based on the bibliographic database Dimensions</article-title>
          .
          <source>Scientific Visualization</source>
          ,
          <year>2020</year>
          , volume
          <volume>12</volume>
          , number 2, pages
          <fpage>127</fpage>
          -
          <lpage>138</lpage>
          , DOI: 10.26583/sv.12.2.10, E-ISSN:
          <fpage>2079</fpage>
          -
          <lpage>3537</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Berberova</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.I.Chernyavskii</surname>
          </string-name>
          , «
          <article-title>Comparative assessment of the NPP risk (on the example of Rostov and Kalinin NPP)</article-title>
          .
          <article-title>Development of risk indicators atlas for Russian NPPs»</article-title>
          ,
          <source>GraphiCon 2019 Computer Graphics and Vision</source>
          .
          <source>The 29th International Conference on Computer Graphics and Vision</source>
          . Conference Proceedings (
          <year>2019</year>
          ), Bryansk, Russia,
          <source>September 23-26</source>
          ,
          <year>2019</year>
          , Vol-
          <volume>2485</volume>
          , urn:nbn:de:
          <fpage>0074</fpage>
          -
          <lpage>2485</lpage>
          -1, ISSN 1613-0073, DOI: 10.30987/graphicon2019-2-
          <fpage>290</fpage>
          -294, http://ceur-ws.org/Vol2485/paper67.pdf, p.
          <fpage>290</fpage>
          -
          <lpage>294</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Berberova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.S.</given-names>
            <surname>Oboimov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kh.Khakimova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.V.</given-names>
            <surname>Zolotarev</surname>
          </string-name>
          , «
          <article-title>Risk-informed security system. The use of surveillance cameras for the particularly hazardous facilities safety»</article-title>
          ,
          <source>GraphiCon 2019 Computer Graphics and Vision</source>
          .
          <source>The 29th International Conference on Computer Graphics and Vision</source>
          . Conference Proceedings (
          <year>2019</year>
          ), Bryansk, Russia,
          <source>September 23-26</source>
          ,
          <year>2019</year>
          , Vol-
          <volume>2485</volume>
          , urn:nbn:de:
          <fpage>0074</fpage>
          -
          <lpage>2485</lpage>
          -1, ISSN 1613-0073, DOI: 10.30987/graphicon2019-2-
          <fpage>321</fpage>
          -325, http://ceur-ws.org/Vol2485/paper74.pdf, p.
          <fpage>321</fpage>
          -
          <lpage>325</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Baluta</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osipov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetverushkin</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yakovenko</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <article-title>Conceptual issues of model representation of conflicts</article-title>
          .
          <source>Proceedings of the International Scientific Conference</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>