=Paper= {{Paper |id=Vol-2763/CPT2020_paper_s4-1 |storemode=property |title=Electronic Training Polygon for Artificial Intelligence Systems |pdfUrl=https://ceur-ws.org/Vol-2763/CPT2020_paper_s4-1.pdf |volume=Vol-2763 |authors=Alexander Karandeev,Victor Baluta,Vladimir Osipov }} ==Electronic Training Polygon for Artificial Intelligence Systems== https://ceur-ws.org/Vol-2763/CPT2020_paper_s4-1.pdf
            Electronic training polygon for artificial intelligence systems
                                         A.A. Karandeev1,2, V.I. Baluta1,2, V.P. Osipov1
                              karalex755@gmail.com | vbaluta@keldysh.ru | osipov@keldysh.ru
                 1
                   Keldysh Institute of Applied Mathematics, Russian Academy of Sciences, Moscow, Russia
                                         2
                                          Plekhanov Russian University of Economics

    In the framework of the concept of neoconflictology, the possibilities and methods of mathematical modelling of conflict dynamics
under uncertainty are considered. Complex contradictory situations, when it is necessary to quickly respond to changes in the situation
and perform some actions in conditions of uncertainty, are often found in different spheres of activity. In this case, the uncertainty may
be due to incomplete knowledge of the situation, the inability to quickly understand and evaluate possible options for its development,
the influence on it from other participants in the process, etc. The result of further developments depends on how the decision will be to
the situation and what actions will follow on its basis. Usually, the effectiveness of behavior in difficult situations is determined by the
level of special training and experience of the decision maker. The trend of using artificial intelligence systems to support decision-
making, including complex conflict situations with a high level of uncertainty, is now more frequent. Focusing on the use of such systems
requires the creation and improvement of specialized tools and technologies, thanks to which the artificial intelligence systems
themselves can form an information base, which is a necessary factor for developing rational solutions in a complex environment. The
creation of a virtual polygons is one of the ways for producing such information base.
    Keywords: neoconflictology, conflict, training polygon, intelligent agent, Artificial intelligence, modelling.

                                                                         other agents in their phase space. The purpose of these
1. Introduction                                                          repeated experiments is to train the agent for rational
    In an antagonistic conflict, the effectiveness of                    action in the given conditions. To achieve this effect, a
achieving goals depends on the adopted strategy of                       mathematical model is designed.
behaviour, the speed of decision-making in a changing
                                                                         2. Basic provisions of the model approach
environment and the level of their optimality, or in other
words, the degree of their compliance with the current                       The developed research technology is based on
situation.                                                               numerical modelling of possible ways to actualize an
    The best way to analyze situations of this kind that                 antagonistic conflict based on the following
demonstrate emerging phenomena or generate unforeseen                    considerations. In the simulation model, each of the
patterns is to model and simulate them [1]. The                          subjects of the conflict can be represented as an
problematic complexity of solving these issues is directly               “intelligent” agent. The theory of agents or multi-agent
related to the level of uncertainty, which is often due to a             systems [4] is a computer theory which seeks to apprehend
lack of time to obtain the information necessary for                     the coordination of competing independent process. An
decision-making in a conflict situation and many other                   agent is thus a computer process [5], which can be
factors.                                                                 considered as autonomous since it is capable of adapting
    Given the modern development of computer methods                     when its environment changes. The basic characteristics of
and tools, it is numerical modelling that assumes the role               such an agent are the formation of an internal model of the
of the tool by which you can explore a variety of conflicts              surrounding world and the presentation of its place in it, a
to identify the role and significance of certain behavioural             system of rules for creating and changing this model,
strategies, determine the list of effective actions in various           decision-making methods for taking actions to respond
situations. One of the methods that greatly simplify the                 adequately in response to a changing situation. Intelligent
modelling process is the creation of an electronic polygon.              agents have targeted behaviour to achieve a certain goal in
Despite the complexity of developing such kind of                        the most optimal way.
software systems, they make it possible to visually                          This process comes down to setting optimization
simulate various situations and see what these or those                  problems and finding their solutions. Based on the
solutions lead to.                                                       described representations, each of the subjects of the
    The main participants in a computer polygon are                      confrontation can be represented in the form of a complex
intelligent agents [2-3] who are trying to achieve their                 system that has certain resources and has its own goals. It
goals. They can have both common goals and completely                    is worth considering that many types of conflicts, such as
opposite. It is worth considering that for each of the                   social and political, cannot be strictly defined [6]. Because
participants in the conflict, their own space of probable                of this, it may be necessary to discard some non-significant
states can be formed, due to its capabilities and ideas about            parameters.
the "world in which it operates." Moreover, the state of the                 The subject's capabilities depend on his position in the
agent can be changed not only due to the actions taken by                phase state space at each current point in time, which is
him, but also under the influence of other agents who are                characterized by a certain set of particular characteristics.
most often the adversaries. The agent’s task is to construct             Otherwise, a change in its state over time depends on the
an optimal trajectory for transition from its initial state to           position of the subject in the phase space of his states,
a given target state.                                                    which also include the physical coordinates of his location.
    Usually, the set of actions that an agent can perform is             Moreover, each subject has its own phase space, only
limited, but each of these actions must be taken into                    partial intersection with the opponent’s space is possible.
account not only affects its own state, but also the state of            The transition of the subject from the initial state to the

Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY
4.0)
target is carried out through a set of intermediate states, the      state to another involves a certain sequence of actions
trajectory of motion in phase space is a result of their             (steps), which is determined by the subject's ideas about
change.                                                              the cost of resources for each of them and an assessment
    Conventionally, the reflection of the model of the               of their proximity to the target state. The diagram shows
intelligent agent in the form of a graph, according to [7-           that each individual action is associated with a certain
10], can be displayed in the form of the circuit shown in            resource of the real world, the availability of which (in the
Fig. 1. The initial position is conditionally shown in a blue        representation of the object) is a factor determining the
circle, and the target in yellow. Choosing a path from one           possibility of its implementation.




                                                Fig. 1. Intelligent agent action planning

    In other words, any change in the parameter that                 initial position of the subject. We also set the second point
describes some of the characteristics of the subject’s               denoting the position of the target. The state of the subject
position in the phase space can be associated with a change          changes as a result of some elementary actions in the state
in some scalar quantity — the transition price, which is             space, while the set of possible actions is limited and is
understood as the conditional cost of the action. Each of            described by a set of rules of behavior, each of which is
the parameters has its own “price scale”, which                      displayed by a vector in the phase space. The sequential
dynamically depends on the current state of the subject and          step-by-step execution of actions forms a certain trajectory
is determined by the value of this parameter to solve the            of the subject in the state space. In a two-dimensional
problem.                                                             interpretation, the task is shown schematically in Fig. 2.

3. Adapting rules of conduct
   In the phase space or in the space of states we define
some arbitrary point, which will be considered as the




                                            Fig. 2. Schematic representation of the problem

    The shortest direction or trajectory from the source to          action vector, is unlikely to coincide with the shortest
the target state is shown by a dashed line. It is obvious that       direction. Clearly, a rational trajectory of movement
the result of the implementation of one or another rule of           should be as close to this line as possible.
behavior, which is displayed in the phase space by some
   But in conditions of inaccurate information about the               an operation can be done “manually”, based on visual
properties of each of the vectors, it is difficult to construct.       estimates. One of the possible results may be, for example,
For the two-dimensional problem shown in figure 2, such                the path options shown in Fig. 3.




                                  Fig. 3. A schematic example of the construction of rational trajectories

    In the presented schematic example, based on a visual              by the degree of approximation of the new state to the
assessment of the set of action vectors, such trajectories             target state obtained as a result of this action from this
can be constructed “manually”. There are at least two                  position. With a positive effect (approaching the goal),
rational options for the trajectories of the transition from           incentive is introduced, with a negative effect (moving
the initial state to the target, and it is enough to use a             away from the goal) - a penalty. At the beginning of the
combination of only two of all possible actions. When                  calculation, all ratings are set close to zero.
comparing the obtained trajectories, it is seen that that the              The next move step is randomly selected from the list
upper one is slightly more preferable in terms of the degree           of all possible actions, but taking into account their rating.
of approximation to the target position, but its                       The probability of choosing an action increases in
implementation requires one more step. To determine the                proportion to the rating. As a result of the calculations, the
best of them, additional evaluation criteria are needed, for           initially uniform distribution of the rating of actions is
example, a comparison of criteria for accuracy and                     gradually deformed towards growing ratings. At each
resource costs. Based on the above examples, it can be                 iteration, the current state in which the subject was located
argued that in a multidimensional phase space, the                     is remembered. Thus, a model of stereotypical situations
construction of the optimal path is possible only with full            is formed in which the effects of the same actions can
knowledge of the situation: where, by what means, with                 manifest themselves in different ways. Each iteration stops
what intermediate results certain actions arising from the             after reaching the target state or after completing a given
rules of behaviour of agents can be realized.                          number of steps. At each subsequent iteration, the ratings
    The following scheme was proposed and tested.                      accumulated in previous experiments are saved and used
Multidimensional phase space and an unordered set of                   when choosing actions. Therefore, the knowledge base is
permissible actions are set. At the same time, the number              accumulated and the conditional agent is “trained” in the
of actions and the dimension of space are determined by                rational choice of actions depending on the position in the
the conditions of a specific task. It is assumed that the basis        state space in which it is located.
for decision-making by the agent in the current situation to
choose a specific action is some a priori assessment of the            4. Results
value of certain actions. Initially, it can be set by experts              Testing of the methodology was carried out on an
or randomly generated.                                                 arbitrary example in a space of 12 measurements with a
    Further refinement of the assessment of the actions that           total number of rules of behavior equal to 50 (Fig. 4).
the agent must carry out is the result of a wide series of                 The fig. 4 shows 33 of 50 actions in a 12-dimensional
numerical experiments, in each of which the trajectory of              state space. The positions of the initial and target states
motion of the agent states in the phase space from the                 were randomly selected as vectors in the phase space. In
source to the target is constructed. In the process of                 addition, in the form of increments of vectors, action
performing calculations, a certain rating is assigned to               vectors were arbitrarily specified.
each perfect action. The rating of each action is determined
                                                   Fig. 4. List of available actions

    Let us dwell in more detail on the data given in the                The data given in the table reflect the obtained picture
table. As already mentioned above, the agent’s training             of the learning outcomes at some intermediate stage of the
technology is demonstrated using a hypothetical example,            experiments. For clarity, a visual analysis introduced the
in which a certain zone of the phase space is                       color of the cells. In this case, color means the directivity
conventionally defined as a range of changes in the                 of the effect, and the intensity of the color is its immediate
coordinate values of the parameters, and the initial and            meaning. In other words, the cells highlighted in red
target position of the agent are determined inside this zone.       indicate that the choice of this action will lead to the loss
Variants of the agent’s possible actions are randomly               of a parameter or resource. Moreover, the more intense the
generated by coordinatewise specifying a certain set of             color, the greater the loss. Positive effects highlighted in
increment vectors. When generating incremental vectors,             green.
the only condition is to limit the length of each of these              Fig. 5 shows the result of a computational procedure
vectors to a value significantly less than the distance             for finding a rational trajectory.
between the initial and target position of the agent in the
phase space.




                                           Fig. 5. Schematic representation of the problem
    The process of qualitatively improving the procedure            application-oriented simulation, prediction and
for choosing rational actions from a given set is visible in        expertizes of conflict interaction, 2018.
the fig. 5. The x-axis shows the number of experiments, in     [4] Bousquet F., Barreteau O., Le C., Mullon C., Weber
each of which a motion pattern limited in the number of             J., An environmental modelling approach. The use of
steps is plotted in the phase space, and the y-axis is the          multi-agent simulations, 1999.
distance from the target state in the adopted metric. It can   [5] Ferber J., La kénétique: dés systems multi-agents à
be seen that the agent training process leads to the                une science de l'interaction, 1994.
achievement of some non-improved level of performance,         [6] Petukhov A., Мalhanov A., Sandalov V., Petukhov
which, however, allows to get close enough to the target            V., Mathematical Modeling of the Ethno-social
state.                                                              Conflicts by Non-linear Dynamics. Proceedings of the
                                                                    7th International Conference on Simulation and
5. Conclusions                                                      Modeling Methodologies, Technologies and
    Proposed technological solution for modeling the                Applications, 2017.
behavior of intelligent agents provides a fairly broad basis   [7] A.Kh. Khakimova, O.V. Zolotarev, M.A. Berberova.
for studying the features of constructing trajectories of           Visualization of bibliometric networks of scientific
achieving goals in phase space and the adaptive behavior            publications on the study of the human factor in the
of agents in a conflict situation.                                  operation of nuclear power plants based on the
     Due to the specifics of the problem, approaches and            bibliographic database Dimensions. Scientific
methods for modeling the phase space of states of a                 Visualization, 2020, volume 12, number 2, pages 127-
conflict environment are proposed, which allow                      138, DOI: 10.26583/sv.12.2.10, E-ISSN:2079-3537.
determining strategies for rational trajectories of goal       [8] M.A.Berberova, K.I.Chernyavskii, «Comparative
achievement. The technology for modeling agent behavior             assessment of the NPP risk (on the example of Rostov
is based on a detailed description of their behavior based          and Kalinin NPP). Development of risk indicators
on a scheme of intelligent transitions between states and           atlas for Russian NPPs», GraphiCon 2019 Computer
behavior models.                                                    Graphics and Vision. The 29th International
    Such a detailed description of complex behavior                 Conference on Computer Graphics and Vision.
models makes it possible to uniformly display in the model          Conference Proceedings (2019), Bryansk, Russia,
of a conflict environment the essential aspects of the              September 23-26, 2019, Vol-2485, urn:nbn:de:0074-
behavior of real-world prototype objects. The proposed              2485-1, ISSN 1613-0073, DOI: 10.30987/graphicon-
scheme for ensuring the adaptive behavior of agents in a            2019-2-290-294,                 http://ceur-ws.org/Vol-
conflict environment is an integral part of the                     2485/paper67.pdf, p. 290-294.
methodological and algorithmic support of an electronic        [9] M.A.Berberova, A.S.Oboimov, A.Kh.Khakimova,
training ground for studying the conflict interaction of            O.V.Zolotarev, «Risk-informed security system. The
complex systems.                                                    use of surveillance cameras for the particularly
    According to the authors, this technology has shown its         hazardous facilities safety», GraphiCon 2019
effectiveness in test tasks and can be recommended for use          Computer Graphics and Vision. The 29th International
as part of an electronic training polygon to test its               Conference on Computer Graphics and Vision.
capabilities in developing strategies for agents' behavior          Conference Proceedings (2019), Bryansk, Russia,
under conditions of uncertainty during their interaction.           September 23-26, 2019, Vol-2485, urn:nbn:de:0074-
    A similar technology can be applied in automated                2485-1, ISSN 1613-0073, DOI: 10.30987/graphicon-
process control support systems, where an operational               2019-2-321-325,                 http://ceur-ws.org/Vol-
assessment of the situation is necessary based on the               2485/paper74.pdf, p. 321-325.
available incomplete or inaccurate data with the               [10] Baluta V., Osipov V., Chetverushkin B., Yakovenko
development of recommendations for rational actions. In             Y., Conceptual issues of model representation of
particular, when creating combat control systems,                   conflicts. Proceedings of the International Scientific
ensuring the safety of facilities, responding in emergency          Conference, 2019.
situations, as well as in control systems for unmanned
                                                               About the authors
vehicles.
                                                                   Karandeev Alexander A., phd student and Junior Scientific
References                                                     Associate of the Keldysh Institute of Applied Mathematics,
                                                               Russian Academy of Sciences. E-mail: KarAlex755@gmail.com
[1] Klügl F., Oechslein C., Puppe F., Dornhaus A., Multi-          Baluta Viktor I., PhD in Technical Science, Docent, Senior
    Agent Modelling in Comparison to Standard                  Scientific Associate of the Keldysh Institute of Applied
    Modelling, 2014.                                           Mathematics, Russian Academy of Sciences. E-mail:
[2] Nechaev Y., Osipov V., Chetverushkin B., Baluta V.,        vbaluta@keldysh.ru
    Ontological synthesis of management decisions in               Osipov Vladimir P., PhD in Technical Science, Docent, Lead
    conditions of antagonistic conflicts, 2018.                Scientist of the Keldysh Institute of Applied Mathematics,
[3] Baluta V., Nechaev Y., Osipov V., Chetverushkin B.,        Russian Academy of Sciences. E-mail: osipov@keldysh.ru
    Conceptual base of a supercomputer platform of