<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic target recognition module development for fire control system based on machine learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victoria Vysotska</string-name>
          <email>Victoria.A.Vysotska@lpnu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roman Romanchuk</string-name>
          <email>roman.v.romanchuk@lpnu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariia Nazarkevych</string-name>
          <email>mariia.a.nazarkevych@lpnu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hetman Petro Sahaidachnyi National Army Academy</institution>
          ,
          <addr-line>Heroes of Maidan 32, 79026 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Stepan Bandera 12, 79013 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Organization and Planning Department of the Training Center “Partnership for Peace” of the International Peacekeeping and Security Center</institution>
          ,
          <addr-line>79026 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Target recognition is a priority in military affairs. This task is complicated by the fact that it is necessary to recognize moving objects, different topography and landscape create obstacles for recognition. Combat actions can take place at different times of the day, accordingly, the lighting angle and general lighting must be taken into account. It is necessary to detect the object in the video by segmenting the video frames and to recognize and classify it. In the work, the authors propose the development of a target recognition module as a component of the fire control system within the framework of the proposed information technology through artificial intelligence use. The YOLOv8 pattern recognition model family was used to develop the target recognition module. The data was collected from open sources, in particular, from video footage posted in open sources on the YouTube platform. The main task of data pre-processing is the classification of three classes of objects on video or in real-time - APC, BMP, and TANK. The dataset is formed using the Roboflow platform based on marking tools and, subsequently, augmentation tools. The data set consists of 1193 unique images - approximately equally for each class. The training was conducted using Google Colab resources. 100 epochs were taken to train the model. The analysis is carried out according to mAP50 (mean Average Precision as 0.85), mAP50-95 (0.6), precision (0.89) and recall (0.75) metrics. Large losses are present because the background was not taken into account in the study - training the module based on validated data (images) of the background without the technique. This will be the next step. It is also necessary to expand the classification of objects of military equipment.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;security</kwd>
        <kwd>privacy</kwd>
        <kwd>moving objects recognition</kwd>
        <kwd>targets identification</kwd>
        <kwd>machine learning</kwd>
        <kwd>YOLO 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Today, the leading armies of the world strive to increase the capabilities of their main
models of equipment and weapons by modernizing the existing fleet or developing the
latest models. Automatic target recognition (Automatic Target Recognition Unit) is the
ability of an algorithm or device to recognize targets or objects based on data received from
sensors, including video surveillance, for example from unmanned aerial vehicles (UAVs),
such as drones, or from video - recorders on combat vehicles. On the other hand, due to the
increased use of UAVs for reconnaissance by the enemy, the security and privacy of many
critical locations may be compromised. Therefore, they are also a legitimate target for fire
control detection.</p>
      <p>Target recognition information technology is a key component in cruise missiles and
UAVs, as well as in the development of combat robots or sapper robots. Automatic target
recognition is used not only in military affairs but also, for example, in the organization of
searching for people/objects (at sea, in the area of natural disasters, fires, etc.).</p>
      <p>The task of automatic target recognition in combat conditions is complicated by several
factors, in particular:
1. Possible movement of the recognized object;
2. The movement of the object (combat vehicle or UAV), from where the video
surveillance and further recognition of the targets originates;
3. Different weather conditions;
4. Different topography and landscape, including forest strips;
5. Presence of other objects that are potentially not targets (buildings,
downed/destroyed combat vehicles, parts of structures such as bridges, etc.);
6. Lighting;
7. Potentially, the object being recognized is not an enemy;
8. Part of the recognized object is hidden behind obstacles;
9. The observation angle for different objects is different (for UAVs from top to bottom,
for combat vehicles not only forward/around, but upwards for UAVs, for example).</p>
      <p>
        As you can see, lighting conditions, different sizes of objects, moving backgrounds and
various background contrasts significantly affect the quality, efficiency and speed of object
recognition from video surveillance [
        <xref ref-type="bibr" rid="ref1 ref2">1-2</xref>
        ]. It is necessary not only to detect the object in the
video by segmenting the video frames, but also to recognize and classify it (for example, a
drone or a bird, a car or a building, etc.), and this is usually during the movement of both the
surveillance object and the object-observer under adverse conditions in real-time.
Detecting flying objects or objects moving in adverse conditions in the video is different
from standard object detection because the size of a stationary/moving/flying and/or
partially hidden object behind another object is constantly changing in frames depending
on its distance and the movement of the observing object. It has problems such as low
resolution, changes in lighting between day and night and unstable background, different
weather conditions. Also, the accuracy of recognition depends on the quality of the
surveillance camera, the selection of which during hostilities is not a controlled process. The
complexity of observation increases when recognizing from 2D (in front of the camera) to
3D (from above from the bottom at different angles at different heights) moving objects,
taking into account scaling and proportions. Similarly, it reduces the accuracy of recognizing
objects that are visually similar to each other and differ in small features or their absence
when viewed from different angles or when the hull is partially hidden behind other natural
objects or buildings (for example, some modifications of the T series tanks). Therefore, the
detection and recognition of stationary/moving/flying/moving and/or partially hidden
objects by other objects in the state of immobility/movement of the observer in different
weather conditions, landscapes, lighting and at different heights have a large scope of
observation and high mobility. There is a strong need for such applications in the real world
because of the size differences within the same object type and the spatial resolution of the
sensor.
      </p>
      <p>Thus, the purpose of the work is to develop a method of target recognition in real-time
as a component of the fire control system, due to the use of artificial intelligence.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>In recent years, during the full-scale war in Ukraine, with the gradual improvement of drone
control technology, UAV remote sensing images and videos have become an important
source of operational data. At the same time, equipping combat vehicles with video
recorders for video surveillance with elements of artificial intelligence and machine
learning for real-time object recognition will increase the level of security of combatants
with appropriate timely responses to the results of target recognition.</p>
      <p>Video frames  video frame segmentation  detection of potential objects  detection
of moving objects  recognition of objects as potentially dangerous  object classification
 object identification.</p>
      <p>
        Today, neural networks and deep learning are commonly used for tasks such as image
segmentation [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3-5</xref>
        ], object detection [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6-8</xref>
        ], and image classification [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9-11</xref>
        ]. Most of the
currently applied deep neural network models, such as PSPNET [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], U-NET [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], RESNET
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and VGG [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], are developed based on manually collected image (non-video) datasets
under favourable conditions, such as MS-COCO [12], VOC2012 [13], VOC2007 [14].
      </p>
      <p>There are two general scenarios for the application of methods of detecting objects by
remote sensing by a drone or based on video surveillance from a car, in particular, data
processing is assumed:


after flight/trip using stationary computers (requires high detection and
identification accuracy).
in real-time during the flight/ride, when the onboard computer on the drone or in
the car, respectively, synchronously processes the video data in real-time. The
parameters of the model must be within a certain scale to meet the requirements for
the operation of the embedded equipment. Once the operating conditions are met,
the detection accuracy of the method should also be as high as possible.</p>
      <p>Therefore, applied neural network-based object detection methods must meet different
requirements for each scenario.</p>
      <p>Thus, neural network methods for detecting objects in drone remote sensing video or
video surveillance in combat vehicles must be able to adapt to the specific characteristics of
this data. They should be designed to meet post-flight/trip data processing requirements,
which can provide high accuracy and recall speed, or they should be designed as
smallerscale parameter models that can be deployed in embedded hardware environments for
real-time processing on drones/ cars. In this work, we propose the application of a neural
network based on the YOLOv8 architecture for the automatic recognition of objects as
potential targets of a fire control system.</p>
      <p>Currently, numerous methods of object detection based on neural networks have been
proposed, in particular, using the YOLO series [15-22]. Unlike the two-step methods, the
one-step method combines object location and classification in one step, achieving real-time
object detection on both desktop and embedded hardware. These methods not only achieve
good identification results, but also offer several improvements in areas such as training
data augmentation methods, network training methods, loss functions, activation functions,
and network model structures.</p>
      <p>
        YOLOX, a neural network model with one-step object detection, is proposed in [18]. In
[19], the authors proposed a neural network model with one-step target detection. In [20],
the authors investigated the optimal speed and accuracy of object detection based on
YOLOv4. The authors in [23] described CSPDarkNet as the backbone structure of the
network, improving the learning ability of convolutional neural networks, and allowing the
network to maintain the accuracy of feature map extraction. In [24], the authors proposed
the CrowdDet method based on a neural network for detecting dense and mutually closed
targets in images. In the neck part of the network, the SPPF module and the PAFPN module
have been introduced [25]. The author still uses CSPDarkNet [23] as the backbone network,
but introduces SiLU as an activation function that solves the gradient dispersion problem
when the input of the ReLU function is negative and the output is 0 [
        <xref ref-type="bibr" rid="ref12">26-27</xref>
        ].
      </p>
      <p>
        There are many studies based on different versions of YOLO, but still, the most important
problem in object identification is the effective detection of small objects and the accuracy
of classification of various moving objects [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17 ref18">28-33</xref>
        ] under different environmental
conditions (for example, an object in different gradations of green colour on the
background, as well as a different spectrum of the green colour [
        <xref ref-type="bibr" rid="ref19">34</xref>
        ]) in a video stream of
different image quality [
        <xref ref-type="bibr" rid="ref20 ref21 ref22 ref23 ref24 ref25 ref26 ref27 ref28 ref29 ref30">35-45</xref>
        ] with subsequent storyboarding and segmentation of the
corresponding images for identification and classification [
        <xref ref-type="bibr" rid="ref31 ref32 ref33 ref34 ref35 ref36 ref37 ref38 ref39 ref40 ref41">46-63</xref>
        ].
      </p>
      <p>Remote sensing images often have large dimensions, complex backgrounds, and a
significant presence of small objects. The proposed solution is focused on optimizing the
accurate detection of small objects at a distance and objects in motion under various
environmental conditions. The advantage of YOLO series networks is the use of multi-level
detection heads, which allow the detection of objects of different sizes from different levels
of feature vectors. Our approach mainly focuses on detecting small objects as well as moving
objects using feature vectors from lower layers that have higher spatial resolution. To
achieve this goal, we use a machine learning module to optimize their semantic
characteristics. Detection heads can obtain feature vectors with both high spatial resolution
and accurate semantic information, thus increasing overall identification accuracy.</p>
      <p>In deep learning, feature extraction methods SIFT (Scale Invariant Feature Transform)
and HOG (Histograms of Oriented Gradients) performed this task by applying some machine
learning algorithms on top of the classifier. Some deep learning methods are applied to
colour images, while others are applied to IR images. Real-time processing of IR images is
simpler because it requires less memory and computing power. It also does not affect
different lighting conditions. However, it is practically impossible to collect a training
dataset for some subject areas (for example, military equipment during the war).
Algorithms of the YOLO family, based on the CNN architecture, are widely used and
wellknown algorithms for solving object detection problems. YOLO v4 and YOLO v5 are mostly
used models. YOLO v4, being a modified version of YOLO3, uses a cross-stage partial
network (CSPNet) in the Darknet, creating a new feature extractor backbone called
CSPDarknet53. To increase the efficiency of the algorithm, YOLOv4 uses a bag of freebies
and a bag of special offers. Total loss of IOU (CIOU), dropout lock regularization and many
expansion approaches. Mish activation, Diou-NMS and modified pathway aggregation
networks are included in the speciality package. But YOLOv5 is different from previous
versions. Here PyTorch is used instead of Darknet. It uses CSPDarknet53 as structural
support. This pipeline removes the redundant gradient information seen in large pipelines
and incorporates gradient transformation into feature maps, which speeds up inference,
improves accuracy, and reduces model size by reducing the number of parameters. It
enhances the flow of information using a Path Aggregation Network (PANet), resulting in
three different feature map outputs for multiscale prediction. This improves the model's
ability to effectively predict small and large objects. The image is sent to PANet for feature
fusion after input to CSPDarknet53 for feature extraction. The processing speed of YOLO v4
and v5 ranges from 45 to 150 frames per second. However, unlike the faster R-CNN, it has
lower recall error and higher localization. Since each grid can only offer two bounding
boxes, it also has trouble detecting nearby objects and small objects. The latest addition to
the YOLO object detection family is the YOLO v8 model. It is the fastest and most accurate
real-time object detector available today. All YOLOv8-based tools outperform previous
object detectors in terms of speed and accuracy.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Models and methods</title>
      <p>The paper discusses a hybrid approach using CNN-LSTM to improve the detection
performance of military equipment with a moving background and different distances, as
well as its TANK/APC/BMP. The main contributions of the article.</p>
      <sec id="sec-3-1">
        <title>1. Collection of images from open sources for YouTube platforms.</title>
        <p>2. Hybrid CNN-LSTM model with hyperparametric tuning using Bayesian optimization
for object detection.
3. Detailed analysis of the YOLOV8 model on different ranges of images and
determination of their accuracy with a certain confidence value.</p>
        <p>Object detection algorithms in deep learning are mainly divided into regional and
regression. The main task of detecting objects of military equipment is to detect an object
in a frame where objects of different target classes are present; therefore, object
classification is a prerequisite for object detection using a bounding box.</p>
        <p>Regression-based algorithms are mainly used for real-time object detection. These are
one-step frameworks based on global regression/classification that directly map image
pixels to bounding box coordinates, reducing time consumption. One of the fastest object
recognition models is YOLO, which can analyse frames at up to 150 FPS for small networks.
Although YOLO is not the most accurate model in terms of Mean Average Accuracy (mAP),
it performed reasonably well during training.</p>
        <p>As part of the detection of objects of military equipment, an experiment was conducted
on automatic recognition and identification of targets of the fire control system. This was
implemented using various object detection models. The experiment is focused on the
implementation of the YOLO v8 algorithm for comparison with other versions.</p>
        <p>Algorithms based on regional offers. A regional convolutional neural network (R-CNN)
is proposed for CNN matching. Compared with models without deep CNNs, R-CNN
significantly improved the detection performance for the mean average precision (mAP). It
has several disadvantages, including costly training in terms of money and time, and worst
of all, high latency (detection time). Building on the performance of R-CNN, fast R-CNN
improves accuracy by speeding up training and testing. Fast R-CNN dramatically reduces
training and testing time, but regional proposals are still generated using traditional
methods that require a lot of time for pre-processing. A faster R-CNN is proposed as a
solution to the region proposal bottleneck problem, which makes region proposals through
a neural network. Fast/faster R-CNN and other object detectors using regional proposal
networks have shown superior performance in many tests. However, they are not always
successful in finding small objects. Current approaches are poorer in terms of repeatability
and generalizability when real-world circumstances are constantly changing because they
depend on specific image data.</p>
        <p>Classification of stationary/moving/flying and/or partially hidden objects of military
equipment by another object under adverse conditions in real-time. Detecting flying objects
or objects moving in adverse conditions in the video is different from standard object
detection because the size of a stationary/moving/flying and/or partially hidden object
behind another object is constantly changing in frames depending on its distance and the
movement of the observing object. It has problems such as low resolution, changes in
lighting between day and night and unstable background, different weather conditions.
Also, the accuracy of recognition depends on the quality of the surveillance camera, the
selection of which during hostilities is not a controlled process. The complexity of
observation increases when recognizing from 2D (in front of the camera) to 3D (from above
from the bottom at different angles at different heights) moving objects, taking into account
scaling and proportions. Similarly, it reduces the accuracy of recognizing objects that are
visually similar to each other and differ in small features or their absence when viewed from
different angles or when the body is partially hidden behind other natural objects or
buildings.</p>
        <p>Deep learning-based object detection methods for various tasks, which include lane
detection, intelligent vehicle systems, detection of moving objects, including military
equipment, etc.</p>
        <p>The observation module for automatic recognition and identification of targets of the fire
control system M is represented by a simulation model via a tuple:

= &lt;  ,  ,  ,  ,  , , , &gt;,
(1)
where I is a set of input data in the form of a video stream from a video camera, I = {i1, i2,
i3, i4}; O is a set of initial data in the form of recognition and identification of objects of
military equipment, O = {o1, o2, o3}; R is basic rules for processing the input data of the video
stream, R = {r1, r2, r3, r4, r5}; U is parameters for processing the input data of the video
stream, U = {u1, u2, u3, u2, u3}; N is a neural network for learning the recognition,
identification and classification of military equipment objects such as TANK/BMP/APC;  is
operator of analysis and storyboarding of input data of the video stream;  is image
processing operator through segmentation and analysis of segmented objects;  is operator
of recognition, identification and classification of objects of military equipment such as
TANK/BMP/APC.</p>
        <p>The</p>
        <p>main processes of the surveillance model for automatic recognition and
identification of fire control system targets are "Video Stream Processing", "Image
Processing", "Machine Learning" and "Object Classification".</p>
        <p>The process of "Processing a video stream" will be described by a superposition:</p>
        <p>MUL =,
MUL=((((CСU, i1), o2, i4), u4), r4),</p>
        <p>MAU =,</p>
        <p>MAU =(((i1, i2, i3), r1, u1), u2),
where  is the operator for recognizing any potential objects in the image (buildings,
bridges, military equipment, etc.); i1 is a set of data from the video stream and images of the
original; i2 is a set of images of military equipment; i3 is a set of landscape background data;
r1 is the rules for framing the video stream into an image; u1 is a set of conditions for forming
images from a video stream; u2 is a set of image analysis requirements, including noise
filtering.</p>
        <p>The process of "Image processing" will be described by superposition:</p>
        <p>MСU =,</p>
        <p>MCU =(((CAU, i2, i3, i4), o1, r2, u3), r3),
where  is the operator for recognizing potential objects of military equipment in the
image; i2 is set of images of military equipment; i3 is a set of landscape background data; i4
is dictionaries of validated images of military equipment; o1 is the set of all recognized
objects in the image; r2 is image segmentation rules; r3 is image segment analysis rules; u3
is a set of image processing conditions.</p>
        <p>The process of "Machine learning" will be described as:
(2)
(3)
(4)
(5)
(7)
(8)
where  is the identification operator of a recognized object of military equipment on
multiple images cut from the video stream; i1 is a set of data from the video stream and
images of the original; i4 is dictionaries of validated images of military equipment; o2 is the
set of all recognized objects of military equipment in the image; r4 is machine learning rules
of the neural network for identification of military equipment; u4 is a set of conditions for
recognition and identification of objects of military equipment.</p>
        <p>The process of "Classification of objects" will be described as:</p>
        <p>MUS =,
MUS=((((CUS, i1), o3, i4), u5), r5),
(9)
(10)
where  is the operator of the classification of the identified object of military equipment
on multiple images cut from the video stream; i1 is a set of data from the video stream and
images of the original; i4 is dictionaries of validated images of military equipment; o3 is the
set of all identified objects of military equipment in the image; r5 is rules for the classification
of military equipment; u5 is a set of requirements for the classification of recognized objects
of military equipment.</p>
        <p>The analysis is carried out using classification or clustering, which segments according
to certain criteria. Although the collection of information occurs automatically, it is still
necessary to implement such studies according to the recognition, identification and
classification of objects in unfavourable conditions in motion and with poor image quality,
and the corresponding processing of the results. The effectiveness of processing the
corresponding background and objects on it also significantly affects the research results
(for example, green on a green background or part of an object hidden behind another
object). One of the most important criteria of such technology is the ability to collect data
depending on the period of the day and season, and their periodicity due to a change in the
background due to the results of active hostilities in a certain area.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments, results and discussion</title>
      <p>In the work, the authors propose the development of a target recognition module as a
component of the fire control system within the framework of the proposed information
technology through artificial intelligence use.</p>
      <p>The YOLOv8 pattern recognition model family was used to develop the target
recognition module. This is the latest version of Ultralitics' popular real-time object
detection and image segmentation product. YOLOv8 comes bundled with the following
prebuilt models:


</p>
      <p>Image classification models pre-trained on ImageNet database with 224 image
resolution.</p>
      <p>Instance segmentation control points trained on the COCO segmentation dataset
with 640 image resolution.</p>
      <p>Object detection control points trained on the COCO detection dataset with 640
image resolution.</p>
      <p>The output is performed at almost 105 frames per second on the GPU of the average
modern laptop, while the ad-large model runs at an average speed of 17 frames per second.</p>
      <p>YOLOv8 uses the PuTorch framework, a framework for developing deep neural networks
from Facebook. YOLOv8 has several advantages over other tools, including:</p>
      <sec id="sec-4-1">
        <title>1. Most services support the provision of computing power.</title>
        <p>2. A large number of methods of application and use of models.
3. High level of accuracy, confirmed by tests on COCO and Roboflow 100 datasets.</p>
        <p>As mentioned above, YOLOv8 achieves high accuracy on the COCO dataset. For example,
the YOLOv8m model - achieves 50.2% mAP when measured on COCO. When compared
against Roboflow 100, a dataset that specifically evaluates model performance in different
domains, YOLOv8 performed significantly better than YOLOv5. In addition, YOLOv8
provides developers with a significant list of features. Unlike other models that split tasks
between many different Python files, YOLOv8 comes with a CLI interface that makes
training the model more intuitive.</p>
        <p>The Google Collaboratory service, better known as "Colab", was used for training. “Colab”
is a cloud version of Jupyter Notebook. Using Colab does not require installing and running
or upgrading your computer hardware to meet Python's CPU/GPU intensive requirements.
In addition, Colab provides free access to computing infrastructure such as storage, RAM,
computing power, graphics processing units (GPUs) and tensor processing units (TPUs).</p>
        <p>The methods used during the study of the formed dataset.
1. Flip – Horizontal (flipping the image object horizontally).
2. Rotation – Between -15 and +15 (rotation of the image object – clockwise or
counterclockwise by a degree from -15 to +15).
3. Brightness – Between -25% and +25% (changing the brightness of the image to
increase the resistance of the model to changes in lighting and camera settings –
from -25% to +25%).
4. Cutout – 3 boxes with 10% size each (cut out a part of the image – 3 boxes of 10%
size each).
5. Bounding Box: Blur – Up to 2.5px
6. Bounding Box: Noise – Up to 15% of pixels.</p>
        <p>The last two points are used to expand the level of the bounding box when
forming/generating new training data, only changing the content of the bounding boxes of
the original image. Image upscaling is the process of increasing the size of a dataset by
manipulating the existing training data. Zooming in helps the model to better generalize to
a wide range of contexts. For example, you can change the brightness or darkness of an
object relative to its background. Or perhaps blur the subject against its background for
tasks that often involve shooting fast-moving subjects. Modifications to the bounding box
alone led to systematic improvements, especially for models that were small datasets
(several thousand photos). You can also change the colours of only objects in the OCR image,
blur only moving objects, such as military equipment in various shades of green, rotate
objects, such as objects in the overhead view, and flip the orientation of objects to create
mirroring effects similar to those present in most camera situations.</p>
        <p>Before the development of the automatic target recognition module of the fire control
system, the workspace is organized and the rules for accessing the data store and processing
relevant data from it are defined. The Google Collaboratory service, better known as
"Colab", was used for training. “Colab” is a cloud version of Jupyter Notebook. The main task
of data preprocessing is the classification of three classes of objects on video or in real-time
- APC, BMP, and TANK. Next, the corresponding data set was created and filled. The data
was collected from open sources, in particular, from video footage posted in open sources
on the YouTube platform (videos from the promotion of Rashka technology in the first days
of the war on the territory of Ukraine and from military parades of Russian equipment).
This process included searching for images and videos of the above-mentioned objects and
marking the corresponding objects. The dataset is formed using the Roboflow platform
based on marking tools and, subsequently, augmentation tools. The data set consists of
1193 unique images - approximately equally for each class. After applying image
preprocessing and argumentation methods, the data set has the following form (Table 1):
train: ../train/images
val: ../valid/images
test: ../test/images
nc: 3
names: ['bmp', 'btr', 'tank']</p>
        <p>The training was conducted using Google Colab resources. 100 epochs were taken to
train the model. The statistical results of neural network training are shown in Fig. 1-2. The
analysis was carried out according to mAP50 (mean Average Precision), mAP50-95,
precision and recall metrics (Fig. 1).</p>
        <p>AP (Average precision) is a popular metric for measuring the accuracy of object detectors
such as Faster R-CNN, SSD, etc. Average Precision calculates the average precision for recall
in the range 0 to 1. It is a measure of the model's precision, taking into account only "easy"
detections. mAP50-95: Average of the average accuracy calculated at different IoU
thresholds ranging from 0.50 to 0.95. It gives a complete picture of the performance of the
model at different levels of detection complexity.</p>
        <p>Precision measures how accurate your predictions are. For example, what percentage of
your predictions are correct? Recall measures how well you find all positive samples. For
example, we can find 75% of all possible positive cases in our best predictions. As we can
see from Fig. 1 The Precision metric gives a larger swing at the beginning and becomes more
similar to the mAP50 at the end as the number of trials increases. The mAP50-95 has bad
values (0.5-0.6) at the end as the number of trials increases. The Recall metric has a
relatively constant value in the range of 0.85-0.75 after half of the conducted experiments
(epochs). This is not a good enough result and needs further research and model training
on a larger dataset of actual data. The Precision metric gives slightly better results - in the
range of 0.85-0.89.</p>
        <p>Figures 3-5 show examples of system operation. Large losses are present because the
background was not taken into account in the study - training the module based on
validated data (images) of the background without the technique. This will be the next step.
It is also necessary to expand the classification for objects of military equipment - which is
exactly T-64, E-72 or T-90.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>Target recognition is a priority in military affairs. This task is complicated by the fact that
it is necessary to recognize moving objects, different topography and landscape create
obstacles for recognition. Combat actions can take place at different times of the day,
accordingly, the lighting angle and general lighting must be taken into account. It is
necessary to detect the object in the video by segmenting the video frames and to recognize
and classify it. The training was conducted using Google Colab resources. 100 epochs were
taken to train the model. The analysis was carried out according to mAP50 (mean Average
Precision), mAP50-95, precision and recall metrics.</p>
      <p>The proposed method can be used to identify objects of a military nature and recognize
targets to create (modernize) modern fire control systems of modern military equipment.
[12] T. Lin, M. Maire, S.J. Belongie, L.D. Bourdev, R.B. Girshick, J. Hays, P. Perona, D.</p>
      <p>Ramanan, P. Doll’ar, C.L. Zitnick, Microsoft COCO: Common Objects in
Context. arXiv 2014, arXiv:1405.0312. URL: http://xxx.lanl.gov/abs/1405.0312.
[13] M. Everingham, L. Van Gool, C.K.I. Williams, J. Winn, Zisserman, A. The PASCAL
Visual Object Classes Challenge (VOC2012) Results. URL:
http://www.pascalnetwork.org/challenges/VOC/voc2012/workshop/index.html.
[14] M. Everingham, L. Van Gool, C.K.I. Williams, J. Winn, A. Zisserman, The PASCAL
Visual Object Classes Challenge (VOC2007) Results. URL:
http://www.pascalnetwork.org/challenges/VOC/voc2007/workshop/index.html.
[15] G. Jocher, A. Chaurasia, J. Qiu, YOLO by Ultralytics. 2023. URL:
https://github.com/ultralytics/ultralytics/blob/main/CITATION.cff.
[16] C.Y. Wang, A. Bochkovskiy, H.Y.M. Liao, YOLOv7: Trainable bag-of-freebies sets new
state-of-the-art for real-time object detectors, in Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–
22 June 2023; pp. 7464–7475.
[17] C. Li, et al. YOLOv6: A single-stage object detection framework for industrial
applications, arXiv 2022, arXiv:2209.02976.
[18] Z. Ge, S. Liu, F. Wang, Z. Li, J. Sun, Yolox: Exceeding yolo series in 2021. arXiv 2021,
arXiv:2107.08430.
[19] G. Jocher, et al. Ultralytics/Yolov5: V6.0—YOLOv5n ’Nano’ Models, Roboflow
Integration, TensorFlow Export, OpenCV DNN Support, 2021. URL:
https://zenodo.org/record/5563715.
[20] A. Bochkovskiy, C.Y. Wang, H.Y.M. Liao, Yolov4: Optimal speed and accuracy of object
detection, arXiv 2020, arXiv:2004.10934.
[21] J. Redmon, A. Farhadi, YOLO9000: Better, faster, stronger, in Proceedings of the IEEE
Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July
2017; pp. 7263–7271.
[22] J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: Unified, real-time
object detection, in Proceedings of the IEEE Conference on Computer Vision and Pattern
Recognition, Las Vegas, NV, USA, 27–30 June 2016, pp. 779–788.
[23] C.Y. Wang, H.Y.M. Liao, Y.H. Wu, P.Y. Chen, J.W. Hsieh, I.H. Yeh, CSPNet: A new
backbone that can enhance learning capability of CNN, in Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA,
14–19 June 2020, pp. 390–391.
[24] X. Chu, A. Zheng, X. Zhang, J. Sun, Detection in crowded scenes: One proposal,
multiple predictions, in Proceedings of the IEEE/CVF Conference on Computer Vision
and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020, pp. 12214–12223.
[25] S. Liu, L. Qi, H. Qin, J. Shi, J. Jia, Path aggregation network for instance segmentation.</p>
      <p>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt
Lake City, UT, USA, 18–23 June 2018, pp. 8759–8768.
[26] X. Glorot, A. Bordes, Y. Bengio, Deep sparse rectifier neural networks, in Proceedings
of the Fourteenth International Conference on Artificial Intelligence and Statistics, Ft.</p>
      <p>Lauderdale, FL, USA, 11–13 April 2011, pp. 315–323.
[57] M. Medykovskyy, P. Lipinski, O. Troyan, M. Nazarkevych, Methods of protection
document formed from latent element located by fractals, in proceedings of IEEE Xth
International Scientific and Technical Conference" Computer Sciences and Information
Technologies"(CSIT), 2015, September, pp. 70-72.
[58] D. B. Goldman, B. Curless, D. Salesin, S. M. Seitz, Schematic storyboarding for video
visualization and editing, Acm transactions on graphics (tog) 25(3) (2006) 862-871.
[59] M. U. Sreeja, B. C. Kovoor, Towards genre-specific frameworks for video
summarisation: A survey, Journal of Visual Communication and Image Representation
62 (2019) 340-358.
[60] V. N. Mandhala, D. Bhattacharyya, B. Vamsi, N. Thirupathi Rao, Object detection
using machine learning for visually impaired people, International Journal of Current
Research and Review 12(20) (2020) 157-167.
[61] N. V. Nguyen, C. Rigaud, J. C. Burie, Digital comics image indexing based on deep
learning, Journal of Imaging 4(7) (2018) 89.
[62] M. Nadeem, H. Shen, L. Choy, J. M. H. Barakat, Smart diet diary: real-Time mobile
application for food recognition, Applied System Innovation 6(2) (2023) 53.
[63] B. Van Eden, B. Rosman, An overview of robot vision, in proceedings of IEEE
Southern African Universities Power Engineering Conference/Robotics and
Mechatronics/Pattern Recognition Association of South Africa (SAUPEC / RobMech/
PRASA), 2019, January, pp. 98-104.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.S.</given-names>
            <surname>Aote</surname>
          </string-name>
          , et al.,
          <article-title>An improved deep learning method for flying object detection and recognition</article-title>
          ,
          <source>SIViP</source>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .1007/s11760-023-02703-y.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Drone-YOLO:
          <article-title>An Efficient Neural Network Method for Target Detection in Drone Images, Drones 7 (</article-title>
          <year>2023</year>
          )
          <article-title>526</article-title>
          . doi:
          <volume>10</volume>
          .3390/drones7080526.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , G. Gkioxari,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mask</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          ,
          <source>in Proceedings of the IEEE International Conference on Computer Vision</source>
          , Venice, Italy,
          <fpage>22</fpage>
          -29
          <source>October</source>
          <year>2017</year>
          , pp.
          <fpage>2961</fpage>
          -
          <lpage>2969</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <article-title>Pyramid scene parsing network</article-title>
          ,
          <source>in Proceedings of the IEEE Conference on Computer Vision</source>
          and Pattern Recognition, Honolulu,
          <string-name>
            <surname>HI</surname>
          </string-name>
          , USA,
          <fpage>21</fpage>
          -
          <issue>26</issue>
          <year>July 2017</year>
          , pp.
          <fpage>2881</fpage>
          -
          <lpage>2890</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Ronneberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brox</surname>
          </string-name>
          , U-net:
          <article-title>Convolutional networks for biomedical image segmentation, in Proceedings of the Medical Image Computing</article-title>
          and
          <string-name>
            <surname>Computer-Assisted</surname>
            <given-names>Intervention - MICCAI</given-names>
          </string-name>
          <year>2015</year>
          : 18th International Conference, Munich, Germany,
          <fpage>5</fpage>
          -
          <lpage>9</lpage>
          October 2015, Springer: Berlin/Heidelberg, Germany, pp.
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Faster</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          :
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          ,
          <source>Adv. Neural Inf. Process. Syst</source>
          .
          <volume>28</volume>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Fast</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          ,
          <source>in Proceedings of the IEEE International Conference on Computer Vision</source>
          , Santiago, Chile,
          <fpage>7</fpage>
          -
          <issue>13</issue>
          <year>December 2015</year>
          , pp.
          <fpage>1440</fpage>
          -
          <lpage>1448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Donahue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Darrell</surname>
          </string-name>
          , J. Malik,
          <article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>
          ,
          <source>in Proceedings of the Conference on Computer Vision</source>
          and Pattern Recognition, Columbus,
          <string-name>
            <surname>OH</surname>
          </string-name>
          , USA,
          <fpage>23</fpage>
          -
          <lpage>28</lpage>
          June 2014, pp.
          <fpage>580</fpage>
          -
          <lpage>587</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Van Der Maaten</surname>
          </string-name>
          , Weinberger,
          <string-name>
            <surname>K.Q.</surname>
          </string-name>
          <article-title>Densely connected convolutional networks</article-title>
          ,
          <source>in Proceedings of the IEEE Conference on Computer Vision</source>
          and Pattern Recognition, Honolulu,
          <string-name>
            <surname>HI</surname>
          </string-name>
          , USA,
          <fpage>21</fpage>
          -
          <issue>26</issue>
          <year>July 2017</year>
          , pp.
          <fpage>4700</fpage>
          -
          <lpage>4708</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in Proceedings of the IEEE Conference on Computer Vision</source>
          and Pattern Recognition, Las Vegas,
          <string-name>
            <surname>NV</surname>
          </string-name>
          , USA,
          <fpage>27</fpage>
          -
          <lpage>30</lpage>
          June 2016, pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          , Zisserman,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          ,
          <source>arXiv</source>
          <year>2014</year>
          , arXiv:
          <fpage>1409</fpage>
          .
          <fpage>1556</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>S.</given-names>
            <surname>Elfwing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Uchibe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Doya</surname>
          </string-name>
          ,
          <article-title>Sigmoid-weighted linear units for neural network function approximation in reinforcement learning</article-title>
          ,
          <source>Neural Netw</source>
          .
          <volume>107</volume>
          (
          <year>2018</year>
          )
          <fpage>3</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Smelyakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Vakulik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Filipov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kotelnykov</surname>
          </string-name>
          ,
          <article-title>Fast Color Images Clustering for Real-Time Computer Vision</article-title>
          and AI System,
          <source>CEUR Workshop Proceedings</source>
          <volume>3664</volume>
          (
          <year>2024</year>
          )
          <fpage>161</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nazarkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kostiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Oleksiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nаconechnyi</surname>
          </string-name>
          ,
          <article-title>Method of Dataset Filling and Recognition of Moving Objects in Video Sequences based on YOLO</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3654</volume>
          (
          <year>2024</year>
          )
          <fpage>265</fpage>
          -
          <lpage>276</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nazarkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kostiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Oleksiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-T.</given-names>
            <surname>Shvahuliak</surname>
          </string-name>
          ,
          <article-title>A YOLO-based Method for Object Contour Detection and Recognition in Video Sequences</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3654</volume>
          (
          <year>2024</year>
          )
          <fpage>49</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [31]
          <string-name>
            <surname>R.M. Peleshchak</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          <string-name>
            <surname>Lytvyn</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          <string-name>
            <surname>Nazarkevych</surname>
            ,
            <given-names>I.R.</given-names>
          </string-name>
          <string-name>
            <surname>Peleshchak</surname>
            ,
            <given-names>H.Y.</given-names>
          </string-name>
          <string-name>
            <surname>Nazarkevych</surname>
          </string-name>
          ,
          <source>Influence of the Symmetry Neural Network Morphology on the Mine Detection Metric, Symmetry</source>
          <volume>16</volume>
          (
          <issue>4</issue>
          ) (
          <year>2024</year>
          )
          <fpage>485</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shirokopetleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dolhanenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chupryna</surname>
          </string-name>
          , S. Smelyakov,
          <article-title>Research of Methods for Image Sharpness Evaluation in Photos of People</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3664</volume>
          (
          <year>2024</year>
          )
          <fpage>255</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Uhryn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ushenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korolenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <article-title>System programming of a disease identification model based on medical images</article-title>
          ,
          <source>in Proceedings of SPIE - The International Society for Optical Engineering</source>
          ,
          <year>2024</year>
          ,
          <volume>12938</volume>
          ,
          <year>129380F</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tsybulia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lavrut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavrut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nazarkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <article-title>Clustering Methods Analysis for Terrain Colors Characteristics Determination</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3387</volume>
          (
          <year>2023</year>
          )
          <fpage>103</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mykhailyshyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rzheuskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Semianchuk</surname>
          </string-name>
          ,
          <article-title>System Development for Video Stream Data Analyzing</article-title>
          ,
          <source>Advances in Intelligent Systems and Computing</source>
          <volume>1020</volume>
          (
          <year>2020</year>
          )
          <fpage>315</fpage>
          -
          <lpage>331</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -26474-1_
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>V.</given-names>
            <surname>Motyka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Stepaniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nasalska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <article-title>People's Emotions Analysis while Watching YouTube Videos</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3403</volume>
          (
          <year>2023</year>
          )
          <fpage>500</fpage>
          -
          <lpage>525</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>O.</given-names>
            <surname>Yakovleva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kovtunenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liubchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Honcharenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kobylin</surname>
          </string-name>
          ,
          <article-title>Face Detection for Video Surveillance-based Security System</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3403</volume>
          (
          <year>2023</year>
          )
          <fpage>69</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>K.</given-names>
            <surname>Dergachov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Krasnov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cheliadin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kazatinskij</surname>
          </string-name>
          ,
          <article-title>Video data quality improvement methods and tools development for mobile vision systems</article-title>
          ,
          <source>Advanced Information Systems</source>
          <volume>4</volume>
          (
          <issue>2</issue>
          ) (
          <year>2020</year>
          )
          <fpage>85</fpage>
          -
          <lpage>93</lpage>
          . doi:
          <volume>10</volume>
          .20998/
          <fpage>2522</fpage>
          -
          <lpage>9052</lpage>
          .
          <year>2020</year>
          .
          <volume>2</volume>
          .13.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>V.</given-names>
            <surname>Barsov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Plakhotnyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kosterna</surname>
          </string-name>
          ,
          <article-title>Research of the method of increasing the object determination accuracy on the low-resolution video stream</article-title>
          ,
          <source>Advanced Information Systems</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ) (
          <year>2021</year>
          )
          <fpage>91</fpage>
          -
          <lpage>97</lpage>
          . doi:
          <volume>10</volume>
          .20998/
          <fpage>2522</fpage>
          -
          <lpage>9052</lpage>
          .
          <year>2021</year>
          .
          <volume>2</volume>
          .12.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sabziev</surname>
          </string-name>
          ,
          <article-title>Determining the location of an unmanned aerial vehicle based on video camera images</article-title>
          ,
          <source>Advanced Information Systems</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ) (
          <year>2021</year>
          )
          <fpage>136</fpage>
          -
          <lpage>139</lpage>
          . doi:
          <volume>10</volume>
          .20998/
          <fpage>2522</fpage>
          -
          <lpage>9052</lpage>
          .
          <year>2021</year>
          .
          <volume>1</volume>
          .20.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>O.</given-names>
            <surname>Tymochko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Larin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Osiievskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Timochko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdalla</surname>
          </string-name>
          ,
          <article-title>Method of processing video information resource for aircraft navigation systems and motion control</article-title>
          ,
          <source>Advanced Information Systems</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ) (
          <year>2020</year>
          )
          <fpage>140</fpage>
          -
          <lpage>145</lpage>
          . doi:
          <volume>10</volume>
          .20998/
          <fpage>2522</fpage>
          -
          <lpage>9052</lpage>
          .
          <year>2020</year>
          .
          <volume>1</volume>
          .22.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>L.</given-names>
            <surname>Moskvych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Riepina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shcherbinin</surname>
          </string-name>
          ,
          <article-title>Application of Innovative Approaches to Video Segmentation in a Criminal Process</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>2870</volume>
          (
          <year>2021</year>
          )
          <fpage>1792</fpage>
          -
          <lpage>1805</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kovaliuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kobets</surname>
          </string-name>
          , G. Shekhet,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tielysheva</surname>
          </string-name>
          ,
          <article-title>Analysis of Streaming Video Content and Generation Relevant Contextual Advertising</article-title>
          ,
          <source>CEUR workshop proceedings 2604</source>
          (
          <year>2020</year>
          )
          <fpage>829</fpage>
          -
          <lpage>843</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>D.</given-names>
            <surname>Reynolds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Messner</surname>
          </string-name>
          ,
          <article-title>Video Copy Detection Utilizing the Log-Polar Transformation</article-title>
          ,
          <source>International Journal of Computing</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ), (
          <year>2016</year>
          )
          <fpage>8</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>M.</given-names>
            <surname>Baharon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdollah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Abu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Abidin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idris</surname>
          </string-name>
          , Secure Video Transcoding in Mobile Cloud Computing,
          <source>International Journal of Computing</source>
          <volume>17</volume>
          (
          <issue>4</issue>
          ) (
          <year>2018</year>
          )
          <fpage>208</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kochan</surname>
          </string-name>
          ,
          <string-name>
            <surname>CN-Unet</surname>
          </string-name>
          :
          <article-title>A Robust Network Based on Deep Convolution for Medical Image Segmentation</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3387</volume>
          (
          <year>2023</year>
          )
          <fpage>29</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hnatushenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kogut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Uvarov</surname>
          </string-name>
          ,
          <article-title>On Satellite Image Segmentation via Piecewise Constant Approximation of Selective Smoothed Target Mapping</article-title>
          ,
          <source>Applied Mathematics and Computation</source>
          <volume>389</volume>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          .1016/j.amc.
          <year>2020</year>
          .
          <volume>125615</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>B.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korniy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lutsyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          ,
          <article-title>Deep Learning for Atmospheric Cloud Image Segmentation</article-title>
          ,
          <source>in proceedings of IEEE 11 th International Conference on Electronics and Information Technologies</source>
          ,
          <string-name>
            <surname>ELIT</surname>
          </string-name>
          <year>2019</year>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>B.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lutsyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korniy</surname>
          </string-name>
          ,
          <article-title>Segmentation of atmospheric cloud images obtained by remote sensing</article-title>
          ,
          <source>in proceedings of International conferences on Advanced in Radioelectronics, Telecomunication and Computer Engineering TCSET</source>
          <year>2018</year>
          , pp.
          <fpage>213</fpage>
          -
          <lpage>216</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          , et al.,
          <article-title>Spatial point patterns generation on remote sensing data using convolutional neural networks with further statistical analysis</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ) (
          <year>2022</year>
          )
          <fpage>14341</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>R.J.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.P.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            <surname>Korniy</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.I. Kerod</surname>
          </string-name>
          ,
          <article-title>Image Segmentation Based on the Evaluation of the Tendency of Image Elements to form Clusters with the Help of Point Field Characteristics</article-title>
          ,
          <source>Cybernetics and Systems Analysis</source>
          <volume>51</volume>
          (
          <issue>5</issue>
          ) (
          <year>2015</year>
          )
          <fpage>704</fpage>
          -
          <lpage>713</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lutsyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <article-title>Detection of pixels corrupted by impulse noise using random point patterns</article-title>
          ,
          <source>Visual Computer</source>
          <volume>38</volume>
          (
          <issue>11</issue>
          ) (
          <year>2022</year>
          )
          <fpage>3719</fpage>
          -
          <lpage>3730</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>B.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lutsyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kosarevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Maksymyuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gazda</surname>
          </string-name>
          ,
          <article-title>Features extraction from multi-spectral remote sensing images based on multi-threshold binarization</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <issue>1</issue>
          ) (
          <year>2023</year>
          )
          <fpage>19655</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>B. T.</given-names>
            <surname>Truong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Venkatesh</surname>
          </string-name>
          ,
          <article-title>Video abstraction: A systematic review and classification</article-title>
          ,
          <source>ACM transactions on multimedia computing, communications, and applications 3</source>
          (
          <issue>1</issue>
          ), (
          <year>2007</year>
          )
          <article-title>3-es</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>J.</given-names>
            <surname>Eggermont</surname>
          </string-name>
          , et al.,
          <article-title>Optimizing computed tomographic angiography image segmentation using Fitness Based Partitioning</article-title>
          ,
          <source>Lecture Notes in Computer Science</source>
          <volume>4974</volume>
          (
          <year>2008</year>
          )
          <fpage>275</fpage>
          -
          <lpage>284</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [56]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nazarkevych</surname>
          </string-name>
          , et al.,
          <article-title>Evaluation of the Effectiveness of Different Image Skeletonization Methods in Biometric Security Systems</article-title>
          ,
          <source>International Journal of Sensors Wireless Communications and Control</source>
          <volume>11</volume>
          (
          <issue>5</issue>
          ) (
          <year>2021</year>
          )
          <fpage>542</fpage>
          -
          <lpage>552</lpage>
          . doi:
          <volume>10</volume>
          .2174/2210327910666201210151809.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>