<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Cybersecurity Providing in Information and Telecommunication Systems, February</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/ELNANO54667.2022</article-id>
      <title-group>
        <article-title>Neural Networks to Recognize Ships on Satellite Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svitlana Popereshnyak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasiya Vecherkovskaya</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liubov Ivanova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute</institution>
          ,”
          <addr-line>37, Prospect Beresteiskyi, Kyiv, 03056</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>24, Bohdana Gavrylyshyn str., Kyiv, 02000</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>28</volume>
      <issue>2024</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>In the course of the work, various digital image processing algorithms were analyzed in detail, and special attention was paid to their use in solving the actual problem of recognizing ships on satellite images. Significant results were achieved in this area through the development of software that allows for solving high-precision recognition tasks based on a specially designed and trained convolutional network. The software development stage also included experimental studies and performance evaluation of the developed solution, which allows us to objectively determine the efficiency and potential capabilities in the context of a particular task. The implementation of the results obtained can help improve the quality and speed of ship recognition on satellite images, which is of great importance in various fields, including maritime and environmental monitoring. The general methodology and developed algorithms can also be applied in other areas of image processing and computer vision. As a result of our research, we are confident in the effectiveness and prospects of using the obtained developments to solve specific problems of object recognition on large volumes of satellite images.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Neural networks</kwd>
        <kwd>machine learning</kwd>
        <kwd>software</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Every year, a huge amount of parallel work and
data that needs to be processed in a certain
way appears in various fields. These tasks are
of the same type, repetitive, or require
constant human concentration to control or
search for an object. Various software products
and machine learning algorithms are used to
simplify people’s work and reduce time.</p>
      <p>Today, machine learning algorithms are in
active use</p>
      <p>Many scientists have studied the use of
neural networks to solve image recognition
problems. The paper [1] reviews the main
methods for solving computer vision problems
of classification, segmentation, and image
processing implemented in CV systems.</p>
      <p>In [2], the convolutional properties of an
autoencoding neural network for object
detection in an image are considered. For
training and testing, datasets were generated
in the form of two-dimensional images with
three color channels.</p>
      <p>Ship images have been studied by scientists
from all over the world, in particular, [3]
presents a sequence of image processing
algorithms suitable for detecting and
classifying ships from nadir panchromatic
electro-optical imagery. In [4], an algorithm
was developed to classify ships according to
size using image processing. The image of the
ship was captured by a stationary camera.</p>
      <p>Classification and segmentation of ships by
analyzing satellite images will help in
searching for objects without human
intervention, because there are many seas and
oceans, and people will not need to look
through every square kilometer to find it. This
will help to find objects faster and reduce the
cost of human labor. This will help to control
the delivery time of certain ships carrying
cargo. Or to control water borders so that an
unwanted object does not cross certain
boundaries. These methods will also help to
control and locate enemy ships [5–7].</p>
      <p>The work aims to study algorithms and
develop software for searching for ships in a
satellite image of the sea area [8, 9].</p>
    </sec>
    <sec id="sec-2">
      <title>2. General Overview of Algorithms and their Comparison</title>
      <p>Convolutional neural networks can be used to
classify and segment digital images, which are
specially designed to handle large amounts of
data such as images. Convolutional neural
networks use convolution to detect local
features in images and pooling to reduce the
dimensionality of the image. They can be
successfully used to classify objects in an image
and segment the image into separate clusters.</p>
      <p>Convolutional neural networks are a type of
neural network commonly used for image and
video processing. These networks are used to
automatically detect image features and
characteristics such as borders, shapes, and
textures.</p>
      <p>The main difference between convolutional
neural networks and fully connected networks
is that convolutional networks use
convolutions to process input images instead
of treating each pixel of an image as a separate
input (Fig. 1).
In Fig. 1, the bottom square is the input photo.
And the top one is the new look of the photo
after going through the convolution (kernel).
And the lines connected between the squares
are the convolution that transforms the
objects. In other words, with the help of
convolution, we reduce the dimensionality of a
digital image for a particular purpose, which
allows us to keep important features and
discard unimportant ones.</p>
      <p>Usually, one convolution is used more than
once. To make it easier to understand the
architecture of a neural network, we use the
word neural layer, which is a certain stage of a
neural network.</p>
      <p>Convolutions are filters that slide over the
input image and perform multiplication and
summation operations to create a feature map
(Fig. 2).
where   is the transposed convolution,  is
the part of the image that is highlighted by the
convolution range, and  —is the value that the
artificial model is looking for.</p>
      <p>The main advantages of convolutional
networks are that they can automatically
detect and utilize local features in an image,
which reduces the number of parameters that
need to be trained and provides faster and
more efficient performance. One of the
advantages of convolutional neural networks
is that they can effectively recognize local
features in images, such as corners, edges, and
textures, reducing the number of parameters
and computational complexity compared to
fully connected neural networks. In addition,
convolutional neural networks can
automatically learn useful features, reducing
the need to manually select features to use.</p>
      <p>Some of the disadvantages of convolutional
networks include high computational
complexity and the ability to overlearn training
data. Also, convolutional networks can have a
complex architecture, which can make them
difficult to understand and develop. The
disadvantages of convolutional neural
networks are the requirement for a large
amount of data for training, as well as the
difficulty of understanding and interpreting
them. In addition, convolutional neural
networks can tend to overlearn, especially
when there is not enough data to train.</p>
      <p>Compared to fully connected networks,
convolutional networks are usually better for
image processing because of their specialized
filters that can detect different types of
features. This reduces the number of
parameters that need to be trained and
improves training efficiency.</p>
      <p>You can also use autoencoders to segment
images and obtain embeddings for image
classification. In addition, contrastive learning
can be used to train embeddings that can be
used for image classification.</p>
      <p>All of these methods can be successfully
used for image classification and
segmentation, depending on the specifics of
the task and the availability of data.</p>
      <p>Below is a comparative table of different
algorithms (Table 1):
After researching the types of algorithms for
classification and segmentation, since neural
networks can be used as a designer, we chose
the convolutional neural network algorithm as
the basis for segmenting the image into 2
classes - ship and non-ship. To do this, we need
to create an auto-encoder that will reduce the
size of the input image and then increase it,
leaving only ships in the image. It performs the
task of this work best, there is also a large
number of digital images, and a GPU
accelerator will be used to train the algorithm,
which will speed up the cloud solutions
learning process many times over.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Input Data Analysis</title>
      <p>The main type of input data is a digital photo
taken from a satellite, they are in jpg format.
There is also another type of data—masks, and
coordinates of ships on the photos, if they are
there, then their format is CSV. The first type of
data, photographs, looks like this (Fig. 3):
These photos show that the size of ships and
their number vary. There are also variants of
photos where there are no ships, but there are
other objects, such as islands, piers, or a part of
the land. And the photos show that the color of
the water is different from each other.
All photos are 768 by 768, but the program also
has a case that converts another size to this
one. The data is displayed as masks (Fig. 4).
them, so different methods need to be used. In
this process, we first took a smaller number of
photos with no ships by about 2 times and
generated a large number of artificial images
on those digital images that have ships.
The total number of objects to be studied, is
192555 photos, but the program artificially
creates additional images from the initial set,
for example, it takes one image and rotates it
vertically or horizontally, and this adds more
objects, but it may be that the model will
overlearn on this set, so you need to
immediately select the number of photos that
the model has never seen. Therefore, in
addition to 192555 objects, we took 15600
more to test the model.</p>
      <p>When training a model, the distribution of
objects is very important. There are cases
when there is an imbalance of classes, as in our
case (Fig. 5). This is critical because it is easier
for the model not to find ships than to find
The number of digital images with ships is
42556. The number of digital images without
ships is 149999. The number of ships in a
digital image is also important (Fig. 6). Since
the model has to understand what data it is
working with, it will adapt to the data for
training. Fig. 6 shows that most digital photos
have one ship each and the distribution is
similar to a logarithmic distribution. Since the
program will generate a large number of
artificial images, this distribution is not a
problem.
To understand how to use masks, you can look
at Fig. 7:
In this case, the yellow translucent color was
chosen to show the location of the ships—their
coordinates. The segmentation task in the
program is to search for ships in a digital image
and select them—to find the coordinates, i.e.
masks.</p>
      <p>If we look at Fig. 8, we can see that the
distribution is exponential in the test set of
photos, and most of the photos have one ship.
However, the rest of the number gradually and
smoothly decreases, which allows us to
correctly estimate the model on the test set.
Data processing in machine learning is an
important process, as proper processing can
boost the model’s results very highly, and
incorrect processing can greatly reduce the
results and quality of the model. However, a
common practice for working with images is to
artificially create additional images based on
the input ones.</p>
      <p>Below is a description that creates the
AirbusDataset class, which allows you to use
various methods to generate images:
The input to the class object is 3 parameters:
1. In_df is an array with a link to photos,
and masks.
2. Transform is an array of functions or
objects of the processing class, they are
given below.
3. Mode has 2 modes ‘train’ and
‘validation’, when you select one of them,
it will generate new photos to the
database with a specific sample.</p>
      <p>As an output, the class will generate new
objects to the database—new photos.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Description of the Metrics Used</title>
      <p>The choice of metric is also important, as the
metric is involved in model training. There are
always two important problems in such tasks.</p>
      <p>It is necessary to choose either the model
will use the most accurate classification, but
this may lead to the fact that the model will
sometimes classify islands, waves, or other
objects as ships. Or choose the other option
when the cost of error is critical and the model
will repeatedly not classify different objects,
but this will result in very small ships (boats)
not being found or not all points being classified
that combines their advantages. It is used to
minimize errors when training a semantic
  =  
step 2.
in the same ship group.</p>
      <p>For this work, we chose the second type,
because this program should help users, so we
don’t want to distract them with unnecessary
noise. Therefore, the BCEJaccardWithLogitsLoss
metric was chosen.</p>
      <p>The BCEJaccardWithLogitsLoss metric is a
combination of two</p>
      <p>metrics: Binary
CrossEntropy (BCE) and Jaccard coefficient, which
are used to evaluate the quality of binary
semantic segmentation. This metric is a good
fit because we need to segment whether a pixel
is a ship or not.</p>
      <p>BCE measures the correspondence between
the predicted and true pixel values of an image. It
does this by calculating the cross-entropy
between
the
predicted
and
true
pixel
distributions. The higher the BCE, the less
accurate the prediction is. The general algorithm
for calculating the BCE metric:
1. Select the initial vector ор  (0) ; take  = 1.
2. Next, you need to generate a random
sample,  1 … , … ,   з  ( ;  ( −1)).</p>
      <p>3. Solve for  ( ) , where</p>
      <p>1
  =1
∑  ( )</p>
      <p>( ;  )
 ( ;  ( −1)) ) 
 (  ;  )
(2)</p>
      <sec id="sec-4-1">
        <title>When convergence is achieved, stop, otherwise, t is increased by 2, and proceed to The</title>
        <p>Jaccard
coefficient
measures the
similarity between the predicted and true
values. This is done by calculating the overlap
area between the sets of pixels corresponding
to the predicted and true regions in the image.
The higher the Jaccard coefficient, the more
accurate the prediction is.</p>
        <p>The</p>
        <p>Jaccard
coefficient
measures the
similarity between sets and is defined as the
measure of the common part divided by the
measure of the union of the sets:
| ∩  |
| ∩  |
 ( ,  ) = | ∪  |</p>
        <p>= | | + | | − | ∩  |
0 ≤  ( ,  ) ≤ 1
(3)
is a predicted value and a real
where А і 
value.
then ( ,  ) = 0.</p>
      </sec>
      <sec id="sec-4-2">
        <title>When</title>
        <p>and 
are
both
empty,</p>
        <p>The Jaccard coefficient is also used to find
similar texts in a large corpus of documents.
BCEJaccardWithLogitsLoss combines BCE and
Jaccard coefficients into a single loss function
segmentation model.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Software and Neural Network</title>
    </sec>
    <sec id="sec-6">
      <title>Architecture</title>
      <p>Creating the architecture is an important step
because
the
number
of layers
is
very
important. If there are a lot of layers, the model
will take a long time to learn and overlearn on
test data, and if there are few layers, the model
may not learn and perform poorly on test data.</p>
      <p>The architecture consists of 2 main parts for
auto-encoding. The first is an encoder, i.e.,
reducing the dimensionality of the object, for
which the down_block class was created. The
other part is the inverse of the previous one.
That is, you need to expand a small object to
the size of the input image—this is the decoder
process, which allows you to create a new
object—a matrix of 0 and 1. Where 1 is the
location of the ship on the map, and 0 is a
segmented non-ship in the digital image.</p>
      <p>To create a neural network, you need to
create a connection and combine the previous
blocks correctly. The NN_Ship_Detection class is
responsible for this.</p>
      <p>You can also see its basic appearance in
Fig. 10.</p>
      <p>down_block1 is the image is added and
the block increases the number of
filters from 3 to 16.
down_block2 increases the number of
filters from 16 to 32.
3. down_block3 increases the number of
filters from 32 to 64.
4. down_block4 increases the number of
filters from 64 to 128.
5. down_block5 increases the number of
filters from 128 to 256.
6. down_block6 increases the number of
filters from 256 to 512.
7. down_block7 increases the number of
filters from 512 to 1024.
8. Then normalize again.
9. up_block1 reduces the number of
filters from 1024 to 512.
10. up_block2 reduces the number of
filters from 512 to 256.
11. up_block3 reduces the number of
filters from 256 to 128.
12. up_block4 reduces the number of
filters from 128 to 64.
13. up_block5 reduces the number of
filters from 64 to 32.
14. up_block6 reduces the number of
filters from 32 to 3.
15. last_conv2 mixes the number of filters
from 3 to 1, where the result is a matrix
with zeros and ones.</p>
      <p>So the main job of this neural network is to
reduce the size of the image so that only the
information about the location of the ships
remains, and the rest of the information is
discarded. Thus, as a result, we get a new
object—a matrix with zeros and ones.</p>
    </sec>
    <sec id="sec-7">
      <title>6. General Discussion</title>
    </sec>
    <sec id="sec-8">
      <title>Evaluation of the Results and</title>
      <p>Since this system is based on the creation of a
neural network, the main part of testing will be
the algorithm itself. Testing neural networks is
an important part of their development and use.
It is the process of evaluating the quality of a
training model on independent test data to
confirm that it works properly.</p>
      <p>Backpropagation testing is the process of
testing the performance of a machine learning
model using a dataset that the model has not
seen before.</p>
      <p>To perform testing on a deferred dataset, you
first need to divide the total dataset into training,
validation, and test samples. The training set is
used to train the model, the validation set is used
to adjust hyperparameters and evaluate the
model performance on unknown data, and the
test set is used to finally evaluate the model
performance.</p>
      <p>After training the model, testing on a deferred
dataset is performed on a test set that consists of
data that the model did not see during training.
The purpose of testing is to evaluate the model’s
performance on new data that the model has not
seen before. The results of the testing can be used
to make decisions about using the model in
realworld applications, such as production tasks or
research.</p>
      <p>When testing on a deferred dataset, it is
important to keep in mind that the test results
may be dependent on the composition of the test
sample. If the test sample does not represent
diverse data, you may experience a carryover
problem where the model performs well on a
50test sample but performs poorly on real data. To
avoid this problem, you should use a test sample
that represents as much diversity as possible.</p>
      <p>After evaluating the results on the deferred
set, you can determine how well the model can
generalize its knowledge to new data. If the
results on the deferred dataset are poor and the
results on the training dataset are very good,
then this may indicate that the model is
overtrained on the training dataset.</p>
      <p>It is also important to keep in mind that the
resulting metrics on the deferred dataset may be
slightly worse than on the training dataset, as the
model has not had the opportunity to learn from
this data and they are assigned the role of
evaluating the overall performance of the model.
However, if the difference between the results on
the training and deferred datasets is significant,
it may be a sign of the model’s poor ability to
generalize its knowledge to new data.</p>
      <p>Testing on a deferred dataset is an important
step in the process of developing neural
networks, as it allows you to check the overall
performance of the model on new, previously
unseen data.</p>
      <p>Evaluation is an important step to check
whether the model is working well and whether
there are moments when the model is
overlearning or underlearning. It is important to
look at each iteration and choose the best one.</p>
      <p>Using the testing method described above, we
checked the quality of the image model
classification. Fig. 11 shows the training process.
The data selected for validation shows a good
result. However, the graph shows that there is a
slight overfitting. Overfitting is an unpleasant
moment when the model memorizes the data it
has been trained on well but performs poorly on
new data. But as you can see on the graph, the
overfitting is not critical and the model copes
well with the result.</p>
      <p>To better understand the model’s estimates,
we need to look at the table (Table 2):</p>
    </sec>
    <sec id="sec-9">
      <title>7. Conclusions</title>
      <p>This paper studies modern methods of
working with digital images. This gives an
understanding of working with artificial
intelligence, namely neural networks. What
types of them are there, what they are used for,
their advantages and disadvantages for
maximum quality of digital image
segmentation?</p>
      <p>The architecture of the neural network was
created. For its construction, convolutional
neural networks in the autoencoder system, a
method for encoding and decoding were used.
The main function of this network is to obtain
a digital image and convert it to a matrix form,
where the elements have values of 0 and 1,
where 0 is not a ship and 1 is a ship. We also
analyzed the architecture components:
advantages and disadvantages, metrics,
activation functions, and the data the model
receives at the input and output. Using this
architecture, we created software for
recognizing ships in digital images.</p>
      <p>We went through all the main stages of
working with the model. The first step was to
search for data, which allowed us to choose the
right methods for working with images.</p>
      <p>The next step was to analyze the data to see
the features of the dataset. It was found that
there was an imbalance of classes in the set,
where there were many more photos without
ships than photos with ships. Then it was
decided to artificially create data—digital
images that had at least one ship were flipped
horizontally and vertically, which made it
possible to increase the number of digital
images with a ship by 3 times.</p>
      <p>Then the data was processed, and an
additional number of images with ships were
created. This set can also be used in other tasks.</p>
      <p>The next step was to create a neural
network architecture using convolutional
methods and encoding and decoding—
autoencoding. For this purpose, we used the
Python programming language and the
PyTorch library. 54 This library is designed to
work with neural networks and makes it
possible to use GPU technology.</p>
      <p>The last stage in the program’s
development was model training and quality
assessment. The model was trained for 2600
iterations. The trained model can be used in
other tasks and programs.</p>
      <p>To measure the quality, we chose the metric
for segmentation—BCEJaccardWithLogitsLoss.
To qualitatively measure the model, we chose the
method of splitting the data into training, test,
and validation data. In the final testing of the test
data, it was found that the trained model was
wrong in 1.5% of cases from the correct value,
which is a good result.</p>
      <p>A possible improvement of the architecture
is to turn the task into a multi-segmentation
one, where the model can show not only the
location of the ship but also its type, for
example, cargo, tourist, military, etc.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>O.</given-names>
            <surname>Zinchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. Zvenigorodsky T.</given-names>
            <surname>Kisil</surname>
          </string-name>
          ,
          <article-title>Convolutional Neural Networks for Solving Computer Vision Problems</article-title>
          ,
          <source>Telecommunication and Information Technologies</source>
          <volume>2</volume>
          (
          <issue>75</issue>
          ) (
          <year>2022</year>
          )
          <fpage>4</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .31673/
          <fpage>2412</fpage>
          -
          <lpage>4338</lpage>
          .
          <year>2022</year>
          .020411
          <string-name>
            <given-names>L.</given-names>
            <surname>Yasenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Klyatchenko</surname>
          </string-name>
          ,
          <source>Convolutional Neural Network Properties Based on an Autoencoder</source>
          , Inf. Technol.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Comput</surname>
          </string-name>
          . Eng.
          <volume>52</volume>
          (
          <issue>3</issue>
          ) (
          <year>2021</year>
          )
          <fpage>77</fpage>
          -
          <lpage>85</lpage>
          . doi:
          <volume>10</volume>
          .31649/1999-9941-2021-52-3-
          <fpage>77</fpage>
          -85.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Buck</surname>
          </string-name>
          , et al.,
          <source>Ship Detection and Classification from Overhead Imagery, Proc. SPIE 6696</source>
          ,
          <article-title>Applications of Digital Image Processing XXX,</article-title>
          <year>66961C</year>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>doi: 10.1117/12</source>
          .754019.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>G.</given-names>
            <surname>Santhalia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <source>Safer Navigation of Ships by Image Processing &amp; Neural Network, Second Asia International Conference on Modelling &amp; Simulation (AMS)</source>
          (
          <year>2008</year>
          )
          <fpage>660</fpage>
          -
          <lpage>665</lpage>
          . doi:
          <volume>10</volume>
          .1109/AMS.
          <year>2008</year>
          .
          <volume>48</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          Theor. Appl. Inf. Technol.
          <volume>100</volume>
          (
          <issue>24</issue>
          ) (
          <year>2022</year>
          )
          <fpage>7390</fpage>
          -
          <lpage>7404</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Khorolska</surname>
          </string-name>
          , et al.,
          <article-title>Application of a Convolutional Neural Network with a Module of Elementary Graphic Primitive Classifiers in the Problems of Recognition of Drawing Documentation and Transformation of 2D to 3D Models,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>