<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Wavelet transform based optimization method for Three- Dimensional computer vision</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svetlana Antoshchuk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Galina Shcherbakova</string-name>
          <email>galina.sherbakova@op.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergey Kondratyev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Usov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daria Koshutina</string-name>
          <email>d.v.koshutina@op.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Odessa Polytechnic National University</institution>
          ,
          <addr-line>Shevchenko Ave., 1, Odessa, 65044</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work aims to develop approximate depth estimation methods for three-dimensional computer vision in unmanned vehicles (UVs) based on a developed wavelet transformation-based optimization method. This method can assess informative features for matching and/or optimizing costs in image analysis, refining discrepancies, and more. Possible solutions are demonstrated for obtaining an approximate depth map by simplifying the calculation of disparity values, traditionally used for forming a depth map in intensity space and using edge description with adjustable detail based on wavelet transformation. The advantage of the developed optimization method over existing algorithms in the wavelet space is the increased speed due to the rational selection of the Haar wavelet support length in the extremum search area. Modeling confirmed the effectiveness of the proposed approach for constructing depth maps and allowed for recommending the proposed method for unmanned vehicles operating under limited computational and energy resources.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Three-dimensional computer vision</kwd>
        <kwd>contour detection</kwd>
        <kwd>wavelet transform</kwd>
        <kwd>optimization</kwd>
        <kwd>image analysis 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The primary source of information about the surrounding environment for solving a multitude of
tasks in various fields such as unmanned vehicles (quadcopters, unmanned cars), robotics, and vision
for the visually impaired (UVs) is video cameras [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Recent advancements in video sensor
technologies, as well as image and video analysis and processing methods, have enabled machine
vision to confidently advance into the realms of automation and security systems, including
household, industrial, and military applications.
      </p>
      <p>Modern machine vision systems not only extract video information but also present it in a form
and quantity that allows for the identification of significant, informative features of objects and
processes. These systems recognize and assess their position and condition, select rational
trajectories for UV movement, and, if necessary, generate corresponding control actions.</p>
      <p>A distinctive feature of machine vision systems is not only the identification and recognition of
an object and its position but also the consideration of the scene and object depth, i.e., the
transformation of a two-dimensional image into a three-dimensional one, where information about
the object is represented not just in units of brightness but with pixel/distance parameters. The
analysis of the evolution of three-dimensional computer vision and the corresponding hardware and
software designed for obtaining images and forming depth parameters has shown that the main
problem is the high cost of the sensors and image processors used, making them largely inaccessible
to the broader consumer market. Therefore, finding solutions for obtaining 3D depth information
about objects in an image characterized by low cost, low power consumption, and acceptable</p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
performance is an important and relevant scientific and practical task for a wide range of the
aforementioned UVs.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Leveraging LiDAR and Stereo Cameras for Efficient Depth</title>
    </sec>
    <sec id="sec-3">
      <title>Information Retrieval in UVs</title>
      <p>
        When solving navigation tasks for UVs or collecting depth information about objects for them, the
acquisition of this information can be redistributed and/or transferred to backup subsystems and/or
systems operating on other physical principles [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For instance, one such approach could be the
combination of data from LiDAR (laser sensors) and video sensors. Currently, the LiDAR-SLAM
technology is actively being researched by Suzuki [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Lee, Savkin, and Vuletic [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Tripicchio et al.
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Bi et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Mansouri et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Vong et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and others.
      </p>
      <p>These researchers have demonstrated the feasibility of using LiDAR for path planning, collision
avoidance, orientation, and navigation.</p>
      <p>
        However, image sensors (video cameras) are characterized by their lightweight, potential
decrease in accuracy under low lighting conditions, and loss (absence) of depth information, which
increases computational load during data processing. LiDAR allows for direct distance measurement,
but the result depends on the reflective properties of the surface and is energy-intensive [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Therefore, a less energy-intensive, interesting, and promising solution for obtaining depth
information about objects and/or scenes in an image is the stereoscopic vision approach. This
approach involves the use of two video cameras with known optical characteristics and parameters
of their relative positioning [
        <xref ref-type="bibr" rid="ref4 ref7 ref8">4, 7, 8</xref>
        ]. Typically, video cameras with similar optical characteristics are
used. These cameras are directed in the same direction, with the distance between the optical centers
of the cameras being much smaller than the distance to the observed objects. Depth information is
extracted by comparing and analyzing this pair of images. This approach allows for modeling image
processing similarly to how it occurs in the vision of many living beings. Just as a given control point
in space is located in different positions in each human eye, this allows the system to calculate the
position of this point in space. As a result, a height map is obtained where objects closer to the
observer are displayed in lighter shades, and those further away are displayed in darker shades.
Several solutions enable the creation of depth maps using local block matching (StereoBM function)
or semi-global block matching (StereoSGBM function) [
        <xref ref-type="bibr" rid="ref10 ref11">10-11</xref>
        ].
      </p>
      <p>Let's consider the method for constructing a depth map, implemented by the StereoBM function.
The algorithm involves the following stages of data processing:
1. Loading images from both video cameras (often using the OpenCV library).
2. Converting images from color to grayscale.
3. Creating a StereoBM object with specified parameters.
4. Calculating the disparity map (the magnitude of the difference in the localization of
corresponding image elements obtained from the right and left cameras). Traditionally, this
is done by matching pixels in the left and right images to calculate a disparity map.
5. Outputting and saving the disparity map.</p>
      <p>In the StereoBM algorithm, stages 1, 2, 3, and 5 are standard and do not provide opportunities to
increase execution speed and thereby reduce energy consumption. Therefore, stage 4, which is aimed
at calculating the disparity map, is more promising. This stage involves the following steps:
1. Selecting blocks. Blocks of a fixed size are highlighted in the image.
2. Comparing blocks. For each block found in the left image, the algorithm tries to find a
matching block in the right image.
3. Calculating disparity. For each block in the left image, a displacement (disparity) is calculated,
indicating how many pixels the corresponding right image is shifted. This disparity is stored
in a disparity map.</p>
      <p>Normalizing the disparity map. The resulting disparity map may contain values other than
depth. Therefore, it is normalized so that the values are within [0, 255] for convenient display.</p>
      <p>
        In this part of the algorithm, the most complex and crucial for the quality of the depth map is the
block comparison stage. The block comparison algorithm searches for a pattern on the left frame
corresponding to a block on the right frame. This is done by algorithmically moving along the
horizontal axis of the right image and comparing it with the corresponding block of the left image.
Various approaches can be used to assess the degree of block matching [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref14">10-14</xref>
        ]:
1. Based on assessing the proximity of pixel intensity values in blocks.
2. Based on assessing the minimum sums of squared differences of intensities (Sum of Squared
      </p>
      <p>Differences, SSD).
3. Based on determining the minimum normalized cross-correlation coefficient (Normalized</p>
      <p>Cross-Correlation).
4. Based on the Semi-Global Matching (SGM) method, which takes into account global and local
features of low-quality images with different intensities or difficult or complex image
acquisition conditions.
5. Based on adaptive window matching. Instead of a fixed block size, algorithms can use
adaptive window mapping to account for changes in texture and intensities in images.
6. Based on the Graph Cuts method, which uses graph structures to calculate the disparity map.</p>
      <p>These algorithms can take global and local image properties into account and improve
accuracy in challenging imaging conditions (e.g., low-light conditions).
7. Based on machine learning methods. These methods rely on machine learning models trained
to predict disparity based on feature detection in images. They can handle images with
uneven texture and variations in illumination.
8. Based on deep learning methods such as Convolutional Neural Networks (CNN). These
neural networks can be trained on large datasets and perform well in complex lighting
conditions.
9. Joint use of sensors implemented on various physical principles, such as LiDAR or infrared
cameras, to improve the accuracy of the disparity map.</p>
      <p>
        Note that the approaches are numbered and sorted by the degree of increase in computational
costs. Approaches N2 to N6 require high energy consumption, which may be critically unacceptable
for BTS. Approaches N7 and N8 rely on neural network calculations, requiring computational
resources and a training "database," which can be challenging to provide for some BTS applications.
Approaches involving the joint use of sensors implemented on different physical principles (e.g., N9)
lead to increased energy consumption. Therefore, for further research, the method of comparing
image blocks in the right and left images (N1) was chosen as the base method, which can ensure the
speed of calculations. However, this method is characterized by low noise immunity, determining its
low quality. This property is especially noticeable in changing light conditions and/or low light
conditions, often encountered in applied problems solved by BTS. Additionally, the analysis allowed
us to establish the following. It should be noted that when developing algorithms for obtaining depth
information about objects and/or scenes, two assumptions about the structure of the observed scene
are used [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The first assumption considers that neighboring pixels have similar disparity values
(the difference between the positions of an object in the left and right images) because the scene
objects are piecewise smooth. The second assumption is based on the idea that homogeneous areas
of the scene correspond to flat surfaces in 3D. The algorithm for constructing a depth map involves
measuring the displacement along the x-axis of each point in the right frame and correlating it with
the corresponding point in the left frame [
        <xref ref-type="bibr" rid="ref15 ref9">9, 15</xref>
        ]. The search for the corresponding point is strictly
along the horizontal line of each frame.
      </p>
      <p>
        Therefore, to correctly determine the distance to objects during vertical calibration (along the
yaxis in the image), the position of the cameras is set such that the horizontal lines of both cameras
coincide. Horizontal calibration (along the x-axis in the image) is done by rotating the cameras
relative to each other until the x-coordinates of points at a distance of more than 10 meters coincide.
Consequently, this method requires pixel-precise positioning of the cameras both horizontally and
vertically. This complicates alignment and reduces the quality of object positioning [
        <xref ref-type="bibr" rid="ref12 ref9">9, 12</xref>
        ]. A
significant argument in favor of using stereoscopic systems in UVs is the availability of numerous
open-source algorithms in the OpenCV computer vision library (the StereoBM class exists) for
implementing stereo vision in languages such as C/C++, Python, Java, Ruby, Matlab, and others [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
The main disadvantages of stereoscopic systems for constructing depth maps for mobile navigation
systems and UVs include the following [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]:
●
●
●
      </p>
      <p>The need for camera calibration. It should be noted that even with perfectly accurate camera
positioning, obtaining a depth map in these systems (e.g., when avoiding obstacles) poses
difficulties due to the necessity of pixel-by-pixel image matching.</p>
      <p>Dependence on the quality of the initial images and/or incorrectly set camera parameters,
lighting, and illumination for each camera.</p>
      <p>Dependence of the number of computational operations and, consequently, performance on
the size and quality of the image.</p>
      <p>
        As the above analysis showed, a number of methods based on the assessment of local and global
image characteristics have been developed to evaluate disparity. Algorithms based on estimating
local image characteristics are fast but low in accuracy, while global methods can improve accuracy
and (often) noise immunity, but are characterized by low performance. The search for a compromise
between accuracy and speed can be aimed at developing methods for assessing disparity based on
optimization algorithms. When assessing disparity based on changes in the intensity of the pixels in
an image line, the coordinates of the extremum with a signal-to-noise ratio below 6 are estimated.
Known methods for searching for an extremum based on estimating the value of the first derivative,
which are traditionally used at this stage, do not work when the signal-to-noise ratio is less than 15.
To search for the minimum under such conditions, the authors developed and investigated an
optimization method based on the wavelet transform (WT) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. However, due to the large number
of stages using several types of wavelet functions, this method is characterized by low performance.
Therefore, improving the method to increase its performance for the application task described above
is relevant.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Positions Assessment of Depth Information for Object Positions</title>
      <p>To reduce the impact of these disadvantages when constructing depth maps from stereo images, an
approximate approach to depth information assessment is proposed. This approach considers the
results of identifying areas of similar intensity or object contours in the images or points of maximum
curvature in the contour description.</p>
      <p>Figure 1 illustrates the main steps of in-depth assessment by identifying areas of similar intensity
in a test image (see Figure 1, a). The proposed approach involves the following:
●
●
●
●</p>
      <p>For a segment of the intensity row, determine the region of maximum intensity on the image
from one of the cameras, for example, the left one (1) (see Figure 1, b).</p>
      <p>Create a brightness envelope template: the intensity value corresponding to the extremum
coordinate of the row and the intensity values of neighboring pixels (Figure 1, c), and locate
a similar fragment in the image from the right camera (2).</p>
      <p>Determine the distance between similar elements and the corresponding disparity value.
Processing is performed row-wise and column-wise. If necessary, the results are combined
using a logical OR scheme.</p>
      <p>To illustrate, brightness envelope templates representing a block of 3 pixels are provided for the
left and right frames of the image (see Figure 1, c). The templates show the intensity distribution
around the extremum in the given image row for the left and right cameras (see Figure 1, b).</p>
      <p>
        Both templates have a "peak" shape (other possible shapes include "trough," "plateau," "rise," "fall,"
etc.). Black squares highlight the brightness levels of each pixel in the blocks, while gray areas
indicate the brightness dispersion boundaries. The disparity between blocks is 4 pixels along the
Xaxis, and the absolute difference in brightness is 2. In real stereo images, the brightness difference
can reach up to 30. This approach is simpler than the well-known StereoBM method [
        <xref ref-type="bibr" rid="ref10 ref12 ref13 ref14">10, 12 - 14</xref>
        ],
which uses two computationally intensive procedures for similar searches: Sum of Squared
Differences (SSD) and Normalized Cross-Correlation (NCC), both of which also exhibit lower
performance. The computational complexity of finding templates can be reduced if object contours
are used in the calculation of disparity. It should be noted that contours are the most informative
part of object images, and their analysis can significantly reduce the number of computational
operations [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        However, it should be noted that in this case, the quality of the depth map depends on the
effectiveness of the contour detection methods used, among which the most noise-resistant methods
are those using the wavelet transform (WT) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. To extract the contours of objects using the wavelet
transform, the summing component of the approximation (vertical) of the smallest scale (a1) of the
discrete wavelet transform is removed by setting its values to zero. Calculation of the inverse
transformation using the detailing component (d1) leads to the selection of contours (horizontally)
in the image thus reconstructed. In a similar way, the vertical boundary of an object can be selected.
The procedure for identifying contours using the wavelet transform can be presented as follows:
f0 1 (1)
d1
where f is the row (column) of the original image; a1 summing component of the wavelet
transform; d1 detailing component of the wavelet transform; f ′ row (column) of the image after
the inverse transformation; K operator for obtaining a contour preparation.
      </p>
      <p>
        Due to the frequency-selective properties of the WT, the ratio of the intensity of the contour of
an object having a smaller size to the intensity of the contour having a larger size decrease as the
scale of the WT increases [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. This pattern makes this technique insensitive to changes in the
intensity of the object and allows you to take its size into account (adjust the detail). This scheme is
proposed to be used to estimate depth by identifying the contours of objects with adjustable detail
in the wavelet transform space (see Figure 2):
●
●
●
●
      </p>
      <p>The rows of the left frame image (see Figure 2, b) and the right frame image are convolved
with wavelet functions of a given scale and contour areas are detected (see Figure 2, c);
The distance between object contours and the corresponding disparity value are determined;
Processing is performed along rows and columns as needed. Results are combined using a
logical OR scheme;
If necessary, the process is repeated for different scales.</p>
      <p>a</p>
      <p>Research has also shown that with a wavelet function support length s = 1, both fine and coarse
image details result in peaks of similar amplitude in the intensity drop area after processing.</p>
      <p>
        As the scale increases, the relative size of the peaks for fine details decreases, while the amplitude
of the peak at the boundaries of large-scale objects increases [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Furthermore, the proposed methodology allows for adjusting the level of image detail, ensuring
high noise resilience, resolution of object contour extraction, and providing an approximate depth
information assessment that meets the requirements for a range of tasks.</p>
      <p>Thus, depth map construction methods can be categorized into local and global groups. Local
method-based algorithms are characterized by their speed and lower accuracy, while global methods
enhance accuracy but are slower.</p>
      <p>To find a compromise between accuracy and performance, a group of methods based on
optimization algorithms is actively being developed.</p>
      <p>
        These methods can be used for evaluating informative features, matching fragments, and/or
optimizing costs in image analysis, as well as refining discrepancies and other related tasks [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Development of the Optimization Method</title>
      <p>
        Analysis has shown that the use of global optimization methods based on exhaustive search leads to
high computational costs, while local methods based on first and second derivative estimates exhibit
low noise resilience [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ].
      </p>
      <p>
        Therefore, for solving a range of optimization problems in image analysis for UVs, the wavelet
transformation (WT) - based approach has been developed, which enhances noise resilience while
providing sufficient accuracy for these tasks. In known optimization methods using WT, the
property of gradient estimation and processing with WT changes sign as it approaches an extremum
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Continuous wavelet transformation is defined by convolution:
 ( ,  0) =
√| | −∞
1</p>
      <p>+∞
∫
 ( )Ψ∗ ( −  0)  ,</p>
      <p>ℎ =</p>
      <p>∫

1</p>
      <p>( )

 −</p>
      <p>   
√
factor</p>
      <p>1 introduced for normalization.
partial wavelet transform:</p>
      <p>where  is the function being tra0nsformed (analyzed); Ψ , 0( )is a two-parameter basis function
derived from the mother wavelet Ψ0( ) through scaling with a scaling factor  ∈  + and translation
with a parameter  0 ∈  ; and Ψ∗ is the complex conjugate function with respect to Ψ (where 
corresponds to the width of the wavelet and  defines the position of the wavelet on the x-axis). The
Iterative Extremum Search Method in Wavelet Transform Space. Let's define the one-dimensional
Operator</p>
      <p>( ) = (ℎ1, … , ℎ )
where  ( ) is the objective function;  is the parameter vector;   −   represents the parameters.</p>
      <p>In known optimization methods using wavelet transform (WT), the property of gradient
estimation and wavelet processing is employed, where the sign of the gradient changes as it
approaches an extremum. In other words, the condition for an optimum is considered to be the
equality of all partial wavelet transforms: WT(c) = 0 .
regular iterative optimization algorithm in the wavelet transform space [19] is as follows:
step size; 
iteration number; 
extremum coordinate.
the optimization requires significant computational resources.</p>
      <p>It should be noted that due to the direction of the search being assessed using wavelet processing,
(2)
(3)
(4)
(5)</p>
      <p>This limits the applicability of the method for practical UV tasks, as it is necessary to ensure
optimization efficiency, especially for multimodal objective functions with low signal-to-noise ratios
existing wavelets is due to the fact that it is characterized by low computational complexity.</p>
      <p>
        When implementing computational procedures using it, multiplication operations by +1 and -1
are performed, which can simplify the hardware implementation of this procedure and increase the
speed of operation [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>However, with a large Haar wavelet function length and a high signal-to-noise ratio of the
objective function, the search may deviate in the wrong direction, moving away from the global
minimum towards a local minimum. This results in reduced search efficiency.</p>
      <p>The proposed work aims to improve search efficiency by selecting the Haar wavelet function
length in the search area.</p>
      <p>To achieve this, it is proposed that after localizing the minimum area of the functional during the
Haar wavelet function search phase Ψ1( ) and narrowing down the search area by defining
constraints   ≤ 0, the Haar wavelet function length should be selected judiciously, taking into
account the characteristics of the extrema in disparity evaluation.</p>
      <p>The method with iterative constraint evaluation using Haar wavelet functions is implemented in
the following sequence:</p>
      <p>
        Step 1. Execute steps 1-5 of the basic wavelet optimization method, taking into account the
iterative scheme [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The coordinates of the extremum are determined as follows:
direction of movement towards the extremum, calculated as follows:
(6)
(7)
(8)
 ( ( [ ],  [ − 1])) = { 1 ,  2 , … ,   }
where  
the result of processing by the  -th variable:

1

 
2
 =− 2
 ≠0
 
=
      </p>
      <p>
        ∑  ( [ ],   +  ) ⋅   ( )
where s is the length of the Haar wavelet support; Ψ ( ) is the step size of the Haar wavelet
discretization; Ψ1 is the Haar wavelet at the k-th start;  = 1, … , 
is the dimensionality of the
parameter vector. This approach allows for determining the range of variation in the extremum [18]
coordinates with  1 being the search error for the optimal start (determined during the preliminary
functional quality studies);  2 being the search error for the practical task's optimum. The search is
conducted considering the following parameters:  [0] as the initial approximation to the optimum
coordinate;  [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as the step size;  as the Haar wavelet discretization step size; s1 as the length of
the Haar wavelet support for the first start Ψ ( ) (determined during the preliminary functional
quality studies); Δ as the step size for changing the length of the Haar wavelet support when
determining the range of extremum coordinates; k = 1 as the start number; n = 1 as the iteration
number; A1 as the value of the minimal height of the registered extremum. The search continues
until the sign of processing    ( ( [ ],  [ − 1])) changes when estimating the direction of
movement towards the extremum.
      </p>
      <p>Step 2. The constraints of the search zone for the next algorithm start g_1(c[n]) and g_2(c[n])
are determined:</p>
      <p>where  ∗[ − 1]is the extremum coordinate at step [
extremum coordinate at step  , where the sign changed (see Figure 1).</p>
      <p>1( [ ]) ≥  ∗[ − 1];  2( [ ]) ≥  ∗[ ]
(9)
  is the</p>
      <p>Step 3. The length of the Haar wavelet support   for the subsequent start is determined by
selecting from the range   = {2,4,6} according to the condition   ≤
| 2( [ ])− 1( [ ])|

.</p>
      <p>Support Length Ψ ( ) 1 = 10
Defined Using Haar Wavelet Ψ ( ) (see Figure 3). When the SNR (signal-to-noise ratio) is reduced
to 2, the proposed method's performance (measured by timer) is on average 1.1 times higher (see
stage of selecting the length of the Haar WF carrier (Fig. 2, b). The sensitivity of the developed
optimization method to local extrema and the starting point of the search using the test Schwefel
function
  ( ) = 418,9829 + (− ⋅ 
√| |)
(10)</p>
      <p>The sensitivity of the developed optimization method to local extrema and the starting point of
the search using the test Schwefel function has been studied. This function has a false global
minimum. The function at x ∈ (−500; 500) has a global minimum fs(x) = 0) at (x = 420.9829.
During the research, the starting point was chosen randomly. The gradient descent method made it
possible to find the minimum closest to the start, while the optimization method with the Haar
wavelet function achieved the global minimum with an error  ≤ 10−2 in 128 out of 150 cases.</p>
      <p>This results in a probability of finding the extremum coordinate of 0.85. The global minimum was
not found when the starting point values were chosen outside the interval  ∈ (−420;
470).
 [ ] are marked with squares.</p>
      <p>
        1
basic method [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]; 2
proposed method
1
basic method [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]; 2
proposed method
(a); relative error in determining the extremum coordinates (b).
      </p>
      <p>Future research in this direction will focus on evaluating the impact of wavelet function support
length on the performance of the maximum search procedure, as well as the effects of the wavelet
discretization step and the step size in iterative searches with Haar wavelets ( ).</p>
    </sec>
    <sec id="sec-6">
      <title>5. Testing</title>
      <p>The developed optimization method on the wavelet transformation base was tested in the context of
constructing a depth map from a test image in the image database [20, 21].</p>
      <p>The length of the wavelet support used in the minimum search after investigation was chosen as
17, the wavelet discretization step was 1, and the step  in the iterative Haar wavelet search was
applied 0,9. The minimum image alignment error was found after 12 iterations (starting the search
at [1; 1]).</p>
      <p>Furthermore, simulation results indicated that modifying the known StereoBM method by
incorporating wavelet transformation for disparity calculation enhances noise robustness, reduces
error in locating the characteristic fragment on the intensity line, and also decreases energy
consumption by nearly a factor of 2 (Table 1).
6. Conclusions
The paper addresses the development of approximate depth estimation methods for three-dimension
computer vision in autonomous vehicles based on an improved optimization method utilizing
wavelet transform. This method can be used for evaluating informative features for fragment
matching and/or optimizing costs in image analysis, as well as for refining inconsistencies, among
other applications. The study demonstrates possible solutions for obtaining an approximate depth
map by simplifying the calculation of disparity values, traditionally used for depth map formation,
within the intensity space and using image contour descriptions with adjustable detail based on
wavelet transform. The advantage of the developed optimization method over existing algorithms
on the wavelet transformation base is its increased efficiency due to the rational choice of Haar
wavelet support length in the extremum search area.</p>
      <p>Future research in this direction will focus on evaluating the impact of wavelet function support
length on the performance of the maximum search procedure, as well as the effects of the wavelet
discretization step and the step size in iterative searches with Haar wavelets.</p>
      <p>This improvement makes the proposed method suitable for autonomous vehicles operating under
constraints of computational and energy resources.
[18] Y. Bodyanskiy, N. Lamonova, I. Pliss, and O. Vynokurova, "An Adaptive Learning Algorithm
for a Wavelet Neural Network," Expert Systems 22.5 (2005): 235-240.
doi:10.1111/j.14680394.2005.00314.x.
[19] G. Shcherbakova, H.-S. Shi, V. Krylov, N. Bilous, and S. Antoshchuk, "Estimation of the
Duration of RR-Intervals of Electrocardiograms by Means of Multi-Start Optimization Based on
Wavelet Transformation," in: IEEE 9th International Workshop on Intelligent Data Acquisition
and Advanced Computing Systems: Technology and Applications, 21-23 September 2017,
Bucharest, Romania. doi:10.1109/IDAACS58523.2023.10348849.
[20] Middlebury Stereo Vision Dataset, "Full-Size Stereo Data and Scene Information," Middlebury
Stereo Vision Project. URL:
https://vision.middlebury.edu/stereo/data/scenes2005/FullSize/Art/Illum1/Exp1/.
[21] Middlebury Stereo Vision Dataset, "Stereo Data Archive," Middlebury Stereo Vision Project.</p>
      <p>URL: https://vision.middlebury.edu/stereo/data/scenes2005/.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Karnati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mehta</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ks</surname>
          </string-name>
          ,
          <source>Manu. "Artificial Intelligence in Self Driving Cars: Applications</source>
          ,
          <article-title>Implications and Challenges,"</article-title>
          <source>Universal Journal of Business and Management</source>
          <volume>21</volume>
          (
          <year>2022</year>
          ):
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .12725/ujbm.61.1.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Konovalenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kuznetsova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Popov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shepelev</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Stepanyan</surname>
          </string-name>
          ,
          <article-title>"New Approaches to the Integration of Navigation Systems for Autonomous Unmanned Vehicles (UAV)," Sensors 18 (</article-title>
          <year>2018</year>
          ):
          <fpage>3010</fpage>
          . doi:
          <volume>10</volume>
          .3390/s18093010.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suzuki</surname>
          </string-name>
          ,
          <article-title>"Integrated Navigation for Autonomous Drone in GPS</article-title>
          and
          <string-name>
            <surname>GPS-Denied</surname>
            <given-names>Environments</given-names>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Journal of Robotics and Mechatronics 30.3</source>
          (
          <year>2018</year>
          ):
          <fpage>373</fpage>
          <lpage>379</lpage>
          . doi:
          <volume>10</volume>
          .20965/jrm.
          <year>2018</year>
          .p0373.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Savkin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Vucetic</surname>
          </string-name>
          ,
          <article-title>"Collision Free Navigation of a Flying Robot for Underground Mine Search and Mapping,"</article-title>
          <source>in 2018 IEEE International Conference on Robotics and Biomimetics (ROBIO)</source>
          , pp.
          <fpage>1102</fpage>
          <lpage>1106</lpage>
          . IEEE,
          <year>2018</year>
          . doi:
          <volume>10</volume>
          .1109/ROBIO.
          <year>2018</year>
          .
          <volume>8665108</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Tripicchio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Satler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Unetti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Avizzano</surname>
          </string-name>
          ,
          <article-title>"Confined Spaces Industrial Inspection with Micro Aerial Vehicles and Laser Range Finder Localization,"</article-title>
          <source>International Journal of Micro Air Vehicles 10.2</source>
          (
          <year>2018</year>
          ):
          <fpage>207</fpage>
          <lpage>224</lpage>
          . doi:
          <volume>10</volume>
          .1177/1756829318757471.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lai</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>"Robust Autonomous Flight and Mission Management for MAVs in GPS-Denied Environments,"</article-title>
          <source>in: 11th Asian Control Conference (ASCC)</source>
          ,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .1109/ascc.
          <year>2017</year>
          .
          <volume>8287144</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Mansouri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kanellakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kominiak</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Nikolakopoulos</surname>
          </string-name>
          ,
          <article-title>"Deploying MAVs for Autonomous Navigation in Dark Underground Mine Environments,"</article-title>
          <source>Robotics and Autonomous Systems</source>
          <volume>126</volume>
          (
          <year>2020</year>
          ):
          <fpage>103472</fpage>
          . doi:
          <volume>10</volume>
          .1016/j.robot.
          <year>2020</year>
          .
          <volume>103472</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Vong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ravitharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Reichl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chevin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <article-title>"Small Scale Unmanned Aerial System (UAS) for Railway Culvert and Tunnel Inspection,"</article-title>
          <source>in: ICRT</source>
          <year>2017</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>1024</fpage>
          <lpage>1032</lpage>
          .. doi:
          <volume>10</volume>
          .1061/9780784481257.102.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.-W.</given-names>
            <surname>Choi</surname>
          </string-name>
          , et al.,
          <article-title>"An Overview of Drone Applications in the Construction Industry,"</article-title>
          <source>Drones 7</source>
          .8 (
          <year>2023</year>
          ):
          <fpage>515</fpage>
          . doi:
          <volume>10</volume>
          .3390/drones7080515.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Hamzah</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Ibrahim</surname>
          </string-name>
          ,
          <article-title>"Literature Survey on Stereo Vision Disparity Map Algorithms,"</article-title>
          J.
          <string-name>
            <surname>Sensors</surname>
          </string-name>
          (
          <year>2016</year>
          ):
          <volume>8742920</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>8742920</lpage>
          :
          <fpage>23</fpage>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2016</year>
          /8742920.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Okae</surname>
          </string-name>
          ,
          <article-title>"Optimization of Stereo Vision Depth Estimation Using Edge-Based Disparity Map,"</article-title>
          <source>in: 10th International Conference on Electrical and Electronics Engineering (ELECO)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1171</fpage>
          -
          <lpage>1175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Richter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosenberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Illmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Buchanan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Notni</surname>
          </string-name>
          ,
          <article-title>"Suitability Study for RealTime Depth Map Generation Using Stereo Matchers in OpenCV and Python," in: Engineering for a Changing World:</article-title>
          <source>Proceedings: 60th ISC</source>
          , Ilmenau Scientific Colloquium,
          <source>Technische Universität Ilmenau, September 04-08</source>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .22032/dbt.58859.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aslam</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Ansari</surname>
          </string-name>
          ,
          <article-title>"Depth-Map Generation Using Pixel Matching in Stereoscopic Pair of Images," (</article-title>
          <year>2019</year>
          ) arXiv:
          <year>1902</year>
          .
          <article-title>03471v3 [cs</article-title>
          .CV]
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Fahmy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Ismail</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Al-Janabi</surname>
          </string-name>
          ,
          <article-title>"Stereovision Based Depth Estimation Algorithm in Uncalibrated Rectification,"</article-title>
          <source>International Journal of Video &amp; Image Processing &amp; Network Security 13.2</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Antoshchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Kondratyev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Y.</given-names>
            <surname>Shcherbakova</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hodovychenko</surname>
          </string-name>
          ,
          <article-title>"Depth Map Generation for Mobile Navigation Systems Based on Objects Localization in Images,"</article-title>
          <source>Herald of Advanced Information Technology 5.1</source>
          (
          <year>2022</year>
          ):
          <fpage>11</fpage>
          -
          <lpage>18</lpage>
          . doi:
          <volume>10</volume>
          .15276/hait.05.
          <year>2022</year>
          .
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Shcherbakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Krylov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Abakumov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Brovkov</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Kozina</surname>
          </string-name>
          ,
          <article-title>"Sub Gradient Iterative Method for Neural Networks Training,"</article-title>
          <source>in: Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications: 6th IEEE Int. Workshop IDAACS</source>
          '
          <year>2011</year>
          , Prague, Czech Republic,
          <volume>15</volume>
          <fpage>17</fpage>
          <lpage>Sept</lpage>
          .
          <year>2011</year>
          : Proceedings, pp.
          <fpage>361</fpage>
          -
          <lpage>364</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W.</given-names>
            <surname>Huan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Shcherbakova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sachenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Volkova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Rusyn</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Molga</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Haar Wavelet-Based Classification Method for Visual Information Processing Systems," Applied Sciences 13.9</source>
          (
          <year>2023</year>
          ):
          <fpage>5515</fpage>
          . doi:
          <volume>10</volume>
          .3390/app13095515.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>