<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Zhuk DV, Tuzikov AV. Reconstruction of a three-dimensional model using two digital images. Informatics</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>method for calculating the disparity value on stereo images in problems of stereo-range metering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A.N. Volkovich</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>United Institute of Informatics Problems of the National Academy of Sciences of Belarus</institution>
          ,
          <addr-line>Surganova str. 6, 220012, Minsk</addr-line>
          ,
          <country>Republic of Belarus</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2006</year>
      </pub-date>
      <volume>1</volume>
      <fpage>16</fpage>
      <lpage>26</lpage>
      <abstract>
        <p>The paper considers the solution of the problem of restoring three-dimensional information based on stereo images. An original combined approach to disparity calculation is proposed, as well as a variant of solving the problem of the heterogeneity of the initial data in calculating the actual metric parameters. The tasks of comparing and image search are the main tasks in computer vision. The quality of the search, low sensitivity to distortions - the fundamental requirements for search algorithms. There are many approaches to solve such tasks, as well as a wide range of technical solutions to the problems of rangу-finding using lasers and other systems. At the same time there are a number of problems that imply the impossibility of active far-range systems use. In the described project is planned to develop and implement effective methods of image sections searching based on characteristics of pixel's local neighborhoods.</p>
      </abstract>
      <kwd-group>
        <kwd>stereo images</kwd>
        <kwd>disparity maps</kwd>
        <kwd>gradient operators</kwd>
        <kwd>satellite photographs</kwd>
        <kwd>long-range systems</kwd>
        <kwd>optical systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Intoduction</title>
      <p>2 sin</p>
      <p>For a long time, the use of stereovision principles in systems has not been considered in practice due to the relative
complexity of implementation and poor quality of digital images. In recent years, the quality parameters of shooting equipment
increased significantly, which suggests the possibility of developing such systems. In addition, the stereovision system can
possess in addition to low-visibility (due to the implementation of the passive calculation technique) also by a functional
allowing the measuring on post-production. It is possible to use stereo-pairs and recalculate the parallax for a multitude of image
points in the overlapping area and the sharply displayed image space in comparation with active range finder systems witch
allows to obtain the distance-range information only at the time of shooting.</p>
      <p>Based on the above facts, it seems extremely relevant to develop both theoretical methods in the field of stereo-range
metering, as well as technical hardware-software solutions.</p>
      <p>Image Processing, Geoinformation Technology and Information Security / A.N. Volkovich</p>
    </sec>
    <sec id="sec-2">
      <title>3. Passive systems of range-finding</title>
      <p>A digital image obtained with a passive stereo system carries only color or brightness information and does not possess any
additional data.</p>
      <p>Images of the real world include a narrow set of colors or luminances, and therefore, when solving the problem of
determining conjugate identification is not for individual points, but for fragments of images. Thus, it is extremely important for
a point taken in one image to know, where at the second image is its conjugate and how to compare these fragments correctly.</p>
      <p>A comparison of the neighborhoods of conjugate points does not yield to strict formalization. It is based on the problem of
identifying images of fragments of the three-dimensional world from images. This task can hardly be adequately described in a
formal way. Significant differences in views lead to the appearance of projective and brightness distortions when shooting. It is
of fundamental importance that these differences depend not only on the geometry of shooting-system, but also on the geometric
and physical characteristics of the surface itself. The location of the light source influences to the surface affects the light
distribution. The position of the surface elements and their properties determine the amount of energy that enters the camera
lenses and the local differences in the brightness of the conjugate fragments of images.</p>
      <p>The significance of the differences depends on the difference in the viewing angles. The more this difference (in particular,
the larger base), the less similar the images becomes. Therefore, all methods of comparing neighborhoods of conjugate points
rely more or less on a formal approach rather than on the character of images. Also important is the possibility of their
preliminary processing, reduction to an epipolar stereopair, construction of efficient descriptors of neighborhoods of conjugate
points for accuracy and speed for comparison.</p>
      <p>In an idealized situation, the values of the similarity function in the scanning process along the line should represent a
onemoment peak value for the desired pixel when the zero similarity value is returned for all other pixels (neighborhoods) of th e
line.</p>
      <p>During processing real graphics data, such combination of values returned by the function of similarity is impossible. But
processing initial data with sufficient information to identify a local region, the graph of the function retains a sufficiently clear
extremum, which allows to identify the desired pixel of the image.</p>
      <p>The main task in the construction of the disparity map is the selection of a variant for comparison of regions in which the
extremum of values of the similarity function will be most pronounced. This involves defining for the point some characteristics
that would uniquely characterize the point of the image. Moreover, the conjugate point on the reference image had identical or
maximally similar values of similar characteristics.</p>
      <p>The final step of calculating disparity is the aggregation of total or averaged values. When aggregating, as the resulting
disparity for the point  of the base image, the value of  , is chosen, where the minimum value of the cost is reached. In the real
situation, it is possible to find several values of  with the same or minimum values of disparity (especially when averaging the
values). The problem of possible ambiguity arises in most methods of constructing disparity maps, which is associated with
optical, mechanical, electronic features of cameras. A solution to multi-valuedness can be the introduction of certain conditions
on the value of disparity, for example, the largest, average or smallest possible.</p>
    </sec>
    <sec id="sec-3">
      <title>4. Combined method of stereo reconstruction</title>
      <p>Usually, in practice, measures are taken based on the sum of the absolute differences or the sum of the squares of the
differences. Both functions (summation over a given window) allow you to calculate the cost effectively enough when the
corresponding pixel of the conjugate image has the closest intensity value. These functions are extremely sensitive to the quality
and parameters of the original data (exposure, glare, overexposure, underexposure, matrix noise, random emissions).</p>
      <p>In the process of carrying out a computational experiment on real-world images, it was determined that correlation methods
of comparing local parts of images give stable results on textured areas and extremely low accuracy in homogeneous areas (there
are no explicit contrasts). During analyze of "alive" systems, we can conclude that in nature the distance to a homogeneous
object is also poorly localized and its binding to the boundaries due to reflex saccadic movements.</p>
      <p>It can be argued that a comprehensive approach to the problem of ranging is required, including both consideration of the
possibilities of increasing the uniqueness of the means of identification, and the development of a combined approach using
different techniques to the neighborhoods of points, depending on their local characteristics.</p>
      <p>Work with digital stereo images allowed us to determine the following combined method: correlation areas with the
maximum uniqueness of points should be applied to areas of images that have contrast objects in their area (brightness
differences), and to homogeneous ones - to bind to remote contrast boundaries by building a set of vectors on several directions.</p>
      <p>Thus, in the preprocessing phase, it becomes important to compile a calculation map based on the proximity to the
contrasting areas. In order to classify a point as being on a brightness difference, the brightness change associated with a given
point must be substantially greater than the change in brightness at the background point. In connection with the specifics of
local calculations, the way to determine the "essential" values to establish a threshold. In turn, the concepts of the first and
second derivatives are used for the quantitative expression of the brightness variation.</p>
      <p>The definition of an image point as a drop point occurs if its two-dimensional derivative of the first order exceeds a certain
predetermined threshold. The calculation of the first derivative of a digital image is based on various discrete approximations of
a two-dimensional gradient. The direction of the gradient vector coincides with the direction of the maximum rate of change of
the function f at the point (х, у).</p>
      <p>The calculation of the gradient of the image consists in obtaining the values of the partial derivatives 
= 
/
 у =
 /
for each point. One of the methods for finding the first partial derivatives 
 у at a particular point is to apply the
following gradient Sobel operator:
  = ( 7 + 2 ∗  8 +  9) − ( 1 + 2 ∗  2 +  3)
 у = ( 3 + 2 ∗  6 +  9) − ( 1 + 2 ∗  4 +  7)</p>
      <p>It is necessary to determine the appropriate masks for the Sobel operator, which identifies horizontal and vertical contours
(brightness differences) for convolution with the original image. It is also possible to change the above formulas that give the
maximum response for contours directed diagonally. Additional pairs of Sobel's masks for detecting gaps in diagonal directions
can be defined as:</p>
      <p>Working with three components of color can be represented as a "cloud" of points in three-dimensional space with the axes
corresponding to the color channels of the image. However, the RGB space is not orthogonal, due to the specificity of the human
visual analyzer, which has a different number of rods and cones that are susceptible to a particular color.</p>
      <p>Since the correlation function of a three-dimensional space is a measure of correlation functions that use the Euclidean
distance, which is correctly calculated in orthogonal systems, one should perform orthogonalization of the space RGB into the
The representation of the RGB base colors, according to the ITU recommendations, in the XYZ space has the following
Therefore the transformation system for translating colors between RGB and XYZ systems can be represented in the
respective channels.
information is seen as obvious.
space XYZ.
correction factors:
following form:
 1  2  3
 4  5  6
 7
 8</p>
      <p>9
0
−1
−2
−2
−1
0
1
0
−1
−1
0
1
2
1
0
0
1
2



:  = 0,64  = 0,33
: 
: 
= 0,29</p>
      <p>= 0,60
= 0,15 
= 0,06
for points lying on the diagonal edge -45 degrees;
for points lying on the diagonal edge +45 degrees.</p>
      <p>For each of the masks the sum of the coefficients equals zero. That means that these operators will validate zero response on
the areas of constant brightness, which is characteristic of the differential operator. The masks considered are used to obtain the
gradient components Gx and Gy. To calculate the magnitude of the gradient these components must be used together:</p>
      <p>= |  | + |  |
Сalculation map is to be constructed on the computed gradient map base. This process implies the recognition of the area as
low-textured in the event that in the search window around the point less than 15% of the area is occupied by contrasts.</p>
      <p>Direct calculation of disparity occurs in several stages. At the first stage, high-textured sections are processed using the
correlation functions of similarity measures, such as, for example, Euclidean distance, cross-correlation, etc.</p>
      <p>In the world practice only brightness information is usually used as criteria for comparing image points. The disadvantage of
this approach is the color interpretation multiplicity for points with the same brightness value. In addition, one should take into
account the fact that the perception of color and monochrome images is uneven. This feature is taken into account in the methods
of degradation of the color model of the image to 256 shades of gray due to the introduction of coefficients applied to the
 = 0.299 ∗  + 0.587 ∗  + 0.114 ∗ 
Most of the images were initially formed by a color sensor in color. Therefore, in order to increase efficiency, the use of color



directions are calculated. A group of multidimensional characteristic vectors is formed, to which a similar approach is applied, as
well as to vectors with color information.</p>
      <p>As the stages are completed, the disparity map is filled. Due to the fact that all operations are in strict accordance with the
calculation card, auto-aggregation of the results of different stages into a single map occurs.</p>
    </sec>
    <sec id="sec-4">
      <title>5. Calculation of the distance to the object on the basis of heterogeneous initial data</title>
      <p>Generalized the principle of determining the position of points in space on the basis of disparity data has been repeatedly
described in the literature. Suppose two cameras L and R are installed in such a way that their  -axes are collinear, and the  and
 , axes are parallel. The centers of the cameras are displaced relative to each other by an amount  , corresponding to the base of
the stereoscopic system. When observing a certain point of the space  the point   , is formed on the left image, and on the right
  .</p>
      <sec id="sec-4-1">
        <title>Considering the similarity of two pairs of triangles, we obtain the equations:</title>
      </sec>
      <sec id="sec-4-2">
        <title>The solution of the system of equations allows one to uniquely calculate the position of a point in space.</title>
        <p>Unfortunately, the given system of equations is not applicable for digital reconstruction, because there is a mixture of
different systems of dimension: disparity value in pixel distance, focal length and base in metric units. However, it is possible to
calculate the distance to the point using the disparity value, the angle of the lens alignment and the base of the stereo system.
a)
b)
c)
d)</p>
        <p>The stereovision system can be represented in the following form:





А, А1 – observation point;</p>
      </sec>
      <sec id="sec-4-3">
        <title>BC – left image;</title>
      </sec>
      <sec id="sec-4-4">
        <title>B1C1 – right image;</title>
      </sec>
      <sec id="sec-4-5">
        <title>BC1 – zone of overlap;</title>
        <p>– horizontal lens opening angle.</p>
      </sec>
      <sec id="sec-4-6">
        <title>The calculation of the distances to the object is the problem of solving the triangle B1АА1.</title>
        <p>Within the system, the base of the triangle is known - the base of the stereo system. Angles at the base can be calculated
through the angle of the lens opening.</p>
        <p>The calculation of the viewing angle to the object of interest for the left camera is performed as follows (Fig. 4c):
∠  1 =  = 90 ± 
(
β – angle of the desired triangle;
α – angle of the horizontal lens opening;
  – the  - coordinate of the image point;
  – the width of the image (Fig 4d).
Знак «+» используется в системе при 
&lt;

2
, «-» при 
&gt;

2
соответственно, а при</p>
        <p>принимаем
=

2
∠  1 =</p>
        <p>, «+» for 
of the system</p>
        <p>The angle of sight is calculated in the same way as the correction for the fact that the sign «-» is used in the system for
&gt;

2
and for 
=

2</p>
        <p>and ∠  1 =  1 = 90.</p>
      </sec>
      <sec id="sec-4-7">
        <title>The third angle can be obtained by the formula:</title>
        <p>By the sine theorem, it is possible to determine the lengths of the sides  1 and 
  1
2
)
)
−
  1 =</p>
        <p>As a result, the distance from the reference (left) camera to the object and the angle relative to the base (plane of the
matrices) of the system are obtained.</p>
        <p>It should be noted that increasing the measured distance increases the sensitivity of the system to the accuracy of the
alignment of the system and the quality of the images. Since the angles at the base of the system take values close to 90° . This
leads to the fact that the values of trigonometric functions change very dynamically.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>During the research, the author made a study of the existing methods for processing stereo images in the tasks of stereo
reconstruction. The algorithms functioning patterns are revealed and the causes of their unstable work are determined. A
combined image processing technique that takes into account the characteristics of local image sections is proposed.
Additionally author examined the problem of the heterogeneity of the initial data necessary for obtaining metric information of
three-dimensional objects in the field of interests.</p>
      <p>The algorithm developed and described was implemented by the author in the form of a program library, which can later be
used in a wide range of applications. Due to the fact that the matrix of distances to image points can be translated into a specific
coordinate system for one or another application system. It should also be noted that the organization of the user's access to the
functions allows for more flexible use of the library.</p>
      <p>In addition, the described technique has found its application in a number of software and hardware and software
developments that are carried out at the United Institute of Informatics Problems of the National Academy of Sciences of</p>
      <sec id="sec-5-1">
        <title>Belarus. Specifically, as element of the mobile topogeodetic system and as program library for ERS system.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>