<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Disparity map estimation with deep learning in stereo vision</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mendoza Guzman V ctor Manuel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mej a Mun~oz Jose Manuel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moreno Marquez Nayeli Edith</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rodr guez Azar Paula Ivone</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Santiago Ram rez Everardo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Autnoma de Ciudad Jurez @alumnos.uacj.mx</institution>
        </aff>
      </contrib-group>
      <fpage>27</fpage>
      <lpage>40</lpage>
      <abstract>
        <p>In this paper, we present a method for disparity map estimation from a recti ed stereo image pair. We proposed, a new neural network architecture based on convolutional layers to predict the depth from the stereo vision images. The Middlebury datasets were used to train the network with a known disparity map in order to compare the error of the estimated map.</p>
      </abstract>
      <kwd-group>
        <kwd>Neural networks stereo vision disparity map</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        One of the techniques that have shown potential for obtaining three-dimensional
(3D) information from two-dimensional (2D) images is the processing of stereo
images [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Images are commonly considered as 2D representations of the real
world (3D). Stereo vision by computer involves the acquisition of images with
two or more cameras moved horizontally to each other. In this way, di erent
views of a scene are recorded and can be processed for di erent applications
such as vehicle tracking [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], aircraft estimation and positioning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and
automatic adaptation systems that cover a wide range of applications [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], including
3D reconstruction or disparity map estimation.
      </p>
      <p>
        Stereo vision tries to imitate the mechanisms that are made in the human
visual system and the human brain. A scene depicted with two horizontally
displaced cameras will get two slightly di erent projections of a scene. If these two
images are compared, additional information can be reached, such as the depth
of a scene. This process of extracting the three-dimensional structure of a scene
from pairs of stereo images is called computational stereo, and the resul is
generally a disparity map which is a map of the depth or distance at which the
objects of a scene are located [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        In recent years, numerous algorithms and applications for the estimation of
disparity maps have been presented. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] a method based on satellite images
is proposed to monitor trees and vegetation. The stereo matching algorithms
are calculated to measure the disparity map based on stereo satellite images.
The estimation of the height of trees and vegetation near the base poles to the
depth map is inversely proportional to the disparity map. In [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] another
application for pedestrian detection is proposed based on the dense disparity map for
smart vehicles. The dense disparity map is used to improve pedestrian detection
performance. The method consists of several steps, detection of obstacle areas
using information of characteristics of roads and detection of columns, detection
of pedestrian areas using a segmentation based on dense disparity maps and
detection of pedestrians using the optimum characteristic.
      </p>
      <p>
        The work of [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] is a investigation for object tracking were a stereoscopic
camera is used to detect objects, which makes it a low-cost solution for tracking
objects. Its objective is to detect objects in a video sequence and track them
throughout the video without prior knowledge about the objects. Calculate the
disparity map using a pair of stereophonic images. Then, the disparity map is
subjected to a depth-based segmentation to detect object blobs and the
corresponding region in recti ed stereo-image is the object of interest.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] an e cient algorithm is presented to optimize the performance of a
stereoscopic vision system and accurately relate the calculated disparity map
with the real depth information.
      </p>
      <p>
        With the new technologies and arti cial intelligence in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] a methodology
for the detection of robust obstacles in outdoor scenes for autonomous driving
applications is proposed using a multi-value stereo disparity approach. The
disparity computation su ers a lot from re ections, lack of texture and repetitive
patterns of objects. This can lead to incorrect estimates, which may introduce
some bias in the obstacle detection approaches that make use of the disparity
map. To overcome this problem, instead of a disparity estimate of a single value,
a new research that uses a diversity of candidates for each point of the image is
proposed. These are selected based on a statistical study characterized by the
performance of di erent parameters: number of candidates and the distance
between them compared to the real value of the disparity. It continues creating a
location map from which the estimation of the obstacles is obtained.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] explain di erent phases to perform the depth estimation: First they
perform an extraction of characteristics, an initial estimation and a nal re
nement of the estimated depth.
      </p>
      <p>Obtain the main characteristics of the two input images, stereo images left
and right, and the nal re nement is done in a main block of the network with
two re nement sub-networks. The network contains a pair of convolution layers
and a pair of deconvolution layers to perform the sampling at the output.</p>
      <p>
        This subnetwork structure generates a depth map through an architecture
that encodes and decodes information, inspired by the DispNetCorr1D network
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        The innovative aspect of the research in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] is that for the rst stage the
network described includes modules to perform deconvolution in addition to the
traditional which leads to estimate the disparity with the same size as the images
that are being used for the input. In the second stage, which is re nement, they
propose residual learning used in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] propose a new optimization method that uses strong smoothing
restrictions obtained in a neural network. The goal for this is to soften the output
disparity map in a robust manner. The rst step in this research was to de ne
the CNN architecture, called DD-CNN, to classify if the disparities are
discontinuous. The training of this architecture was carried out with real data from
Middlebury stereo data [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. In the next step they de ne an energy function
composed of a term of data obtained with the method of [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and a term that
penalizes disparity di erences.
      </p>
      <p>
        Finally in [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] a network is developed with a dual structure. Each of the
structures takes an input image that passes through a nite set of layers followed by
a normalization and a recti ed linear unit. In their experiments, di erent lters
were tested per layer and the parameters were shared between the two
structures. For the training, small kernels of the images extracted from a set of pixels
were randomly used. Providing a diverse set of examples and it was considered
an e cient method in memory.
      </p>
      <p>
        For this research it is proposed a new architecture based on convolutional
networks, for the training of the network we will use the stereo images from the
Middlebury database [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the disparity map results obtained in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Theory</title>
      <sec id="sec-2-1">
        <title>Convolutional Network</title>
        <p>Convolutional networks are composed of a speci c type of connections for data
processing that have known properties. Examples can be highlighted from time
series, which is considered a grid of one dimension with a regular time
dened, and images, which is the same case as time series with more than one
dimension ordered at speci c points. CNN (Convolutional Neuronal Network)
has demonstrate a high level of success in eld applications. It is called the
convolutional network because it performs the mathematical operation called
convolution within neurons. Convolution is a specialized type of linear
operation. Convolutional networks are simply neural networks that use convolution
instead of the general multiplication of matrices in at least one of their layers.</p>
        <p>In its most general form, convolution is an operation in two functions of an
argument of real value.</p>
        <p>In machine learning applications, the input is usually a multidimensional
array of data, and the core is usually a multidimensional array of parameters that
are adapted by the learning algorithm.</p>
        <p>The convolution takes advantage of three important ideas that can help to
improve a learning system: dispersed interactions, shared use of parameters and
equivalent representations. In addition, convolution provides a means to work
with entries of variable size. Traditional neural network layers use matrix
multiplication through a parameter matrix with a separate parameter that describes
the interaction between each input unit and each output unit. However,
convolutional networks often have scattered interactions. This is achieved by making
the kernel smaller than the input.</p>
        <p>CNN are usually developed in the following stages: First, a de ned number of
convolutions are made at the same time to produce a group of linear activations.
In the next stage, each activation is executed with an activation function that
is not linear. This stage could be de ned as the detector stage. The third stage
uses a grouping function with the objective of modifying the output of each layer.</p>
        <p>A grouping function replaces the output of the network at a given location
with a summary statistic of the nearby outputs.</p>
        <p>
          In all cases, the grouping helps to make the representation almost invariant
for small translations of the entry. The invariance to translation means that, if
we translate the entry by a small amount, the values of most of the grouped
outputs do not change [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Proposed architecture</title>
      <p>The proposed convolutional network can be seen in Figure 2, the architecture
of the network is as follows: The inputs, pair of stereo images, are individually
processed through 3 convolutional layers, the rst 2-D convolutional layer is 32
lters with a 3x3 kernel returning the same size of the images. The output of
the rst layer continue with a MaxPooling with a size of 2, followed with
another convolutional layer of 62 lters with a 3x3 kernel and the output with a
MaxPoling also with a size of 2 and nally a convolutional layer with a size of 92
lters and a 3x3 kernel is applied to obtain an image in its original size of 96x96.
This process is applied independently to each image, the outputs are combined
to apply another 3 convolutional layers.</p>
      <p>
        The rst convolutional layer has 62 lters with a 3x3 kernel followed by a
UpSampling with a size of 2x2. At the exit, another convolutional layer with a
size of 22 lters and a 3x3 kernel is applied to the output and another
UpSampling with a size of 2x2, nally the last convolutional layer of 1 lter and a 2x2
kernel is applied. All the activation functions used in the convolutional layers is
the recti ed linear unit (ReLU).
In this research we used the images of the database of Middlebury [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
images of the database are in color, and consist of a pair of stereo images as show in
Figure 3. The original database images were pre-processed, a change was made
in the size of the images in such a way that all the images had a size of 92x92,
the main idea of leaving the square images is to speed up the computations and
to lower memory requirements.
      </p>
      <p>Aloe left image</p>
      <p>Aloe right image
(1)
(2)</p>
      <p>Since the neural networks need a large amounts of data to work e ectively,
data augmentation was used to increase the number of images in the data set.
For data augmentation, operations of translation, rotation, and scaling were used
to increase the database to 500 images.
3.2</p>
      <sec id="sec-3-1">
        <title>Metrics</title>
        <p>The metrics that will be used for the evaluation of this research will be the
Peak Signal to Noise Ratio (PSNR) and the Structural Similarity Index (SSIM).
These metrics have been used as a reference point for the comparison of input
images and output images in the evaluation of image quality. The PSNR uses
the Mean Square Error (MSE), the MSE is calculated between the average of
the original intensity and the intensity of the output image and is given by:
M SE =</p>
        <p>1
N M</p>
        <p>M 1 N 1
X X e(m; n)2
m=0 n=0</p>
        <p>Where e(m; n) is the di erence of the error between the original image and
the output image PSNR is the mathematical measure of image quality based on
the pixel di erence of two images. And it is de ned by:</p>
        <p>P SN R = 10log</p>
        <p>
          s2
M SE
Where s = 255 for an 8-bit image [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          For the SSIM, Wang[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], proposed the Structural Similarity Index as an
improvement of the Universal Image Quality Index (UIQI). The SSIM is calculated
as follows.
        </p>
        <p>The Input and output images are divided into blocks then the blocks are
converted into vectors, two means, two standard derivations and one covariance
value are computed from the images.</p>
        <p>Then the luminance, contrast, and structure comparisons based on statistical
values are computed, the structural similarity index measure is given by:
SSIM (x; y) =
( x + y + c1)( x2 +
(2 x y + c1)( xy + c2)
y2 + c2)</p>
        <p>
          Where x y denotes the mean values of original and distorted images. And
x y denotes the standard deviation of original and distorted images, and xy
is the covariance of both images, c1 and c2 are constants. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>The image quality MSSIM is obtained by calculating the mean of SSIM
values given by:
(3)
(4)</p>
        <p>P
M SSIM = 1 X SSIMj</p>
        <p>P</p>
        <p>j=1</p>
        <p>
          Where p is the number of sliding windows [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Experiments were made using a computer with microprocessor Intel (R) Core
(TM) i5-2410M of 2.30 GHz, 8GB of RAM GPU NVIDIA CUDA GeForce 315M
with a CNN training of 8 to 16 hours.</p>
      <p>For the evaluation of our convolutional neuronal network, the following
images were used: a) Bowling g 4, b) Midd g 5, c) Lamp g 6, d) Monopoly g
7 and e) Baby g 8. The left image and the right image were used as input for
each estimation of the disparity map.</p>
      <p>V. Mendoza et al.</p>
      <p>Bowling left image</p>
      <p>Bowling right image</p>
      <p>Disparity map estimation with deep learning in stereo vision
Lamp left image Lamp right image</p>
      <p>Baby left image</p>
      <p>Baby right image</p>
      <p>In the table 1 the PSNR and SSIM results are shown between the original
disparity map and the estimated disparity Map with our convolutional neuronal
network. With these values it can be seen that PSNR values are not the most
optimal but with the SSIM it shows optimal similarity values between the images.</p>
      <p>In the images 9, 10, 11, 12 and 13 it can be clearly seen how the convolutional
neural network performed in the estimation of the disparity map. The output
images of our network show low de nition of the edges with respect to the original
disparity map, however the accuracy of the estimate has an acceptable level.
In this research, it was demonstrated how a new architecture of a convolutional
network can estimate the disparity map between stereo images. With the
obtained results it can be observed how a post processing of the output images
could help the de nition of the edges in the images, which seems to be the main
problem to be solved as a next step in the investigation in order to obtain results
more precise.</p>
      <p>The limitation of hardware was another problem for this research, the
training times of the convolutional neuronal network was from 8 to 16 hours. In
addition, it can be concluded how applications for stereo vision systems can
be solved by convolutional neural networks, as future work we plan to apply
this neural network to stereo vision in real time video for obstacles detection to
continue searching applications in stereo vision systems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>"Encyclopedia of Science and Technology"</article-title>
          ,
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          ,
          <year>2009</year>
          , pp.
          <fpage>594</fpage>
          -
          <lpage>596</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>M. Z. Brown</surname>
            , D. Burschka,
            <given-names>G. D.</given-names>
          </string-name>
          <string-name>
            <surname>Hager</surname>
          </string-name>
          ,
          <article-title>"Advances in Computational Stereo"</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>25</volume>
          ,
          <year>2003</year>
          , pp.
          <fpage>993</fpage>
          -
          <lpage>1008</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Scharstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Szeliski</surname>
          </string-name>
          , Middlebury Stereo Vision Page, available: http://vision.middlebury.edu/stereo
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R.</given-names>
            <surname>Canals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roussel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Famechon</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Treuillet</surname>
          </string-name>
          ,
          <article-title>A biprocessor-oriented visionbased target tracking system,"</article-title>
          <source>in IEEE Transactions on Industrial Electronics</source>
          , vol.
          <volume>49</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>500</fpage>
          -
          <lpage>506</lpage>
          ,
          <year>Apr 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Approach to Position and Orientation Estimation in Vision-Based UAV Navigation,"</article-title>
          <source>in IEEE Transactions on Aerospace and Electronic Systems</source>
          , vol.
          <volume>46</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>687</fpage>
          -
          <lpage>700</lpage>
          ,
          <year>April 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Shih</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Tsai</surname>
          </string-name>
          , \
          <article-title>A Two-Omni-Camera Stereo Vision System With an Automatic Adaptation Capability to Any System Setup for 3-D Vision Applications," in IEEE Transactions on Circuits and Systems for Video Technology</article-title>
          , vol.
          <volume>23</volume>
          , no.
          <issue>7</issue>
          , pp.
          <fpage>1156</fpage>
          -
          <lpage>1169</lpage>
          ,
          <year>July 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>D.</given-names>
            <surname>Scharstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Szeliski</surname>
          </string-name>
          .
          <article-title>High-accuracy stereo depth maps using structured light</article-title>
          .
          <source>In IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR</source>
          <year>2003</year>
          ), volume
          <volume>1</volume>
          , pages
          <fpage>195</fpage>
          -
          <lpage>202</lpage>
          , Madison,
          <string-name>
            <surname>WI</surname>
          </string-name>
          ,
          <year>June 2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>D.</given-names>
            <surname>Scharstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Pal</surname>
          </string-name>
          .
          <article-title>Learning conditional random elds for stereo</article-title>
          .
          <source>In IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR</source>
          <year>2007</year>
          ), Minneapolis, MN,
          <year>June 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>H.</given-names>
            <surname>Hirschmller</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Scharstein</surname>
          </string-name>
          .
          <article-title>Evaluation of cost functions for stereo matching</article-title>
          .
          <source>In IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR</source>
          <year>2007</year>
          ), Minneapolis, MN, June 2007
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. I.Goodfellow,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          . Deep Learning MIT Press http://www.deeplearningbook.org,
          <year>2016</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A.</given-names>
            <surname>Yusra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Al-Najjar</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Soong</surname>
          </string-name>
          .
          <article-title>Comparison of Image Quality Assessment: PSNR, HVS</article-title>
          , SSIM, UIQI.
          <source>International Journal of Scienti c and Engineering Research</source>
          ,
          <year>August 2012</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>B. Zhou</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Universal</given-names>
            <surname>Image Quality Index</surname>
          </string-name>
          ,
          <source>IEEE Signal Processing Letters</source>
          , vol.
          <volume>9</volume>
          , pp.
          <fpage>81</fpage>
          -
          <lpage>84</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>A.</given-names>
            <surname>Qayyum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Naufal B. Muhammad Saad</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Abdullah</surname>
            and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <article-title>Disparity Map Estimation Based on Optimization Algorithms using Satellite Stereo Imagery</article-title>
          ,
          <source>IEEE International Conference on Signal and Image Processing Applications (ICSIPA)</source>
          , pp.
          <fpage>127</fpage>
          -
          <lpage>132</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. L.
          <string-name>
            <surname>Chung-Hee</surname>
            and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Dongyoung</surname>
          </string-name>
          ,
          <article-title>Dense Disparity Map-based Pedestrian Detection for Intelligent Vehicle</article-title>
          ,
          <source>IEEE International Conference on Intelligent Transportation Engineering</source>
          , pp.
          <fpage>1015</fpage>
          -
          <lpage>1018</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>A. K. Wasim</surname>
            ,
            <given-names>R. P</given-names>
          </string-name>
          <string-name>
            <surname>Dibakar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bhisma</surname>
            and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rasana</surname>
          </string-name>
          , 3D Object Tracking Using DIsparity Map,
          <source>International Conference on Computing Communication and Automation (ICCA2017)</source>
          , pp.
          <fpage>108</fpage>
          -
          <lpage>111</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Okae</surname>
          </string-name>
          ,
          <article-title>Optimization of Stereo Vision Depth Estimation using EdgeBased Disparity Map</article-title>
          ,
          <source>10th International Conference on Electrical and Electronics Engineering</source>
          , pp.
          <fpage>1171</fpage>
          -
          <lpage>1175</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ge</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Lobaton</surname>
          </string-name>
          ,
          <article-title>Obstacle Detection in Outdoor Scenes based on MultiValued Stereo Disparity Maps</article-title>
          ,
          <source>IEEE Symposium Series on Computational Intelligence (SSCI)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Learning for Disparity Estimation through Feature Constancy</article-title>
          ,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>N.</given-names>
            <surname>Mayer</surname>
          </string-name>
          , E. Ilg,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hausser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cremers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Brox</surname>
          </string-name>
          .
          <article-title>A large dataset to train convolutional networks for disparity, optical ow, and scene ow estimation</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>4040</fpage>
          -
          <lpage>4048</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>J. Pang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>J. SJ.</given-names>
          </string-name>
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Cascade Residual Learning: A Two-stage Convolutional Neural Network for Stereo Matching</article-title>
          ,
          <string-name>
            <surname>ICCVW</surname>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>770778</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. G. Song,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Su</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Training a Convolutional Neural Network for Disparity Optimization in Stereo Matching</article-title>
          .
          <source>In Proceedings of the 2017 International Conference on Computational Biology and Bioinformatics (ICCBB</source>
          <year>2017</year>
          ). ACM, New York, NY, USA, pp.
          <fpage>48</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Scharstein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschmller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kitajima</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krathwohl</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nei</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Westling</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>High resolution stereo datasets with subpixel-accurate ground truth</article-title>
          .
          <source>Proc. German Conf. Pattern Recognit</source>
          .
          <source>(GCPR)</source>
          .
          <volume>31</volume>
          -
          <fpage>42</fpage>
          , (Jan.
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zbontar</surname>
            , J. and LeCun,
            <given-names>Y.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Stereo matching by training a convolutional neural network to compare image patches</article-title>
          .
          <source>arXiv: 1510</source>
          .
          <fpage>05970</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <given-names>W.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Schwing</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Urtasun</surname>
          </string-name>
          ,
          <article-title>E cient Deep Learning for Stereo Matching, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas</article-title>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>5695</fpage>
          -
          <lpage>5703</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>