<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Orthophoto Map Feature Extraction Based on Orthophoto Map Feature Extraction Based on Neural Networks Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zdenek Horak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milos Kudelka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vaclav Snasel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vit Vozenilek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zdenek Horak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milos Kudelka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vaclav Snasel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V t Vozen lek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departme1n7t. olifstCoopmadpuut1e5r</institution>
          ,
          <addr-line>S7c0i8en3c3e, OFEstIr,aVvaS-BPo-ruTbeac,hnCizceaclhURnievpeurbsiltiyc of Ostrava</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Departmenttor.f SGveoobgordapyh2y6,</institution>
          ,
          <addr-line>F7a7c1ul4t6y OofloSmcioeunc,e,CPzaelcahckRyepUunbilviecrsity, Olomouc tr. Svobody</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>zednetnoefk.Gheoorgarka</institution>
          ,
          <addr-line>phmyi,lFoasc.ukltuydeolfkSac,ienvacec,lPava.lascnkayseUln</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>216</fpage>
      <lpage>225</lpage>
      <abstract>
        <p>In our paper we use neural networks for tuning of image feature extraction algorithms and for the analysis of orthophoto maps. In our approach we split an aerial photo into a regular grid of segments and for each segment we detect a set of features. These features describe the segment from the viewpoint of general image analysis (color, tint, etc.) as well as from the viewpoint of the shapes in the segment. We also present our computer system that support the process of the validation of extracted features using a neural network. Despite the fact that in our approach we use only general properties of an images, the results of our experiments demonstrate the usefulness of our approach.</p>
      </abstract>
      <kwd-group>
        <kwd>orthophoto map</kwd>
        <kwd>image analysis</kwd>
        <kwd>neural network</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>use a neural network for tuning these algorithms. The features are then detected
automatically. To assess the quality of the detection we use our own application
that incorporates the neural network again. This application visualizes how the
system assesses a particular images. In the case of discovered inaccuracies, the
user can retroactively affect the parameters of automatic feature detection.</p>
      <p>In the following section we discuss related approaches. The third section
contains a description of features in our detection system, while the fourth section
recalls some basics of tools and techniques used. The fifth section is focused on
our experiment with orthophoto maps.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related approaches</title>
      <p>
        The basis for all methods and algorithms for analyzing the orthophoto maps
is digital image processing. Digital image processing is a set of technological
approaches using computer algorithms to perform image processing on digital
images ([
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]). Digital image processing has many advantages over
analogue image processing. It allows a much wider range of algorithms to be applied
to the input data and can avoid problems such as the build-up of noise and
signal distortion during processing ([
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). Digital image processing may be modeled
in the form of multidimensional systems rather than images that are defined
over two dimensions (perhaps more). Some research deals with a new
objectoriented classification method that integrates raster analysis and vector analysis
(e.g. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). They combine the advantages of digital image processing (efficient
improved CSC segmentation), geographical information systems (vector-based
feature selection), and data mining (intelligent SVM classification) to interpret
images from pixels to objects and thematic information.
      </p>
      <p>
        Many different approaches dealing with the detection and extraction of
manmade objects can be found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These are mainly methods focused on
automatic road extraction and automatic building extraction. A summary and
evaluation of methods and approaches from the field of automatic road
extraction can be found in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], while for the field of building extraction see [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. For
more recent approaches from the field of Object-Based Image Analysis (OBIA)
you can see e.g. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a detailed summary of existing methods is described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Authors of this paper have participated in the development of a
commercial Document Management System, which is used in several institutions of the
Government of Czech Republic. Experiments based on image dataset from one
of these institutions are described in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Image features</title>
      <p>Our approach is to describe any image in terms of image contents and in the
concepts which are familiar to the users. In the following we present features we
are capable of detecting. Some of the features are related to the whole image
only, but many of them can also be used to describe some parts of the image.
3.1</p>
      <p>Color features
According to used colors we are able to find out whether the image is gray-scaled,
and if not, whether the image is toned into some specific hue. Also we can say
if the image is light, dark or if the image is cool or warm.</p>
      <p>– grey-scaled images
– color-toned images
– bright or dark images
– images with cool or warm color tones</p>
      <p>
        The last group of features is color features. We want to describe the image in
terms of colors in the same way as a human will, but it does not suffice only to
count the ratio of one color in the image or in some area of the image. A more
complex histogram is also not enough. We should consider things like dithering,
JPEG artifacts and the subjective perception of colors by people. Using color
spatial distribution, color histograms and below mentioned shapes recognition
we are also able to detect the background color. The colors we are currently able
to detect are:
– red, green, blue, yellow, turquoise, violet, orange, pink, brown, beige, black,
white and gray
– background color
Color features detection Low-level color features were detected using a
combination of their spatial distribution and a comparison with their prototypes. The
first version of the system contained prototypes that were constructed manually
using our subjective perception. However this approach was not general enough,
therefore we have created a set of training image patterns with manually
annotated color features. To deal with human perception we have averaged the
annotation results among several annotators. Using this set we have trained the
artificial single-layer feed-forward neural network (see [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) to confidently
identify the mentioned features. This network was very similar to the network
used in the whole application, which is described below in detail.
      </p>
      <p>As an input we have used the pixels of particular patterns in different color
models (as different models are suitable for different color features). Trained
neurons (their input weights and hidden threshold) were then transformed (using
the most successfull color model) into the color feature prototypes (see fig. 1). We
detect all of the mentioned features as fuzzy degrees, but for selected applications
we scale them down to the binary case.</p>
      <p>
        Image segmentation To obtain more precise information about the processed
image, we have decided to employee an image segmentation technique. Using
the Flood fill algorithm (with eight directions, for details see [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) we were able to
separate regions with same (or almost same) color. But to be able to index these
shapes, we need to describe them. We have calculated the center of this shape
and using this point and different angles we have sliced the shape into several
regions (see 16 regions in figures 3, 4, 5). For each region, we have computed
the maximum distance from the center. Following the changes of this distance
(peaks, regularity) we are able to distinguish between different basic shapes
(rectangle, circle, triangle, etc.).
      </p>
      <p>Of course, this approach is not general. We use it only for bigger shapes and
we ignore possible holes within the shapes. Because we use mostly downsampled
versions of source images, we can guarantee the effectivness of processing. And
because we use high quality downsampling, our results are similar to a person’s
first glance. The shapes we are able to detect are: line, restangle, circle, triangle
and quad.</p>
      <p>At this moment, we detect shapes separately, but to the resulting description
we save only information, whether at least one shape of such kind has been
detected (i.e. the image contains one triangle) or whether there are multiple
shapes of such kind (i.e. the image contains more triangles).</p>
      <p>Currently we are thinking of using obtained distances not only for shape
identification, but also for shape description. The same shape can be scaled,
moved or rotated on different images, but the description using relative distance
changes is still the same (up to index rotation). The more different angles we
use, the more precise description we obtain.</p>
      <p>
        Anomalies Since the orthophoto maps are created from long distance, interesting
objects are often relatively small and vaguely bound in the image. For this reason
we have incorporated the concept of anomalies (see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for a recent survey).
As an anomaly we consider:
– a shape formed by similar pixels,
– which – due its size – cannot be reliably classified as being one of the
previously mentioned shapes and
– has other than background color.
      </p>
      <p>As you will see in the experiment section, this concept became very important
in our approach. Figure 2 contains highlighted samples of various shapes detected
in the orthophoto maps.
The Neural network (or more precisely artificial neural network) is a
computational model inspired by biological processes. This network consists of
interconnected artificial neurons which transform excitation of input synapses to output
excitation. Most of the neural networks can adapt themselves. There are many
different variants of neural networks. Each variant is specific in its structure
(whether the neurons are organized in some layers, whether the neurons can be
connected to themselves, etc.), learning method (the way the neural network is
adapted) and neuron activation function (the way the neuron transforms input
excitation to output excitation) and its parameters.</p>
      <p>
        For our purposes we use a structure consisting of an input layer of
neurons, several inner hidden neuron layers and one output layer of neurons. As the
learning method we use supervised learning, where the network is presented
repeatedly with specific samples, which are propagated towards the network
output. This output is compared with expected results and the network is (using
calculated error) adapted to minimize this error. The learning finishes after a
predefined number of learning epochs or if the error rate decreases under a
predefined constant. After the learning phase, the neural network can be presented
with another group of samples and provides its output. For more details on
neural networks consult [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or see [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] for this particular case.
      </p>
      <p>Our particular network is illustrated in figure 8. We have decided to use
classical bipolar-sigmoid (because we needed to represent both positive and negative
examples) as an activation function of neurons (having β = 2):
Simple backpropagation has been used as a learning algorithm:
2
f (x) = 1 + e−βx − 1
4wij (t + 1) = η
σE
σwij
+ α4wij (t)</p>
      <p>The basic idea of this algorithm is to calculate the total error E of the network
(computed by comparing real outputs of the network with expected ones) and
then change the weights 4wij (t + 1) of the network to minimize this error. The
learning rate parameter η controls the speed of weight changes. To speed up
learning, we use momentum α – which updates the weight in each step also with
the value from the previous step 4wij (t).
5</p>
    </sec>
    <sec id="sec-4">
      <title>Feature validation</title>
      <p>To verify that our set of features is capable of representing the user point of
view on the images content, we have created a web application for image
suggestion. In the first step, the user is presented with several random images from
the data. He/she marks these images as interesting (or not interesting). Using
this process the user search profile is created. In the second step the application
tries to understand this profile (a set of positive and negative examples) using
an artificial multilayer feed-forward neural network. In the last step, the trained
network is presented with the whole dataset and suggests images which may be
potentially interesting to the user.</p>
      <p>The user can clarify his/her profile by marking further images and the process
is repeated. We have used part of the profile for training and the rest for the
validation of the profile to verify the meaningfulness of this profile. The score of
presented images is an indication for users to add more positive (if the overall
score is too low) or negative (the overall score being too high) examples.
Network parameters Parameters of the neural network have been selected as
follows (see fig. 8): 610 input neurons (input activation represents the degree
of individual feature presence), 5 hidden neurons and one output neuron
(representing the degree of image acceptance). The learning rate was η = 0.1 and
momentum α = 0.1. The maximum number of iterations per learning epoch was
set to 1,000. The number of input neurons correspond to the number of features
in different regions of the image. Remaining parameters have been selected
after several attempts of being subjectively the best. A larger number of hidden
neurons often caused the overtraining of the network (good performance on the
training samples with very limited ability of generalization). A larger number of
iterations produced no significant improvement. Lower values failed to comply
with user judgments.</p>
      <p>Application description This application (see fig. 8) has been created as an
ASP.NET Web application on the Microsoft .NET platform utilizing several
other technologies such as CSS/JavaScript to improve user experience. Most of
the computation time is used during the image dataset indexing, which is done
only once and can be precomputed offline. The indexing of particular image takes
on average 0.76 seconds and can be easily paralelized as the indexing of every
particular image is a completely independent.</p>
      <p>The neural network is recreated with every request, but in high-load
environment can be stored between the requests. Application memory contains indexed
image signatures only, therefore the whole application is well scalable. In our
testing environment we have been able to run this application easily on an Intel
2.13 GHz processor and 4 GB RAM with a dataset containing several thousand
images. Clearly the process of running the neural network with particular image
signatures in every step of recommendation has its computational limits, but
these limits lie far beyond the boundaries of the purpose of our experiment.</p>
      <p>We have performed several user testing sessions where we have selected the
presented set of features as being the most suitable for our purposes. Using this
process we made sure that normal users can understand image analysis systems
based on selected features and these users were normally able to find expected
results after giving two or three positive and negative examples. The discrepancy
between the user’s expectations and the output of the system is a suggestion for
another iteration of feature detection tuning.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and future work</title>
      <p>Using several mathematical models and methods, such as neural networks, we
have developed and described a system which can analyze orthophoto maps,
detect user-oriented features in the maps and visualize the structure of the region.
In the future work we will investigate the similarities between different image
segments and sources of these similarities. We would like to consider different
kinds of maps and use our approach on a much wider landscape region.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment</title>
      <p>This work is supported by Grant of Grant Agency of Czech Republic No.
205/09/1079.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arbib</surname>
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>The handbook of brain theory and neural networks</article-title>
          , The MIT Press (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baltsavias</surname>
            <given-names>E.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruen</surname>
            <given-names>A</given-names>
          </string-name>
          .,
          <string-name>
            <surname>Van Gool L.</surname>
          </string-name>
          :
          <article-title>Automatic extraction of man-made objects from aerial and space images (III) (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Blaschke</surname>
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Object based image analysis for remote sensing</article-title>
          ,
          <source>ISPRS Journal of Photogrammetry and Remote Sensing, Elsevier</source>
          , vol.
          <volume>65</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>2</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Blaschke</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lang</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hay</surname>
            <given-names>G.J.</given-names>
          </string-name>
          :
          <article-title>Object-based image analysis: spatial concepts for knowledge-driven remote sensing applications</article-title>
          , Springer Verlag (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chandola</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Anomaly detection: A survey</article-title>
          ,
          <source>ACM Computing Surveys (CSUR)</source>
          , vol.
          <volume>41</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>1</fpage>
          -
          <lpage>58</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Glassner</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Fill'Er Up</surname>
          </string-name>
          !,
          <source>IEEE Computer Graphics and Applications</source>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>85</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gonzalez</surname>
            <given-names>R. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woods R</surname>
          </string-name>
          . E.:
          <article-title>Digital image processing</article-title>
          . London, Pearson Education,
          <volume>954</volume>
          <fpage>pages</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Horak
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Kudelka</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Snasel</surname>
          </string-name>
          <string-name>
            <surname>V.</surname>
          </string-name>
          :
          <article-title>FCA as a Tool for Inaccuracy Detection in Content-Based Image Analysis</article-title>
          ,
          <source>IEEE International Conference on Granular Computing</source>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>228</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lee</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seung</surname>
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Algorithms for Non-Negative Matrix Factorization</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          , vol.
          <volume>13</volume>
          , pp.
          <fpage>556</fpage>
          -
          <lpage>562</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Li</surname>
            <given-names>H. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gu H</surname>
          </string-name>
          . Y.,
          <string-name>
            <surname>Han</surname>
            <given-names>Y. S.</given-names>
          </string-name>
          et al.:
          <article-title>Object-oriented classification of high-resolution remote sensing imagery based on an improved colour structure code and a support vector machine</article-title>
          ,
          <source>International journal of remote sensing</source>
          , vol.
          <volume>31</volume>
          (
          <issue>6</issue>
          ), pp.
          <fpage>1453</fpage>
          -
          <lpage>1470</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lillesand</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiefer</surname>
          </string-name>
          , R.:
          <source>Remote Sensing and Image Interpretation</source>
          , New York, John Wiley and Sons, 724 pages (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mayer</surname>
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Automatic object extraction from aerial imagery-A survey focusing on buildings, Computer vision</article-title>
          and image understanding,
          <source>Elsevier</source>
          , vol.
          <volume>74</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>138</fpage>
          -
          <lpage>149</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mena</surname>
          </string-name>
          , JB:
          <article-title>State of the art on automatic road extraction for GIS update: a novel classification</article-title>
          ,
          <source>Pattern Recognition Letters</source>
          , vol.
          <volume>24</volume>
          (
          <issue>16</issue>
          ), pp.
          <fpage>3037</fpage>
          -
          <lpage>3058</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Pitas</surname>
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Digital image processing algorithms</article-title>
          and applications, New York, Wiley, 419 pages (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Riedmiller</surname>
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Advanced supervised learning in multi-layer perceptrons-From backpropagation to adaptive learning algorithms</article-title>
          ,
          <source>Computer Standards &amp; Interfaces</source>
          , vol.
          <volume>16</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>265</fpage>
          -
          <lpage>278</lpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Umbaugh S</surname>
          </string-name>
          . E.:
          <article-title>Computer imaging: digital image analysis and processing</article-title>
          , London, Taylor &amp; Francis, 659 pages (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Yegnanarayana</surname>
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Artificial neural networks</article-title>
          ,
          <source>PHI Learning Pvt. Ltd</source>
          . (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>