<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>COMBINING SATELLITE IMAGERY AND MACHINE LEARNING TO PREDICT ATMOSPHERIC HEAVY METAL CONTAMINATION</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Uzhinskiy</string-name>
          <email>auzhinskiy@jinr.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gennady Ososkov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Goncharov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marina Frontsyeva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Joint Institute for Nuclear Research</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Sukhoi State Technical University of Gomel</institution>
          ,
          <country country="BY">Belarus</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>351</fpage>
      <lpage>358</lpage>
      <abstract>
        <p>Air pollution has a significant impact on the European and Asian countries. More than nine out of 10 of the world's population - 92%, lives in places where the air pollution exceeds safe limits, according to the research of the World Health Organization. There are a lot of regional and international environment control programs. They use different techniques and tools but, as a result, they all intend to understand what is the current situation and how does it evolve. Generally, to get some indexes, researches take samples and analyze them. For the natural reasons, sampling is carried out rarely, and the dimension of the sampling grid can be very big. In such a situation, modeling can be a right choice. In our research, we have tried to predict atmospheric heavy metal contamination combining satellite images and machine learning. Data sources for model training were satellite images from Google Earth Engine platform and sampling data from Data Management System of UNECE International Cooperative Program (ICP) Vegetation. We obtained satisfactory results in the prediction of Sb for Norway and Mn for Serbia.</p>
      </abstract>
      <kwd-group>
        <kwd>prediction</kwd>
        <kwd>ecological monitoring</kwd>
        <kwd>satellite images</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Air pollution is the fourth largest threat to human health, behind high blood pressure, dietary
risks and smoking. The health risks of breathing dirty air include respiratory infections and
cardiovascular diseases, stroke, chronic lung diseases, and lung cancer. In the study of the World Bank
and the Institute for Health Metrics and Evaluation (IHME) the economic cost of air pollution was
calculated. It was found that air pollution led to one of 10 deaths in 2013, which cost for the global
economy about $225 billion in loss of labor income [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>There are many regional and international environment control programs. They use different
techniques and tools but as a result, they all intend to understand what is the current situation and how
does it evolve. Generally, to get some indexes researches take samples and analyze them. For the
objective reasons a sampling is carried out rarely, and the dimension of the sampling grid can be very
big. In such a situation, modeling can be a right choice. Our idea is to use real-life information about
heavy metals concentration and indexes taken from the satellite images to train the special statistical
model. After that, the model together with the new satellite indexes can be used to predict
contamination for any part of the study area at any time.</p>
      <p>It is clear that getting the indexes from the satellite images is a much easier process than a
field sampling. A sampling and analysis of an area like Moscow Region can take 4-6 month.
Gathering of indexes for the same area can be done for a few days.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data sources for model training</title>
      <p>To obtain reliable results we should have real contamination data and additional parameters.
The sources of the contamination data are environment control programs. The most perspective source
of the additional parameters for the models is satellite images done in various spectra.</p>
      <sec id="sec-2-1">
        <title>2.1. UNECE ICP Vegetation</title>
        <p>The aim of the UNECE International Cooperative Program (ICP) Vegetation in the framework
of the United Nations Convention on Long-Range Transboundary Air Pollution (CLRTAP) is to
identify the main polluted areas of Europe, produce regional maps and further to develop the
understanding of the long-range transboundary pollution [2] . Atmospheric deposition study of heavy
metals, nitrogen, persistent organic compounds (POPs) and radionuclides is based on the analysis of
naturally growing mosses through moss surveys carried out every 5 years [3]. Nowadays the UNECE
International Cooperative Program (ICP) Vegetation is realized in 39 countries of Europe and Asia.
Mosses are collected at thousands of sites across Europe, and their heavy metals (since 1990), nitrogen
(since 2005), POPs (persistent organic compounds, pilot study in 2010) and radionuclides (since 2015)
concentrations are determined. A total of 13 elements are reported for the Atlas (As, Cd, Cr, Cu, Fe,
Hg, Ni, Pb, V, Zn, Al, Sb, and N). Results are reported as a number of sampling sites, minimum,
maximum and median concentrations in mg/kg. Specialists of the Joint Institute for Nuclear Research
(JINR) developed a cloud platform (ICP Vegetation Data Management System, DMS, dms.jinr.ru)
consisting of a set of interconnected services. The platform provides ICP Vegetation participants with
the modern unified system of collecting, analyzing and processing of biological monitoring data [4].
More than 6000 sampling sites from 47 regions of different countries are presented at the DMS now
from the Moss Survay 2015-2016.</p>
        <p>As the developers of the DMS, we have an agreement with ICP Vegetation participants. Due
to this circumstance, we can use these data in our research.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Google Earth Engine</title>
        <p>Satellite programs like LandSat, MODIS, Sentinel provides free access to their data. One can
search their database and find necessary images. Special software such as ENVI or ERDAS can be
used to process images after that. Such an approach is not too comfortable, because images are of
gigabyte size, and we should have a few of them to cover the region. Some software exists to search
through the image archives, but the functionality of these programs is rather poor, and they often work
with only one satellite data source.</p>
        <p>
          We used Google Earth Engine (GEE) – a cloud-based platform for the planetary-scale
environmental data analysis to get satellite image indexes. The purpose of the Earth Engine is to
perform highly interactive algorithm development at global scale, push the edge of the envelope for
Big Data in remote sensing, enable high-impact, data-driven science and make substantive progress on
global challenges that involve large geospatial datasets [
          <xref ref-type="bibr" rid="ref2">5</xref>
          ]. There are more than 100 satellite programs
and modeled datasets. GEE has the JavaScript online editor to create and verify code online and
Python API to communicate with user’s applications.
        </p>
        <p>There are as common programs presented at GEE - Landsat, Modis, Sentinel, as pretty
specific - the MOD11A2 V6 product that provides an average 8-day land surface temperature (LST) in
a 1200 x 1200 kilometer grid, VIIRS Nighttime Day/Night - Monthly average radiance composite
images using nighttime data from the Visible Infrared Imaging Radiometer Suite (VIIRS) Day/Night
Band (DNB) or the MOD13A2 V6 – product that provides two Vegetation Indices (VI): the
Normalized Difference Vegetation Index (NDVI) and the Enhanced Vegetation Index (EVI), see
examples at fig 1.</p>
        <p>To calculate the satellite image index with GEE, we undertake a few steps. We define a
region, for example, 2 square kilometers with the center at the known sampling site point. Then we
choose satellite program/product, for example, MODIS/006/MOD09A1 and define some filters like
dates, weather, region, etc. After that, we get a collection of images and combine them with the
median function. At the moment, we have an image of the region. For the image, we have a set of
spectral channels, and GEE provides functionality to execute some mathematical functions (max, min,
median, etc) for each channel. For example, we can calculate max(NDVI) for our 2 km2 area based on
MODIS satellite images.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Correlation of the contamination and satellite images</title>
      <p>We have created a piece of software that takes coordinates of the sampling site from DMS and
calculates indexes from different satellite program images. Then the correlation between the
contamination in a region and indexes is defined. We had analyzed information for seven countries
where the number of sampling sites were more than 200: Norway, France, Germany, Sweden,
Rumania, Serbia, and Iceland. We used more than 20 satellite programs/products, and different types
of indexes for them to find a correlation. We see that connection between elements and indexes varies
from region to region, and we can find several regions where few elements have correlations more
than 0.5 with the number of indexes, see table 1.</p>
      <p>Nature of such correlations can be an objective of the independent study. We consider the
following possibilities: 1. Strait influence of aerosol heavy metal particles on the images in exact
spectrum. 2. Indirect connection through technogenic and anthropogenic factors. 3. Indirect
connection through landscape features and vegetation characters. 4. The mix of first three factors.</p>
      <p>To use some statistical methods for prediction, we should find for an element at some region
six or more indexes which have a satisfactory connection with element concentration, but a weak
cross-correlation. This is a complicated task because different programs can make images in analogous
or similar spectra, so the satellite image indexes of such programs can be strongly correlated. We
manage to find eight or more indexes satisfactory correlated with Sb for Norway, Mn for Serbia and U
at Romania.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Machine learning methods and algorithms</title>
      <p>We used two types of statistical approaches: regression and classification. For each of them,
we tried a tree-based model and artificial neural networks. Tree-based models include gradient
boosting, decision trees, random forest and bagging [6]. The neural network approach assumes a usage
of multi-layer perceptron with two hidden layers.</p>
      <p>
        The goal of the regression task is to predict a single output, which represents the
concentration. The common metric for regression tasks is the mean-squared error. The output layer of
the neural network regressor is a single neuron with linear activation. The linear activation is used
because concentration values are unbounded. We can reduce the regression problem to the
classification task because a prediction of the exact concentration value in the particular point is not
mandatory. The K-Nearest Neighbours (KNN) [
        <xref ref-type="bibr" rid="ref3">7</xref>
        ] algorithm was trained on the data of contamination
for selection K data clusters, each of which corresponds to a particular level of contamination.
Contamination values were replaced with the labels of corresponding clusters after the training of the
KNN to pass them to the input of classifiers. The output layer of neural network classifier has K
neurons equal to the number of clusters. Softmax activation on the output layer of the neural network
classifier computes the distribution among classes to identify which of them can be chosen as model’s
response depending on the input data. We use the cross-entropy loss as the minimized metric during
the training process of the neural network classifier. The predicted labels indicate the level of the
contamination. We also tried to add the weights of the classes, but it did not have any significant
impact on the prediction efficiency.
      </p>
      <p>
        To find optimal parameters for the tree-based models we performed a special procedure
named a grid-search cross-validation. This procedure consists of a total check of all possible
combinations of the parameters and finding one, which allows obtaining a minimal cross-validation
loss. Training data were permuted and splited into ten equal portions. Nine of them were used for
training and the remaining one for the evaluation. We repeated the procedure ten times per each
training epoch. The finding of the optimal parameters of the neural networks even with only two
hidden layers is a very time-consuming task. Thus, for the parameters selection of the neural network
models, we used Tree-structured Parzen Estimator Approach (TPE). TPE allows doing a smaller
number of iterations than the grid-search, while the results are better than random search [
        <xref ref-type="bibr" rid="ref4">8</xref>
        ]. The full
list of explored parameters is available at [9].
      </p>
      <p>In addition, we tried a few new models and improvements: we applied the MinMax
normalization not only to the input values, but also to the concentrations – it allowed us to minimize
the binary-cross entropy loss in the case of neural networks regressor instead of the mean-squared
error method which works well only when dealing with the normally distributed data. Also, we tried to
use the robust scaling of the input features. It performs a subtraction of the median value and scales
the data according to the quantile range (defaults to IQR: InterQuantile Range). The IQR is the range
between the 1st quartile (25th quantile) and the 3rd quartile (75th quantile). The value of the
meansquared error for the best regression algorithms measured on the test subset of the data was about
0.0035 (± 0.0015).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Modeling results</title>
      <p>Having the models, we tried to predict the contamination for some regions. We focused on Sb
for Norway and Mn for Serbia. One can see in fig. 2 a color-graduated map of Sb distribution in
Norway and Mn distribution in Serbia. Various сolors represent concentrations in mg/kg of the
element in taken at that point sample. In table 2, one can also see some statistics about element
distribution.</p>
      <p>Sb sources are mostly anthropogenic ones (traffic and industrial). We have a few programs in
the list that represent temperature, radiance and other signs of the human population. 8 indexes are
given in table 3, that we choose to train the model for Norway. Sources of Mn for Serbia could have a
natural, agricultural, or industrial origin. As one can see, we used much smaller analyzed areas in this
experiment than for Sb in Norway. In table 4, for 9 indexes are given that we choose to train our
model for Serbia. We have removed two indexes with correlation ~0.52 from Serbia experiment to get
better results.</p>
      <p>We trained models on the real-life data and satellite image indexes. After that, we calculated
the new indexes to get the prediction. We had 1198 points at the south part of Norway. We applied our
models, and gradient boosting regressor gives the best results, which presented in fig 3.</p>
      <p>One can see that the model shows tendencies very well. We tried the same model on the north
part of Norway and the results were the same.</p>
      <p>We have calculated indexes for 958 points covering the whole Serbia and some trans-border
region. We used an automation to choose the best model. To the first set of 742 points, we added 216
points geographically close to the sampling sites points (their longitude was moved on 0.001 aside).
Then we done 1000 iterations of training and prediction. On each step, we verified is the new model
statistically closer to the real-life data. At the end, we chose the best model from different model types.
Gradient boosting regressor gave the best result, see fig. 4.</p>
      <p>The model shows tendencies and peak values clearly. If we can find indexes with better
correlation with real data, the accuracy of the model can be increased even more.</p>
      <p>We made also some experiments with U for Romania but results were not satisfying, in our
opinion. We are going further to investigate where the problem is in indexes, or it is in models.</p>
    </sec>
    <sec id="sec-6">
      <title>4. Conclusion</title>
      <p>Indexes from satellite images combined with the special statistical models can be used to
predict the atmospheric contamination of some heavy metals in some regions. The connection between
elements and indexes varies from region to region, so there is no unified solution or a unified model.
Modeling can open new horizons for the contamination analysis. Researchers will be able to monitor
the evaluation of the situation when it is needed, get detailed information about the areas of interests,
check the situation in the cross-border areas, partly automate environment control process
(automatically run the model and get the notification when the contamination level is higher than the
critical level).</p>
      <p>We keep on searching connection of the contamination and satellite indexes and testing new
models and approaches.
[9] Uzhinskiy A., Ososkov G., Goncharov P., Frontasyeva M., Perspektivy ispol'zovaniya
kosmosnimkov dlya prognozirovaniya zagryazneniya vozdukha tyazhelymi metallami [Perspectives of
using a satellite imagery data for prediction of heavy metals contamination] //
Computer Research and Modeling, 2018, vol. 10, no. 4, pp. 535-544, ISSN 2076-7633</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Rosamond</given-names>
            <surname>Hutt</surname>
          </string-name>
          , Keith Breene, 7 shocking facts about air pollution//World Economic Forum,
          <year>2016</year>
          . Available at: https://www.weforum.org/agenda/2016/10/air
          <article-title>-pollution-the-true-cost-in-numbers/ [2] United Nations Economic Commission for Europe (UNECE) Conventions</article-title>
          and Protocols [Electronic resource]: http://www.unece.org/env/treaties/welcome.html.
          <source>(Accessed</source>
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
          <year>2018</year>
          ) [3]
          <string-name>
            <given-names>UNECE</given-names>
            <surname>International Cooperative</surname>
          </string-name>
          <article-title>Programme on Effects of Air Pollution on Natural Vegetation</article-title>
          and Crops [Electronic resource]: http://icpvegetation.ceh.ac.uk/.
          <source>(Accessed</source>
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
          <year>2018</year>
          ) [4]
          <string-name>
            <surname>Ososkov</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frontasyeva</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uzhinskiy</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kutovskiy</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nechaevsky</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <article-title>Cloud platform for data management of the environmental monitoring network:</article-title>
          <source>UNECE ICP Vegetation case // CEUR Workshop Proceedings</source>
          , Vol.
          <volume>1787</volume>
          ,
          <year>2016</year>
          , Pages
          <fpage>224</fpage>
          -
          <lpage>229</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Google</given-names>
            <surname>Earth Engine</surname>
          </string-name>
          [Electronic resource]: https://earthengine.google.com/.
          <source>(Accessed</source>
          <volume>1</volume>
          .
          <fpage>10</fpage>
          .
          <year>2018</year>
          ) [6]
          <string-name>
            <surname>Hastie</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Trees Bagging Random Forests</surname>
          </string-name>
          and Boosting // Stanford University -
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Alsabti</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ranka</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            <given-names>V.</given-names>
          </string-name>
          <article-title>An efficient k-means clustering algorithm - 1997.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bergstra</surname>
            <given-names>J. S.</given-names>
          </string-name>
          et al.
          <source>Algorithms for hyper-parameter optimization // Advances in neural information processing systems</source>
          .
          <source>- 2011</source>
          . - P.
          <fpage>2546</fpage>
          -
          <lpage>2554</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>