<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Machine Learning Approach to Monitor Air Quality from Tra c and Weather data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudio Rossi LINKS Foundation</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy claudio.rossi@linksfoundation.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giacomo Falcone LINKS Foundation</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy giacomo.falcone@linksfoundation.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alessandro Farasin Politecnico di Torino, Italy &amp; LINKS Foundation</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Carlotta Castelluccio Microsoft</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowing the amount of air pollutants in our cities is of great importance to help decision makers in the de nition of e ective strategies aimed at maintaining a good air quality, which is a key factor for a healthy life, especially in urban environments. Using a data set from a big metropolitan city, we realize the uAQE: urban Air Quality Evaluator, which is a supervised machine learning model able to estimate air pollutants values using only weather and tra c data. We evaluate the performance of our solution by comparing the predicted pollutant values with the real measurements provided by professional air monitoring stations. We use the predicted pollutants to compute a standard Air Quality Index (AQI) and we map it into a set of ve qualitative AQI classes, which can be used for decision making at the city level. uAQE is able to predict the AQI class value with an accuracy of 0.8.</p>
      </abstract>
      <kwd-group>
        <kwd>Air Quality</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>BRNN</kwd>
        <kwd>Weather</kwd>
        <kwd>Tra c</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Works</title>
      <p>Air pollution introduces into the atmosphere chemicals, particulates, or biological materials that causes
discomfort, disease, or death to humans and to other living organisms alike. More than 5.5 million people worldwide
are dying prematurely every year as a result of air pollution exposure [1]. This fact con rms that air pollution
is one of the world's largest environmental health risks. Most of these deaths are occurring in rapidly developing
economies, e.g., China and India, but also in European metropolitan cities, e.g., Naples, Turin and Milan, which
have an air pollution index among the highest ones according to recent rankings 1.</p>
      <p>Road transport is one of the main causes of air pollutants emissions, accounting for the 14% of the total
emissions in European countries 2. In recent years, advanced after-treatment technology (Particulate Matter traps)
Copyright © by the paper's authors. Use permitted under Creative Commons License Attribution 4.0 International (CC
BY 4.0).
have been implemented by car manufacturers in response to increasingly tight European emission policies.
Consequently, emissions of the main regulated pollutants from road transport, i.e., Nitrogen Oxides (N Ox) and
Particulate Matter (PM), have been reduced despite the increased vehicle activity. Speci cally, urban N Ox
emissions from tra c has been reduced by 16% between 2000 and 2010, mainly due to the introduction of the
Euro 4 and the Euro 5 standards for passenger cars (both petrol and diesel), which were applied from early 2005
and late 2009, respectively [3].</p>
      <p>Other human activities having a strong impact on air quality are industrial processes, farming, heat and air
conditioning, and other types of transport (trains, airplanes,etc.).</p>
      <p>It is a well-known fact that weather phenomena have a strong impact on air pollutants because once pollutants
are emitted into the air, they propagate into the atmosphere according to weather conditions, e.g., turbulence
mixes pollutants into the surrounding air, and wind carries them away from the source location. Conversely, when
the air near the surface of the earth is cooler than the air above (a phenomenon called temperature inversion)
there is very little air mixing. Since cool air is heavy, it will not to move up to mix with the warmer air above.
Thus, any pollutants released near the surface will get trapped and build up in the cooler air layer.
Municipalities struggle to predict the e ect of tra c policies, e.g., total tra c block, stop of most pollutant
vehicles, on the air quality because there is a lack of easy-to-use tools that can estimate the air pollution taking into
account also the meteorological predictions. Furthermore, the availability of air quality measurement stations
in a city is very limited due to economic constrains. A professional station requires a non negligible investment
(about 200k e per installation) and it has a high maintenance cost (about 30ke per year) [2].
Because of its importance, the estimation of the air quality has been subject to some studies. In [2], Microsoft
researchers proposed a semi supervised learning approach able to predict P M10 and Nitrogen Dioxide (N O2)
emissions at an higher spatial resolution with respect to the one achieved by the installed air quality sensors by
coupling other data sources such as tra c ows, the structure of the road network, meteorological conditions
and point of interest locations. Their solution is complementary to ours, and it can can be used to improve the
spatial resolution of the uAQE. Other relevant studies include the [4] and [5], which present a set of learning
methods able to predict N Ox concentrations from past observations and weather conditions. In [6], the authors
studied Delhi's P M2:5 concentrations and its correlation with the vehicular tra c and with the weather
conditions. However, the proposed model makes several empirical assumptions and it includes parameters speci c to
the city of Delhi. Hence, it cannot be re-used for our purpose.</p>
      <p>To help decision maker in keeping under control the air quality we propose uAQE: urban Air Quality Evaluator,
which is a set of supervised machine learning model able to predict air pollutants values in a urban environment
using only weather and tra c data. We train our models with data taken from Milan, building one model for
each air pollutants. Our work is di erent from all the above mentioned approaches because we aim to predict
pollutants without requiring data from air quality stations. Note that we train one model for each air
pollutants, namely Nitrogen Dioxide (N O2), Ozone (O3), Carbon Monoxide (CO), Benzene (C6H6), Total Nitrogen
(N2), Particulate Matter (P M10), Sulfur Dioxide (SO2), Particulate Matter (P M2:5), Black Carbon (BC), and
Ammonia (N H3). We present the accuracy of each model using the pollutants as measured by professional air
stations. Following a regional standard, we use the predicted pollutants to compute an Air Quality Index (AQI)
which is then mapped it into a set of ve qualitative classes that are used to manage air quality policies at city
level. We nally asses the classi cation accuracy achieved by uAQE obtaining a value of 0.8.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Input data</title>
      <p>Our data has been collected in the city of Milan during two months (Nov.- Dec. 2013), and it contains three
distinct data categories:
{ Weather: we have six di erent weather stations placed within the city limit. Each station has a unique ID,
type, location, and it features a set of co-located sensors. Each sensor measures a di erent meteorological
phenomena. This information has been obtained thanks to ARPA (Agenzia Regionale per la Protezione
dell'Ambiente) 3.
{ Tra c: through xed video cameras already installed for tra c access control at 52 locations in the central
area of Milan (Cerchia dei Bastioni) the local authority obtained the plate number of transiting vehicles,
from which the vehicle characteristic could be extracted from the o cial database, i.e., the Motorizzazione
civile, which holds the information of all Italian vehicles. Note that we received anonymized data, i.e., with</p>
      <sec id="sec-2-1">
        <title>3 http://ita.arpalombardia.it/ITA/qaria/doc RichiestaDati.asp</title>
        <p>hashed plate numbers and with no information about the vehicle owner. Therefore, only the technical details
of each vehicle has been made available to us. These data have been provided as open data by the city of
Milan.
{ Air: we take the measurements of three di erent air stations located within the city limits. Each station
features multiple co-located sensors, each of which measures a single air pollutant. Also these measurements
are directly provided as open data by ARPA, who is the o cial source of this kind of data.
The locations of weather stations, air stations and xed cameras are shown in Fig. 1.</p>
        <p>The weather station data contain wind direction (degree), wind speed (m/s), temperature (Celsius degree),
relative humidity (%), precipitation (mm), global radiation ( W=m2), net radiation ( W=m2), and atmospheric
pressure (hPa). The tra c data include each vehicle passage at each gate, for which the location and the
timestamp of each passage is known. For each passage, the vehicle characteristics are given, namely the European
emission standard category (EURO category from 1 to 6), the vehicle type (i.e., bus, freight, transport, people
transport or not available), the fuel type (i.e., petrol, diesel, electric, LPG, hybrid or missing), the presence of
the Diesel Particle Filter (DPF) and the vehicle length expressed in mm.</p>
        <p>The air pollution data contain ten di erent agents: N O2 ( g=m3), N H3 ( g=m3 ), C6H6 ( g=m3), SO2
( g=m3), BC ( g=m3), CO ( g=m3), N2 (ppb), P M10 ( g=m3), P M2:5 ( g=m3), O3 ( g=m3).
We compute the hourly air quality index as de ned by Piedmont index AQI because is the only example of
operational use of an air quality index in Italy4. The AQI uses only three pollutants, namely N O2, P M10, O3,
and it is formulated as follows:</p>
        <p>IP M10 =</p>
        <p>INO2 =
I8hO3 =</p>
        <sec id="sec-2-1-1">
          <title>Vmed24hP M10</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>VrifP M10</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>VmaxhNO2</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>VrifNO2</title>
        </sec>
        <sec id="sec-2-1-5">
          <title>Vmax8hO3</title>
          <p>Vrif8hO3
100
100
100
IAQI =</p>
          <p>
            IP M10 + max(INO2 ; IO3 )
2
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            )
(
            <xref ref-type="bibr" rid="ref4">4</xref>
            )
We observe that in the considered data set O3 never exceeds the maximum value established for preserving
human health (i.e., 120 g=m3), whereas N O2 exceeds its hourly maximum value (i.e., 200 g=m3) only in few
cases (&lt; 5%). Conversely, P M10 exceeds the the daily maximum value (i.e., 50 g=m3) in 50% in the cases.
We map the computed AQI in the ve classes de ned by the Piedmont region, namely Optimal (0 AQI &lt; 50),
Good (50 AQI &lt; 75), Fair (75 AQI &lt; 100), Average (100 AQI &lt; 125), Not Very Healthy (125 AQI &lt;
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>4 http://www.arpae.it/cms3/documenti/aria/IQA.pdf</title>
        <p>150), Unhealthy (150 AQI &lt; 175), Very Unhealthy (AQI 175). In our data set, we observe that there are
no AQI values in the Optimal level and very few in the Good one, while the most part of values ( 80%) fall
between Fair and Not Very Healthy levels. We show in Fig. 2 the temporal evolution of the AQI.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Feature Construction</title>
      <p>Our aim is to use a supervised machine learning approach to predict pollutants from tra c and weather data
under the hypothesis that we do not have air sensors to directly measure air pollutants. We show the empirical
cumulative distribution function (ecdf) of the considered weather data in Fig.3 and the distribution of the
tra c features in Fig.4. For sake of readability, the ecdfs related to the air pollutants are shown in Fig.7, in the
Appendix.</p>
      <p>The aforementioned data categories, namely weather, tra c and air pollutants, are merged at hourly resolution,
according to the maximal temporal resolution of both weather and air data. Then, measurements produced by
di erent sensors of the same type in the same hour are averaged.</p>
      <p>Finally, the hourly passages at vehicle gates are counted, according to four di erent sets: (i) the EURO class
(EURO), (ii) the vehicle type (Vtype), (iii) the fuel type (Ftype), (iv) existence of Diesel Particulate Filter
(DPF).</p>
      <p>In order to perform all data manipulations, we use R and the plyr library, which provides data aggregation
operators. As a nal step, we lled a small percentage of missing values (&lt; 1%) in the air and weather data
by polynomial interpolation using the spline function of the zoo library. Conversely, because of the remarkable
percentage of NA values for each group in the tra c features and due to not available information, we lled
missing values following the probability distribution of data.</p>
      <p>
        The nal feature set is composed by the following variables:
{ Time: day of week (
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7">1-7</xref>
        ), hour (
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1-24</xref>
        ). This is to consider the regular patterns of human activities, which are
framed within the day and within the week;
{ Hourly passages: counts of total passages and aggregated count by EURO class, vehicle type, fuel type,
existence of particulate lter. Because we compute the total passages, we remove one category from each
aggregation to avoid creating features which are linear combination of other ones while reducing the number
of total features;
{ Hourly weather phenomena averages: wind direction, wind speed, temperature, relative humidity,
precipitation, atmospheric pressure.
      </p>
      <p>In order to consider the e ect of the past on the current pollutants levels, for each tra c and weather feature
f (t) (wind direction excluded) we add another feature f p(t) equal to the sum of f (t) over the last x time slots
(f p(t) = Ptt==ii x f (t)). We evaluated increasing values of x starting from 1 and and we empirically found the
best value to be 12.</p>
      <p>Studying the correlation between weather and pollutants (Fig. 5) we noticed that temperature, relative humidity,
precipitation wind speed, and atmospheric pressure are the most signi cant ones, having an average absolute
correlation of 0.40, 0.20, 0.13, 0.51, 0.51 with the pollutants considered in the AQI computation, respectively.
Conversely, all tra c features results less correlated and they are not shown for brevity.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Pollutants prediction and and evaluation of AQI</title>
      <p>We implemented several machine learning models using the caret package. In particular, we tested algorithms for
regression, including Generalized Linear Model (GLM), Random Forest (RF), Support Vector Machines (SVM)
and Arti cial Neural Networks (ANN). We tested all algorithms with default hyper-parameters and with the same
random seed, uniformly selecting in time the 70% of the samples as training set, and leaving the remaining 30%
for the test set. We trained all models with a 5-fold cross-validation and we computed the model performances in
terms of mean squared error (MSE) for pollutants. For brevity, we report only the results of the GLM, which we
considered as the baseline and of the ANN, which is the model that performed best. Arti cial Neural Network
(ANN) [8] are inspired by biological nervous systems such as the human brain, which process information through
a large number of highly interconnected processing elements (neurons). ANNs can be used in several applications,
such as pattern recognition or data classi cation, and they are a supervised machine learning technique.
Speci cally, we chose a particular type of ANN, namely the BRNN model (Bayesian Regularization of Neural
Networks) [7], because it is more robust than standard back-propagation networks and it can reduce the need for
lengthy cross-validation. Bayesian regularization is a mathematical process that converts a nonlinear regression
into a \well-posed" statistical problem in the manner of a ridge regression [7]. The main model parameter is the
number of neurons n to be used. In order to de ne the optimal n, we incrementally evaluated the model accuracy
starting with n = 1 and incrementing it in steps of 1 until 20. Therefore, we empirically nd the best value of
n = 9, after which the performance improvement can be considered negligible.</p>
      <p>For each pollutant, we compare the BRNN model with the GLM performances, obtaining for the BRNN an
improvement of the average relative error between 36% and 61% over the GLM. The performance comparison is
fully reported in Table 1.</p>
      <p>Using the predicted values of N O2, O3 and P M10, we computed the AQI value, and then we mapped it into the
classes described in Section 2 (i.e. Optimal, Good, Fair, Average, Not Very Healthy, Unhealthy, Very Unhealthy).
Finally, we computed the accuracy of the estimated AQI class, which is reported in Table 2.</p>
      <p>Our model predicts AQI with a class accuracy of 0.8, which we evaluate as satisfactory (see also the scatter plot
in Fig.6), especially considering that the distance of the classi cation error is never greater than one, meaning
that when the model predicts an erroneous class it is never beyond the adjacent one, e.g., the model can predict
Fair instead of Good but it never predicts Average or any worst condition instead of Good.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future works</title>
      <p>In this paper we used tra c and weather data in order to predict the air pollution in a metropolitan city. We
designed and implemented a set of machine learning models to predict single pollutants that we used to compute
qualitative air quality classes based on a standardized Air Quality Index (AQI). The performance of our best
model, (BRNN with 9 neurons), achieves an AQI class accuracy of 0.8.</p>
      <p>Future works will include the evaluation of our approach on a bigger dataset, an improvement of the feature
set, and the evaluation of several scenarios (e.g., including partial or complete tra c block, di erent weather
conditions, etc.) in order to evaluate the impact of local tra c policies on the air quality.</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix: Pollution agents Ecdf Graphs</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Amos</surname>
          </string-name>
          , \
          <article-title>Polluted air cause 5.5 million deaths a year new research says"</article-title>
          ,
          <source>BBC NEWS, Science and Environment</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,H. Hsieh,\
          <article-title>U-air: When Urban Air Quality Inference Meets Big Data"</article-title>
          , Microsoft Research Asia, ACM,
          <year>2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>I.</given-names>
            <surname>Sundvor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. Castell</given-names>
            <surname>Balaguer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Viana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Querol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Amato</surname>
          </string-name>
          , G. Mellios, C. Guerreiro,\
          <article-title>Road tra cs contribution to air quality in European cities"</article-title>
          ,
          <source>ETC/ACM</source>
          , 2012
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>I.</given-names>
            <surname>Juhosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Makrab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ttha</surname>
          </string-name>
          , \
          <article-title>Forecasting of tra c origin NO and NO2 concentrations by Support Vector Machines and neural networks using Principal Component Analysis"</article-title>
          ,
          <source>Simulation Modelling Practice and Theory</source>
          , vol.
          <volume>16</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>1488</fpage>
          -
          <lpage>1502</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>R.</given-names>
            <surname>Berkowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Palmgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hertel</surname>
          </string-name>
          , E. Vignati, \
          <article-title>A Study on E ects of Weather, Vehicular Tra c and Other Sources of Particulate Air Pollution on the City of Delhi for the Year 2015"</article-title>
          ,
          <source>Journal of Environment Pollution and Human Health</source>
          , vol.
          <volume>4</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>R.</given-names>
            <surname>Gopalaswami</surname>
          </string-name>
          ,\
          <article-title>Using measurements of air pollution in streets for evaluation of urban air quality meterological analysis and model calculations"</article-title>
          ,
          <source>Science of The Total Environment</source>
          , vol.
          <volume>189</volume>
          -
          <issue>190</issue>
          , pp.
          <fpage>259</fpage>
          -
          <lpage>265</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>F.</given-names>
            <surname>Burden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Winkler</surname>
          </string-name>
          ,\
          <article-title>Bayesian regularitazion of neural networks"</article-title>
          ,
          <source>PubMed</source>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>C.</given-names>
            <surname>Stergiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Siganos</surname>
          </string-name>
          ,\
          <article-title>Neural networks"</article-title>
          , Imperial College London,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>