<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Estimating Urban Ultra ne Particle Distributions with Gaussian Process Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jason Jingshi Li</string-name>
          <email>jason.li@anu.edu.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arnaud Jutzeler</string-name>
          <email>R@Locate14</email>
          <email>arnaud.jutzeler@ep</email>
          <email>arnaud.jutzeler@ep .ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Boi Faltings</string-name>
          <email>boi.faltings@ep</email>
          <email>boi.faltings@ep .ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arti cial Intelligence Laboratory, EPFL</institution>
          ,
          <addr-line>Lausanne, 1025</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>College of Engineering and Computer Science., The Australian National University</institution>
          ,
          <addr-line>Canberra, ACT 0200</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <fpage>145</fpage>
      <lpage>153</lpage>
      <abstract>
        <p>Urban air pollution have a direct impact on public health. Ultra ne particles (UFPs) are ubiquitous in urban environments, but their distribution are highly variable. In this paper, we take data from mobile deployments in Zurich collected over one year with over 25 million measurements to build a high-resolution map estimating the UFP distribution. More speci cally, we propose a new approach using a Gaussian Process (GP) to estimate the distribution of UFPs in the city of Zurich. We evaluate the prediction estimations against results derived from standard General Additive Models in Land Use Regression, and show that our method produces a good estimation for mapping the spatial distribution of UFPs in many timescales.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Tra c junctions, industrial installations and urban canyons all contribute to the high spatial and temporal
variability of air pollution in urban areas. Small-scale spatial distribution of ambient air pollution have
traditionally been studied with Land-Use-Regression (LUR) summarised in Hoek et al. (2008). It uses land-use and tra c
characteristics of a particular grid region as explanatory variables to learn to estimate pollution concentrations
under a Generalized Additive Model (GAM). The learnt model is then used to predict pollution levels for all
locations with the available land-use information.</p>
      <p>In this paper, we propose a novel approach of estimating urban ultra ne particle levels across di erent temporal
aggregates from measurements collected from the trams. Similar to standard models in land use regression, it
estimates the pollution levels within di erent grid-cells in the urban environment from a set of land-use features.
Our model is based on constructing a Gaussian Process described in Rasmussen and Williams (2006), with
additional consideration to spatial features in the covariance matrix. Following the practice in previous work
of Hasenfratz et al. (2014), we evaluate the models (GAM, pure land use and mixed spatial-land-use) using
standard random 10-fold cross validation.</p>
      <p>The outline of this paper is as follows: we begin with a summary of the background to the paper: the data, the
traditional models used in land-use regression, and introducing Gaussian Process Regression. We then introduce
a new approach for estimating UFP levels, and evaluate it against the previous approach over a benchmark
dataset.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        The Aggregate Datasets
The data were selected from UFP measurements collected on Zurich trams between April 2012 and March 2013 as
part of the OpenSense project and the sensing methodology is described in Li et al. (2012b) an
        <xref ref-type="bibr" rid="ref6">d Hasenfratz et al.
(2014</xref>
        ). The data were partitioned into 13,200 grid cells of size 100m 100m. The pro le of a typical grid cell,
such as the one containing Centralplatz in Zurich, is shown in Fig 2. We can see that instead of being tted to a
normal distribution (solid line shows the best- t), the measurements ts much better as log-normal distribution
(dotted line). This is consistent with literature on particle count concentrations in urban environments described
in M lgaard et al. (2012). The data were captured and transmitted in real time to a back-end server running
Global Sensor Network (GSN) by Aberer et al. (2006), and removed to a local database to be preprocessed and
aggregated before entered into the model.
      </p>
      <p>
        Several preprocessing steps were used before the data were prepared for the model, including removal of
measurements within the indoor tram depot, measurements with bad GPS data, and measurements with
extraordinary high levels &gt;100'000 particles per cm3. These steps were described in
        <xref ref-type="bibr" rid="ref6">detail in Hasenfratz et al.
(2014</xref>
        ), with the purpose of avoiding bias due to erroneous measurements.
      </p>
      <p>
        We then aggregate the data within the di erent grid cells according to the di erent time windows, such
as yearly, seasonal, monthly, biweekly, weekly, daily and half-daily. This is done to understand the trade-o
between long and short term aggregate data. In order to evaluate and compare our results to previous work, we
followed the convention of selecting only the 200 grid cells with the highest measurements count for the purpose
of modelling and vali
        <xref ref-type="bibr" rid="ref6">dation, as Hasenfratz et al. (2014</xref>
        ) showed that the state-of-the-art models produced the
most reliable predictions when only the top 200 grid cells with the highest measurement count are considered.
2.2
      </p>
      <p>Land Use Regression
In literature, land-use regression models are used to assess intra-urban air pollution distributions, and a
comprehensive review of these techniques can be found in Hoek et al. (2008). They typically combine monitoring
of air pollution at 20-100 locations spread over the study area, and develop a model using predictor variables
obtained through geographic information systems (GIS). The predictor variable generally include some tra c
information, population density, designated land use and features of the landscape such as attitude and slope.
Due to the cost of deployments, studies usually last 1-2 weeks in duration.</p>
      <p>
        For particulate matter such as PM2:5 and PM10 and UFPs, Generalized Additive Models (GAMs) have been
used in land use regression to study their spatial and temporal variability. It typically use the following equation
to model the relationship between the pollution level p and a set of explanatory variables A1; : : : An.
ln(p) = a + s1(A1) + s2(A2) +
+ sn(An) +
(1)
where a is known as the intercept, the error term, and s1 : : : sn are typically smooth regression splines with
an upper limit of 3 on the degree of freedom. In this paper, we use the GAM
        <xref ref-type="bibr" rid="ref6">data from Hasenfratz et al. (2014</xref>
        )
as a benchmark to compare our model predictions.
2.3
      </p>
      <p>
        Gaussian Process Regression
Also known as Kriging, Gaussian process regression (GPR) has been extensively used for decades in Geostatistics
to model various spatial phenomena such as soil concentrations, weather-related events, etc., and in-depth
overviews can be found in Cressie and Cassie (1993) and Rasmussen and Williams (2006). Similar to other
non-parametric approaches, GPR does not require prior structural knowledge about the phenomenon. Indeed,
the idea is precisely that structure is directly inferred from the data. Furthermore, GPR outputs statistical
predictions and thus represents an adequate candidate to model phenomena that are inherently noisy and which
one can only observe through noisy instruments. Recently it has been successfully applied in many machine
learning tasks such as bioinformatics in Chu et al. (2005), sensor calibration in Monroy et al. (2012) and
crowdsourcing Venanzi et al. (2013) It still represents a very active ongoing research area as seen in e.g. Bonilla et al.
(2010), Cao et al. (2013) an
        <xref ref-type="bibr" rid="ref6">d Nguyen and Bonilla (2014</xref>
        ). To allow the reader to have a better understanding of
our models, in the following we will provide a very brief technical overview of Gaussian Process Regression.
      </p>
      <p>A Gaussian Process (GP) is used to model a phenomenon that takes place in a certain input space X Rd.
We formally write f (x) where x 2 X the function that models the phenomenon. The general idea is to assume
that the function f (x) is a speci c realization of a prior Gaussian Process GP, which is the generalization of
a multivariate normal distribution to an in nity of random variables, that is to say a distribution over whole
functions. A GP is fully de ned by its mean function m(x) and its covariance function k(x; x0) (also called
kernel) that are the generalization of the mean vector, respectively the covariance matrix of a multivariate
normal distribution.</p>
      <p>Regression with a GP is typically performed as follows. In general, we can only make from the phenomenon
noisy observations yi = f (xi) + i where the additive noise is also assumed to be Gaussian N (0; n2). By
using the marginalization property of GPs and the additive nature of the noise we know the joint distribution
of the observations y at locations X and the values f at test points X to be:
y
f</p>
      <p>
        N
mm((XX)) ; k(Xk;(XX +;Xn)2I)
k(X; X )
k(X ; X )
Our rst GP model uses only land-use variables as features to generate predictions on the mean UFP
concentration measured by the sensors within the respective grid cells in the timeframe of the speci ed dataset. They
follow from the features use
        <xref ref-type="bibr" rid="ref6">d in Hasenfratz et al. (2014</xref>
        ). The model takes a vector xLU containing the land-use
variables values of a certain 100m 100m grid cell as input. These land-use features were taken from the
following sources:
Then for those test points X the regression consists in computing the predictive distribution p(f jy). Fortunately
by the conditioning property of a joint multivariate Gaussian distribution this expression is tractable and even
admit a closed formula. It results in another multivariate Gaussian distribution. For any single test points
x 2 X the predictive mean and variance are given by:
f (x ) = m(x ) + k(x ; X)(k(X; X) +
n2I) 1(y
      </p>
      <p>m(x))
V[f (x )] = k(x ; x )
k(x ; X)(k(X; X) +
n2I) 1k(X; x )</p>
      <p>The main challenge is to create and choose prior mean and covariance functions that carry adequate
assumptions about the phenomenon. We describe in detail how we derived such functions in the following section.
k(xLU; x0LU) =
f2 exp
(xLU
x0LU) where M = diag(`LU ) 2
(2)
(3a)
(3b)
(4a)
(4b)
Swiss Federal Statistical O ce
Canton of Zurich government</p>
      <p>{ Average daily tra c volume
OpenStreetMaps.org
{ Population density, industry density, building heights, heating type, terrain elevation, terrain slope
{ Main road type, distance to next major road, distance to major tra c signal</p>
      <p>As we wanted to start with no particular a priori structural knowledge, only very simple mean functions
were tried such as the trivial xed 0 function and a constant c. Deriving a suitable covariance function was,
however, a bit more complex. Indeed, to be valid a covariance function must be positive de nite. It is common
practice to start from well-known parametrized families of positive de nite functions and t the parameters (that
in the scope of GPR are called hyperparameters) using the data. All the covariance functions that were tried
are stationary that is to say every points of the space shows the exact same covariance structure with its own
surroundings or more formally we have k(x; x0) = k(x x0). Stationary covariance functions such as squared
exponential, and various avours of the Matern class were tried, each one carrying di erent assumption about
the smoothness of the process. Finally from preliminary tests results, a constant mean function and a squared
exponential covariance function were selected. We note that the chosen covariance function was the one that
carried the strongest smoothness assumptions. The prior GP is thus de ned by:</p>
      <p>m(xLU) = c
The f2 is the magnitude hyperparameter, and the `LU are the length-scale hyperparameters that determine the
relevance of some or other land-use variables. To learn the values of all the hyperparameters = (c; f2; `LU ; n2)
one can either use optimization or sampling techniques. In our case, we used the standard approach that consists
in optimizing the log marginal likelihood:
Even though the explanation of the phenomenon given by the land-use variables may already be quite good, it
is very likely that part of it still elude us because of some contributions to the phenomenon that are badly or
not at all re ected in the variables. To address this matter, we tried to incorporate geographical informations
into the model with the hope that such missed contribution will at least be partly explained locally.</p>
      <p>The problem with parametric models such as GAM is that we cannot easily add geographic informations into
the model in a sensible way. For example if we naively add the longitude and latitude as covariates, we would
be making very strong assumptions rather unrealistic.</p>
      <p>However, with GPR (and this is why it has been extensively used in Geostatistics) it is natural to include
such informations in the reasoning. This is done by including a consideration for geographical distance in the
covariance function. We call our second model a mixed spatial-land-use model, which is a variant of the rst
one in which we added a term in the covariance structure. We also tried di erent isotropic kernels to be this
additional term. From the preliminary experiments the following covariance function was selected:
k( xxLSU ; xx0L0SU ) =
f2LU exp
It is worth noting that it is the exponential function, the less smooth of the considered covariance functions, that
was chosen to be the additional term in function of the geographical distance. The values of the hyperparameters
= (c; f2LU ; f2S ; `S ; `LU ; n2) were once again xed using marginal likelihood maximization.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluations</title>
      <p>
        We implemented our own Java framework to perform GPR. However, the conjugate gradient optimizer, used
to maximize the log marginal likelihood, was taken from the Matlab toolbox GPML v.2 (see Rasmussen and
Nickisch (2010)) and translated in Java. Most of linear algebra operations were carried out using EJML1 library.
The experiments were conducted on a server with 64 AMD Opteron processing cores and 96 GB of RAM. In the
experiments, we compared the following three di erent type of models on the UFP datasets described earlier.
1. GAM A General Additive Mo
        <xref ref-type="bibr" rid="ref6">del from Hasenfratz et al. (2014</xref>
        );
2. GP LU Our land-use only GP model;
3. GP LUXY Our mixed spatial-land-use GP model.
      </p>
      <p>
        From the benchmarking data supplie
        <xref ref-type="bibr" rid="ref6">d by Hasenfratz et al. (2014</xref>
        ), we get 989 datasets comprise of 597
halfdaily, 309 daily, 44 weekly, 23 biweekly, 11 monthly, 4 seasonally and a single yearly aggregated dataset from
measurements taken from Zurich trams between April 2012 and March 2013. For each aforementioned type of
model, we trained yearly to half-daily models to predict mean pollution level within grid cells (in particle count
per cm3). We evaluated the quality of the models predictions using standard randomised 10-fold cross validation.
That is, for each dataset, we randomly partitioned the data into 10 equal parts, and iteratively we used 9 parts
as training set of the model to generate predictions to be compared against the 1 remaining part.
      </p>
      <p>
        Fig. 3 shows the satellite image of the urban area covered in the deployment, the output of the pollution map
for the season of summer in 2012, and the comparison of the prediction against ground truth of the same season
under random 10-fold cross validation. Fig.4 shows the scatter plots of model predictions against ground truth
1http://code.google.com/e cient-java-matric-library/
across all time scales, where all predictions of the same time scale are located on the same plot. It is worthy
to note that similar to the previous model presente
        <xref ref-type="bibr" rid="ref6">d in Hasenfratz et al. (2014</xref>
        ), our models also show little to
no bias, as evident from the fact that across all cases the linear regression lines (in green) are very close to the
optimal 1-to-1 lines. It indicates the absence of systematic model errors.
4.1
      </p>
      <p>RMSE
First we compare the Root Mean Square Error (RMSE) of the predictions derived from the models under
random 10 fold cross validation (Fig. 5). It is a standard metric of predictive power for measuring the accuracy
of prediction models. It is obtained by:</p>
      <p>RM SE =
s</p>
      <p>PiN=1(pi</p>
      <p>N
gi)2
(6)
where pi denotes the ith prediction, gi the ground truth of the ith prediction, and N the total number of
predictions. In Fig. 5, the plot on the left displays the overall mean of the RMSE, while the box-plot on the right
displays the minimum, lower quartile, median, upper quartile and maximum of the average RMSE of the whole
10-fold validation tests on all the datasets of the same time scale. The yearly data came from a single dataset,
thus it is presented as a single value. It shows that as expected, the higher temporal resolution leads to higher
uncertainty in the prediction, the GP models outperforms GAM across all temporal resolutions, and the mixed
spatial-land-use model produced less error than the land-use only model.
3./]trcm 6000
a
p
[
ESM 4000
R
2000</p>
      <p>0
2R
1
0.8
0.6
0.4
0.2
0
F2AC 0.9
1.2
1.1
1
0.8
0.7
0.6</p>
      <p>GP_LU
GP_LUXY
GAM
yearly
seasonal
monthly
biweekly
weekly
daily
halfdaily
yearly
seasonal
monthly
biweekly
weekly
daily
halfdaily
2R
4.2</p>
      <p>
        R2 Score
We then compare the R2 coe cient, also known as the coe cient of determination of the model predictions
(Fig. 6). It indicates how well the observed outcomes are replicated by the model predictions as the proportional
variation of outcomes explained by the model. Its formula is given by:
where pi denotes the ith prediction, gi the ground truth of the ith prediction, and g is the mean of the ground
truth. In Fig. 6 observe that the variance of the R2 score also increases as the the time scale shrinks across all
models. We see that the results of GP models in general have a higher R2 than GAMs, and introducing the
spatial covariance in the GP model also improves the R2 score across all time scales.
Finally, we compare the F AC2 score of the model predictions (Fig. 7). It measures the fraction of data points
that lie inside the factor of two area. It is a robust measure of prediction as it is not overly in uenced by high
and low outliers. It is derived by:
We implemented two schemes based on Gaussian Process for estimating mean UFP concentrations in urban areas
of Zurich, Switzerland. We show that they provide an alternative to GAM approaches in land-use regression,
and there is a general trade o between the length of the time scale and the quality of the model predictions. We
also show that across the timescales the proposed GP models presents an improvement on the current state of
the art. The resulting maps may be useful for application such as assessing population exposure to air pollutants
similar to that of Carroll et al. (1997), uncover areas of high air pollution for persons with allergies, or evaluate
the trustworthiness of measurements contributed by a community of sensors as described in Li et al. (2012a) an
        <xref ref-type="bibr" rid="ref6">d
Faltings et al. (2014</xref>
        ).
      </p>
      <p>Possible future work includes moving away from a grid-based model to make use of urban spatial features
described in Li et al. (2012b), developing models that handles di erent aspects of sensor reliability and
measurement bias; detecting and ltering spurious measurements, and combining meteorological information and real
time data to produce the best real-time estimations for individual exposure analysis and route planning. Our
approach based on Gaussian Process Regression is very general, and it is interesting to see if it can be generalised
to particulate dispersion outside urban environments to applications such as bush- re detection; and whether it
can be applied to estimating other air-borne or water-borne pollutant dispersions.</p>
      <p>Acknowledgements
We thank our collaborators at ETHZ David Hasenfratz and Olga Saukh for supplying the benchmarking data
and model from their previous work. This work is supported by OpenSense project funded by NanoTera.ch, and
the ARC Discovery Project (DP120103758) \Arti cial Intelligence Meets Sensor Networks".
K. Aberer, M. Hauswirth, and A. Salehi. A middleware for fast and exible sensor network deployment. In</p>
      <p>VLDB, 2006.</p>
      <p>K. Aberer, S. Sathe, D. Chakraborty, A. Martinoli, G. Barrenetxea, B. Faltings, and L. Thiele. OpenSense:</p>
      <p>Open community driven sensing of environment. In ACM IWGS, 2010.</p>
      <p>Edwin V Bonilla, Shengbo Guo, and Scott Sanner. Gaussian process preference elicitation. In NIPS, pages
262{270, 2010.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Yanshuai</given-names>
            <surname>Cao</surname>
          </string-name>
          , Marcus A Brubaker, David Fleet,
          <string-name>
            <given-names>and Aaron</given-names>
            <surname>Hertzmann</surname>
          </string-name>
          .
          <article-title>E cient optimization for sparse gaussian process regression</article-title>
          .
          <source>In NIPS</source>
          , pages
          <volume>1097</volume>
          {
          <fpage>1105</fpage>
          ,
          <year>2013</year>
          . URL http://papers.nips.cc/paper/5087-efficient
          <article-title>-optimization-for-sparse-gaussian-process-regression</article-title>
          .
          <source>pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>RJ</given-names>
            <surname>Carroll</surname>
          </string-name>
          , R Chen,
          <source>EI George, TH Li</source>
          , HJ Newton,
          <string-name>
            <given-names>H</given-names>
            <surname>Schmiediche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Ozone exposure and population density in harris county, texas</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          ,
          <volume>92</volume>
          (
          <issue>438</issue>
          ):
          <volume>392</volume>
          {
          <fpage>404</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Wei</surname>
            <given-names>Chu</given-names>
          </string-name>
          , Zoubin Ghahramani, Francesco Falciani, and David L Wild.
          <article-title>Biomarker discovery in microarray gene expression data with gaussian processes</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>21</volume>
          (
          <issue>16</issue>
          ):
          <volume>3385</volume>
          {
          <fpage>3393</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Noel AC</surname>
          </string-name>
          <article-title>Cressie and Noel A Cassie. Statistics for spatial data</article-title>
          , volume
          <volume>900</volume>
          . Wiley New York,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Boi</given-names>
            <surname>Faltings</surname>
          </string-name>
          , Jason Jingshi Li,
          <string-name>
            <given-names>and Radu</given-names>
            <surname>Jurca</surname>
          </string-name>
          .
          <article-title>Incentive mechanisms for community sensing</article-title>
          .
          <source>IEEE Transactions on Computers</source>
          ,
          <volume>63</volume>
          (
          <issue>1</issue>
          ):
          <volume>115</volume>
          {
          <fpage>128</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Hasenfratz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Saukh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Walser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hueglin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fierz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Thiele</surname>
          </string-name>
          .
          <article-title>Pushing the spatio-temporal resolution limit of urban air pollution maps</article-title>
          .
          <source>In Proceedings of the 12th International Conference on Pervasive Computing and Communications (PerCom'14)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Gerard</given-names>
            <surname>Hoek</surname>
          </string-name>
          , Rob Beelen, Kees de Hoogh, Danielle Vienneau, John Gulliver, Paul Fischer, and
          <string-name>
            <given-names>David</given-names>
            <surname>Briggs</surname>
          </string-name>
          .
          <article-title>A review of land-use regression models to assess spatial variation of outdoor air pollution</article-title>
          .
          <source>Atmospheric Environment</source>
          ,
          <volume>42</volume>
          (
          <issue>33</issue>
          ):
          <volume>7561</volume>
          {
          <fpage>7578</fpage>
          ,
          <year>2008</year>
          . ISSN 1352-
          <fpage>2310</fpage>
          . doi: http://dx.doi.org/10.1016/j.atmosenv.
          <year>2008</year>
          .
          <volume>05</volume>
          .057. URL http://www.sciencedirect.com/science/article/pii/S1352231008005748.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Jason</given-names>
            <surname>Jingshi Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Boi</given-names>
            <surname>Faltings</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Radu</given-names>
            <surname>Jurca</surname>
          </string-name>
          .
          <article-title>Incentive schemes for community sensing</article-title>
          .
          <source>In The 3rd International Conference in Computational Sustainability</source>
          ,
          <year>2012a</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Jason</given-names>
            <surname>Jingshi Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Boi</given-names>
            <surname>Faltings</surname>
          </string-name>
          , Olga Saukh, David Hasenfratz,
          <string-name>
            <given-names>and Jan</given-names>
            <surname>Beutel</surname>
          </string-name>
          .
          <article-title>Sensing the air we breathe - the opensense zurich dataset</article-title>
          .
          <source>In Proceedings of the 26th AAAI Conference on Arti cial Intelligence (AAAI12)</source>
          , Toronto, Canada,
          <year>July 2012b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Bjarke</surname>
            <given-names>M lgaard</given-names>
          </string-name>
          , Tareq Hussein, Jukka Corander, and
          <string-name>
            <given-names>Kaarle</given-names>
            <surname>Hmeri</surname>
          </string-name>
          .
          <article-title>Forecasting size-fractionated particle number concentrations in the urban atmosphere</article-title>
          .
          <source>Atmospheric Environment</source>
          ,
          <volume>46</volume>
          (
          <issue>0</issue>
          ):
          <volume>155</volume>
          {
          <fpage>163</fpage>
          ,
          <year>2012</year>
          . ISSN 1352-
          <fpage>2310</fpage>
          . doi: http://dx.doi.org/10.1016/j.atmosenv.
          <year>2011</year>
          .
          <volume>10</volume>
          .004. URL http://www.sciencedirect.com/science/article/pii/S1352231011010491.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>J</given-names>
            <surname>Monroy</surname>
          </string-name>
          , Achim Lilienthal,
          <string-name>
            <given-names>J</given-names>
            <surname>Blanco</surname>
          </string-name>
          , Javier Gonzalez-Jimenez, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Trincavelli</surname>
          </string-name>
          .
          <article-title>Calibration of mox gas sensors in open sampling systems based on gaussian processes</article-title>
          .
          <source>In IEEE Sensors'12</source>
          , pages
          <fpage>1743</fpage>
          {
          <fpage>1746</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Trung</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Nguyen and Edwin V Bonilla</surname>
          </string-name>
          .
          <article-title>Fast allocation of gaussian process experts</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Carl</given-names>
            <surname>Edward</surname>
          </string-name>
          Rasmussen and
          <string-name>
            <given-names>Hannes</given-names>
            <surname>Nickisch</surname>
          </string-name>
          .
          <article-title>Gaussian processes for machine learning (gpml) toolbox</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>11</volume>
          :
          <fpage>3011</fpage>
          {
          <fpage>3015</fpage>
          ,
          <year>December 2010</year>
          . ISSN 1532-
          <fpage>4435</fpage>
          . URL http://dl.acm.org/citation.cfm?id=
          <volume>1756006</volume>
          .
          <fpage>1953029</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Carl</given-names>
            <surname>Edward Rasmussen and Christopher K. I. Williams</surname>
          </string-name>
          .
          <article-title>Gaussian Processes for Machine Learning</article-title>
          . The MIT Press,
          <year>2006</year>
          . ISBN 0-262-18253-X.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Venanzi</surname>
          </string-name>
          , Alex Rogers, and Nicholas R Jennings.
          <article-title>Crowdsourcing spatial phenomena using trust-based heteroskedastic gaussian processes</article-title>
          .
          <source>In First AAAI Conference on Human Computation and Crowdsourcing</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>World-Health-Organization</surname>
          </string-name>
          .
          <article-title>Air quality and health</article-title>
          .
          <source>In Fact Sheet No. 313</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>