<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ioannina, Greece
$ mohammad.abboud@uvsq.fr (M. Abboud);
karine.zeitouni@uvsq.fr (K. Zeitouni); yehia.taher@uvsq.fr
(Y. Taher)
 https://pages.david.uvsq.fr/kzeitouni/ (K. Zeitouni)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Enriching fixed stations air pollution monitoring with opportunistic mobile monitoring</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohammad Abboud</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karine Zeitouni</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yehia Taher</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DAVID Lab</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>UVSQ - Université Paris-Saclay</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Versailles</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The deteriorating air quality in urban areas, particularly in developing countries, has led to increased attention being paid to the issue. Daily reports of air pollution are essential to efectively manage public health risks. Pollution estimation has become crucial to expanding spatial and temporal coverage and estimating pollution levels at diferent locations. The emergence of low-cost sensors has enabled high-resolution data collection, either in fixed or mobile settings, and various approaches have been proposed to estimate air pollution using this technology. This study aims to enhance the data from fixed stations by incorporating opportunistic mobile participatory monitoring (MPM) data. The research question is: "How can we enrich fixed station data using MPM?" To overcome the limited availability of MPM data, we reuse existing data for periods with similar pollution maps observed by the fixed stations. The combined fixed and mobile data is then subjected to interpolation methods to generate more accurate pollution maps. The efectiveness of our approach is demonstrated by experiments conducted on a real-life dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Air Quality Monitoring</kwd>
        <kwd>Opportunistic Mobile Participatory Monitoring</kwd>
        <kwd>Low-cost sensors</kwd>
        <kwd>Data integration</kwd>
        <kwd>Spatial interpolation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1https://gogreenroutes.eu/</title>
        <p>ing a range of challenges, including monitoring and esti- our methodology. Section 4 presents the
implementamating air pollution. The current research contributes to tion and experimental results. In sections 5 and 6, we
this project by utilizing fixed and mobile sensor data to summarize our findings and suggest future directions for
broaden air pollution estimates’ geographical and tem- research.
poral coverage.</p>
        <p>Researchers have utilized fixed stations and mobile
sensor data to estimate pollution maps. Some studies have 2. Related Work
relied exclusively on fixed stations [ 4, 5, 6], while
others have applied air pollution estimation methods used Researchers have shown interest in the problem of
estiin fixed stations to low-cost mobile sensor data [ 7, 8]. mating pollution for several years. The problem has been
However, recent research proposes combining data from examined in the literature from various perspectives and
ifxed and mobile sensors [ 9, 10], which raises several scales. While meso-scale air quality modeling systems,
unresolved questions. Firstly, what are the most efec- such as CHIMERE[12], are the most commonly used,
urtive methods to use deterministic methods, geostatistical ban scale models utilizing Computational Fluid Dynamic
methods, or machine/deep learning models? Secondly, (CFD) simulations have also been proposed. However,
what features should be considered during the pollution their computational complexity limits their applicability
estimation process? Lastly, how should we address the to a wide area [13]. In addition to these model-driven
challenges of merging data from fixed and mobile sen- approaches, data-driven methods have become popular
sors, considering the diferences in their resolution and due to the increased use of monitoring stations, including
spatiotemporal coverage? traditional fixed networks, denser networks of low-cost</p>
        <p>This paper presents a novel approach to assessing air ifxed sensors, and low-cost mobile devices. In this
dispollution levels in the city of Versailles by utilizing data cussion, we will focus on data-driven approaches that
from both fixed and mobile sensors. Prior studies that expand spatial and temporal coverage. This section
sumintegrate fixed and mobile sensor data or solely rely on marizes the conducted studies on pollution estimation
mobile sensing typically involve targeted campaigns fo- and interpolation for various measurements.
cused on specific routes or deploying sensors on buses or Over the years, numerous techniques have been
sugtrams following fixed paths. In contrast, our methodology gested for approximating or interpolating pollution
levleverages a mobile crowd-sensing (MCS) approach. MCS els in areas without monitoring stations. Although air
[11], as a new paradigm, harnesses data acquired by vol- quality estimation methods are typically intended for
unteers using sensor-enhanced mobile devices with GPS stationary sites, they can also be modified to
accommocapabilities while carrying out their daily routines, result- date information obtained from mobile and stationary
ing in non-persistent data collection and limited outdoor sensors. These techniques can be divided into five
catedata samples as most activities are indoors. Our research gories: Land Use Regression (LUR), Dispersion Models,
question centers on the incorporation of MCS/MPM2 Deterministic Interpolation Methods, Geostatistics, and
data to supplement fixed station data and estimate air ML/DL Algorithms.
pollution levels across the city. In their study cited as [4], the authors employed a</p>
        <p>This work proposes a methodology to augment air deep learning method for predicting the concentration
pollution monitoring stations with randomly collected of PM2.5 in Beijing, China. Their approach involves
data from mobile sensing devices (MCS). Our goal is using a CNN-LSTM neural network to increase the
spato improve the accuracy of pollution maps by utilizing tiotemporal coverage by incorporating historical
polludata from both fixed and mobile sensors, thus increasing tant data, meteorological data, and PM2.5 concentrations
spatial and temporal coverage. Our approach is based from nearby monitoring stations. The proposed approach
on the assumption that similar data can be found for can capture the spatiotemporal characteristics by
comifxed stations at diferent periods. Specifically, we aim bining the convolutional neural network and
long-shortto identify clusters of diferent fixed station data, match term memory network. The study evaluated the proposed
them with MCS data at corresponding times, and combine approach against other deep learning methods. Notably,
them to generate more data samples and improve the this paper focused on predicting future PM2.5
concenpollution map. This method results in enhanced pollution trations rather than estimating or interpolating missing
estimation. values using only fixed monitoring stations and other</p>
        <p>The remainder of this paper is organized as follows: ifxed features in the model.</p>
        <p>Section 2 reviews related work and diferent approaches In [5], the use of LUR methods by Habermann et al.
discussed in the literature. Section 3 details and explains to visualize NO2 pollution concentration distribution is
discussed. LUR is employed due to its reliance on air
pollutant concentration trends. The authors built a LUR
model based on land use, demographic, and geographical
2Please note that MPM (opportunistic mobile participatory
monitoring) and MCS (opportunistic mobile crowd sensing) are used
interchangeably in this paper
features with NO2 measurements as the dependent vari- tator. The approach utilized mobile sensor data without
able. Kriging was then used to visualize the LUR-NO2 relying on additional features and incorporated the
Consurface for each point. The model predicted almost 60% vLSTM structure within the decoder based on a previous
of NO2 variability, although the authors note limitations study [8].
of LUR methods in their paper. In [9], the authors introduced HazeEst, a machine</p>
        <p>A Multi-AP learning network was introduced in [6] learning-based approach that combines sparse fixed
stafor estimating pixel-wise pollution based on fixed-station tions with dense mobile sensor data to estimate hourly air
measures and features such as land use, trafic, and me- pollution surfaces. The method utilized air pollution,
temteorology. The authors classified features into micro, poral, and spatial features and merged fixed and mobile
meso, and macro views and used a fully convolutional data by averaging mobile sensor measurements hourly.
network (FCN) to simulate multiple pollutants. The Multi- The approach implemented several regression methods,
AP network outperformed other methods in various ex- such as SVR, DTR, and RFR.
periments, although the authors acknowledge data con- Song et al. proposed the Deep-Maps approach [10] to
straints, seasonality, and model extension as potential estimate PM2.5 measures. The method combined mobile
challenges. sensor data with fixed stations’ data to expand spatial</p>
        <p>Guo et al. proposed a high-resolution air quality map- coverage and utilized a machine learning framework that
ping approach for multiple pollutants in [7]. The method adapts gradient-boosting decision trees with local
feauses a dense monitoring network and combines dense tures such as land use and meteorological data.
Neighbornetworks and machine learning techniques. The authors ing features captured spatiotemporal correlations among
took advantage of micro-station monitoring systems with urban features, while macro features represented
pollumultiple sensors, as well as land use and meteorological tion measurements from sites outside the study area.
data. XGBoost algorithm was used to estimate pollu- In [18], Zhang et al. proposed machine learning
retion concentration at diferent grids with fine granularity. gression models to predict real-time localized air
qualHowever, the monitoring phase relied on dense network ity, utilizing multiple static and IoT mobile sensors of
data collection. the same type to monitor air quality efectively. The</p>
        <p>The paper by Cassard et al. [14] introduces an engine approach developed gradient boosting, SVR, and RFR
rethat predicts air quality for PM2.5 and PM10 concentra- gression models to estimate pollution, where the gradient
tions in the United States. The authors employed fixed boosting model was most responsive to sudden changes.
and low-cost sensors near road networks and used traf- The results indicated that the hybrid network had better
ifc data to build features. They utilized the five nearest outcomes for all selected dates.
oficial monitoring stations, the five closest low-cost sen- Existing approaches in the literature that use fixed
sors, as well as road and trafic features. A convolutional and/or mobile data have typically conducted targeted
layer was tailored for low-cost sensors, and all features data collection campaigns on specific roads or outdoor
were combined and flattened before being passed through places. However, this work aims to use MCS data to
a fully connected layer. The authors considered three enhance fixed stations’ data without relying on directed
prediction models, including using only oficial stations, data collection campaigns or outdoor data collection.
only low-cost sensors, or a combination of both. While
integrating high-quality data from oficial monitoring
stations with low-cost sensors can improve pollution 3. Methodology
estimation, the authors acknowledge that more spatial
coverage is a potential limitation. In this section, we will present our proposed
method</p>
        <p>In [15], the authors utilized geostatic methods with ology for enhancing fixed station measures with data
data collected from low-cost mobile sensors deployed obtained through mobile crowd sensing. We may have
on top of trams (OpenSense [16]). The study compared very few samples from various outdoor locations when
kriging and deterministic methods such as IDW, where using mobile crowd sensing. Our proposed solution aims
kriging approaches (simple kriging, ordinary kriging, and to address the question of how to leverage MCS data to
kriging with external drift) were found to be superior. improve fixed station measures and estimate air
polluAlthough geostatistical methods do not require external tion.
data, machine learning methods that combine diferent Air pollution levels can vary significantly from one
data types have demonstrated better performance for place to another and may change rapidly due to various
pollution estimation. factors such as meteorological conditions, trafic, and</p>
        <p>In [17], the authors proposed a deep autoencoder land use. Despite these diferences, it is possible to group
model to recover spatiotemporal pollution maps by sep- these changes into clusters that reflect pollution levels
arating the processes of pollution generation and data during specific time periods.
sampling using an encoder, decoder, and sampling imi- Our methodology is based on the hypothesis that fixed
station measures that fall within the same pollution
cluster could share similar MCS data. To test this hypothesis,
we will cluster fixed station measurements and use the
dates and periods to identify relevant MCS data. We will
then use this data to enrich pollution maps and estimate
pollution levels by combining fixed and MCS data.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Algorithm 1: Pollution estimation using Fixed</title>
        <p>and MCS data
Input: Hourly Fixed Stations data, Hourly average</p>
        <p>MCS</p>
        <p>Output: Enriched pollution estimation map
1 Create diferent snapshots of pollution maps
based on hourly fixed station data.
2 Apply a clustering algorithm to group those
snapshots into clusters.
3 Select the date and periods within each cluster.
4 For each cluster, compute the mean of its
pollution maps, and use it as the representative
map for that cluster.
5 Select hourly average MCS data matching the
periods extracted from each cluster.
6 Enrich each representative map with its</p>
        <p>corresponding MCS data.
7 Apply the interpolation method to estimate
pollution on top of the enriched map.</p>
        <p>Our approach is detailed in Algorithm 1, which
outlines the following steps. First, we generate snapshots
from the fixed station data, considering the data from all
ifxed stations for each timestamp as the current state of
pollution. Next, we apply a clustering algorithm (such
as K-means) to identify all similar snapshots. We then
calculate each cluster’s mean per fixed station,
forming a new map representing the cluster. For
example, if entries 1 and  in Figure 1 are grouped in one
cluster, then the representative vector of this cluster
is ([14.25, 16.6, 5.6, 17.95, 5.3, 3.8, 15.1, 17.3]). These
steps are illustrated in Figure 1.</p>
        <p>For MCS data, we begin by calculating the hourly
average. Then, using each cluster’s date and time periods, we
extract the relevant MCS data. The selected data enriches
the representative map, as shown in Figure 2. Finally, we
apply an interpolation technique to generate an air
pollution estimation map, as demonstrated in Figure 3.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Implementation and</title>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <sec id="sec-3-1">
        <title>This section presents the data and methods utilized in our implementation, followed by a discussion of the experiments and results. Our study was conducted in Versailles, within the geographical boundaries of a specified</title>
        <p>bounding box (2.08170001, 48.79231, 2.1540488, and
48.8283) that covers the Versailles region. The area of
this bounding box is approximately 32 square
kilometers, and we partitioned the map into grids with varying
spatial resolutions.
4.1. Data
Our approach to air quality data collection involves two
types: fixed station measures and MPM data. eLichens 3
has deployed eight fixed stations in Versailles, which
provide the air quality index and measures of particulate</p>
      </sec>
      <sec id="sec-3-2">
        <title>3https://www.elichens.com/</title>
        <p>matter (  1.0,   2.5, and   10), as well as esti- for simplicity. We note that additional features, such as
mates of  2, 3, temperature, and humidity. These trafic and meteorology, can also improve performance.
sensors provide hourly aggregations, resulting in one
representative record per station for each timestamp. 4.3. Experiments</p>
        <p>MPM data is collected through the Polluscope 4
project, which conducted multiple campaigns to col- The experiments were carried out on real-life data
collect mobile sensory data. Participants are recruited to lected in Versailles city. Within the context of the
collect surrounding pollutant concentration and geo- GoGreen Routes project, eLichens deployed eight fixed
location for one week, 24 hours a day while going stations in Versailles. Meanwhile, MPM data were
colabout their daily activities. The sensors in their multi- lected as part of the Polluscope project. We utilized data
sensor boxes collect time-annotated measurements of between October and December from both fixed and
mo  1.0,   10,   2.5,  2,   (), bile sensors. The fixed stations produced roughly 1087
temperature, and relative humidity. To address typical hourly average records.
quality issues with low-cost sensors (outliers, noise, and Concerning the MPM data, the experiments afirmed
missing values), the data is preprocessed, and thoroughly the reduction in data when we limited it to an outdoor
screened [19]. context. At the start, we had 11500 minutes of outdoor</p>
        <p>We combine fixed station measurements with MCS records during the collection period, out of a total of
outdoor data by classifying samples based on previous 642200 records, which comprised only 1.7% of the
colresearch to identify micro-environments and selecting lected data. After filtering only records within the
boundonly outdoor periods [20, 21]. However, this represents ing box, we were left with approximately 2538 outdoor
less than 10% of the data due to the majority of time spent records out of 103062 records, which equates to roughly
indoors. To deal with data scarcity, we propose grouping 2.4% of the data collected in Versailles. We observed a
by similarity after calculating the hourly average of the rapid increase in pollution levels for all fixed stations after
MCS data to align with the fixed station data. November 26. Furthermore, we noticed that fixed sensor
S4 had numerous peaks and missing values, implying
4.2. Methods that the observations were unreliable. Therefore, we
conducted experiments with and without S4 measurements.</p>
        <p>For both experiments, we applied the same procedure.</p>
        <p>Firstly, we loaded data from all available stations,
specifically the PM2.5 dimension. Secondly, we removed all
missing values and kept only records with measurements
from all available stations. Finally, we normalized the
data using min-max normalization.</p>
      </sec>
      <sec id="sec-3-3">
        <title>4https://polluscope.uvsq.fr/</title>
        <p>As illustrated in Section 3, our method for estimating
pollution involves adapting unsupervised learning and
geostatistics techniques. The process consists of three
phases. The first phase clusters similar snapshots from
ifxed stations using PM2.5 measures (but the same
process could apply to any other pollutant) where at each
timestamp , the measurements from the eight fixed
stations (1, 2, ..., 8) are considered as a snapshot. 4.3.1. First Experiment
Thus, the input of the clustering algorithm is a set of
snapshots. The K-means algorithm is applied, and the
Elbow method is used to determine the optimal number
of clusters. Each cluster is represented by an aggregated
record that shows the mean values of the stations in that
cluster.</p>
        <p>Moving on to the second phase, we select all the date
and time values of the diferent snapshots in each cluster.</p>
        <p>We use these values to select samples from MPM data,
which is then hourly averaged and aggregated over cells.</p>
        <p>Using each cluster’s representative map with the selected
MPM data, we generate an enriched map for that cluster.</p>
        <p>In the third phase, we estimate pollution in uncovered
areas using enriched maps. While many approaches, such
as machine and deep learning methods like CNN-LSTM,
ConvLSTM, and auto-encoder, have shown promising
results, we use geostatistics methods such as ordinary
kriging interpolation and deterministic interpolation IDW</p>
      </sec>
      <sec id="sec-3-4">
        <title>Once the data was preprocessed and prepared, we utilized</title>
        <p>K-means clustering to partition it into distinct clusters.</p>
        <p>For each record, there were eight measures associated
with the eight stations in this particular experiment. By
applying the elbow method, we designated  = 10,
forming 10 clusters.</p>
        <p>After grouping the fixed station measures’ records, we
calculated a representative map for each cluster. The
next step is to select all the date and time values in each
cluster, which will be used to query MPM data.</p>
        <p>On the other hand, for MPM data, we first selected the
PM2.5 dimension from the preprocessed data. Then we
split the area of interest into grids where each cell has a
1KM X 1KM granularity. The study area is approximately
32  2; thus, we used nine columns and five rows to
split the area of interest into 32 cells. Unfortunately, the
eight stations fall in one cell, number 29. This, on the
one hand, afects the accuracy as the accurate fixed
stations’ measurements are all in one cell. However, on the
IDW
OK
other hand, it shows the strength of our approach since
the performed MPM enrichment allowed us to estimate
pollution even if we have few fixed station measures.</p>
        <p>We merged the MPM data for each cluster, sharing
the exact date and time values. For clusters 0, 3, and 8, Figure 4: Enriched maps
we did not find any MPM data that was collected at the
same time as those clusters. We kept only clusters 1, 2,
4, 5, 6, 7, and 9. We merged the Mobile data with the
representative map of each cluster (fixed data) to get the IDW
enriched maps having MPM data and fixed station data OK
at resolution 1KM x 1KM x 1h.</p>
        <p>The final step is to interpolate missing values. We
use Inverse distance weighting (IDW) and the Ordinary
kriging approach. For each cluster, we applied the two
methods. For validation, we use leave-one-out
validation, where we try to interpolate the cell’s value having
the fixed stations, as it is considered the ground truth.</p>
        <p>Mean absolute error (MAE) and root mean squared error
(RMSE) are used as metrics for validation.</p>
        <p>We repeated the same experiment while varying the
spatial resolution. We split the area of interest into cells
of 500m X 500m. Now the fixed stations fall within two
cells. We repeated the same procedure and applied the
same approaches. Table 1 reports the results of MAE and
RMSE for the diferent splits.</p>
        <p>1KM X 1KM
MAE RMSE
3.7 4.8
3.4 3.9
500m X 500m
MAE RMSE
2 2.6
3 4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Therefore, applying interpolation is irrelevant for this</title>
        <p>cluster. However, as shown in the figure, the MPM data
has enriched the other clusters. Initially, all the clusters
have located in one 1KM x 1KM cell. The figure shows
the importance of the proposed method and how MPM
values change with the change of fixed station measures.
The light yellow color shows a low pollution level, while
the dark red corresponds to high pollution measures.</p>
        <p>Map plot before and after interpolation is shown in
ifgure 5 We chose those 2 clusters to visualize the impact
of interpolation when we have a low pollution level map
as shown in the top part of figure 5, and a high pollution
level in the bottom part. The interpolation is performed
with 1KM x 1KM resolution. The plots show the
superiority of the kriging method over the IDW method.
4.3.2. Second Experiment</p>
      </sec>
      <sec id="sec-3-6">
        <title>We repeated the whole experiment while removing sen</title>
        <p>sor S4. Now for each record, we have seven measures
corresponding to the seven available fixed stations. Again,
with the help of the elbow method, we set K=8, and we
have 8 clusters.</p>
        <p>Unfortunately, we did not have MPM data at the same
periods in cluster 5, as cluster 5 contains only eight
records. We merged MPM data and fixed station data
for other clusters 0, 1, 2, 3, 4, 6, and 7, and we got the
enriched maps.</p>
        <p>The same procedure and methods as in the previous
experiment were applied. We have a grid split of 1KM X
1KM and another split of 500m X 500m. The results are
reported in table 2.</p>
        <p>The following results correspond to the second
experiment (after removing S4 measures). Figure 4 shows the
enriched clusters for the second experiment. We have
eight clusters if we count cluster 5. As aforementioned,
cluster 5 has only 8 records and no matching MPM data.</p>
        <p>Moreover, figure 6 shows the plots of maps before and
after interpolation at 500m X 500m resolution. Again, from similar projects worldwide [22], [3]. One challenge
clusters 1 and 7 are chosen to show the impact of inter- we face is the distribution of fixed stations, which are
polation on spatial measures with high and low pollution all located in small areas. We hope to distribute them
levels. Based on the plots, kriging interpolation preserves better to expand spatial coverage and include additional
the original measurements and estimates pollution levels features such as meteorological data, trafic data, land
in uncovered spots. use features, and other relevant factors that impact air
pollution. Overall, we believe incorporating deep
learning models will be critical to achieving greater accuracy
in our research.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Conclusion</title>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Future work</title>
      <p>Over the past few years, monitoring air pollution through
ifxed stations and inexpensive and portable sensors has
become a popular topic. Due to the constant concern
over air quality in urban areas, improving the air quality
index has become crucial in dealing with the challenges of
urbanization. Several studies have attempted to estimate
pollution levels using fixed stations, mobile sensors, or a
combination of both, using diferent methodologies and
possibly requiring additional features.</p>
      <p>In this study, we present our approach, which involves
combining fixed station data with mobile participatory
sensing (MPM) data collected by individuals during their
daily activities rather than at specific outdoor locations.</p>
      <p>This type of data collection presents a challenge, as only
10% of the time is spent outdoors, resulting in a scarcity
of MPM data. To address this issue, we augment the
MPM data by clustering periods with similar pollution
maps based on fixed station measurements, which we
hypothesize represent the overall conditions. We aggregate
identical measurements into a single map for each cluster
and use all corresponding periods to select the relevant
MPM data, which we merge with the aggregated map to
enhance the air pollution monitoring data. Finally, we
employ interpolation techniques to estimate pollution
levels in uncovered areas.</p>
      <p>We tested our approach on a real-life dataset and
obtained acceptable results. However, future work can be
done to improve the model’s accuracy and performance
by adopting more advanced interpolation methods and
features.</p>
      <p>The main focus of this study is to enhance air monitoring
ifxed stations by incorporating mobile sensor data
collected from the public. Previous projects have typically
conducted targeted mobile sensing campaigns in specific
areas or along particular paths. In contrast, our study
utilizes opportunistic MPM data to supplement fixed station
data.</p>
      <p>Our initial challenge was to determine how to
integrate MPM data with fixed station data to estimate air
pollution. We formulated a hypothesis that periods of air
pollution where fixed stations’ measurements fall within
the same cluster could share similar MPM data. To test
our hypothesis, we clustered the fixed stations’ data and
merged them with MPM data. Our experiments
conifrmed the validity of our hypothesis, and we believe that
this methodology could improve the accuracy of fixed
station data.</p>
      <p>While we achieved acceptable results using basic
interpolation techniques, we anticipate that more advanced
geostatistics, machine learning, and deep learning tech- Acknowledgments
niques could enhance the performance even further.
Convolutional networks could expand spatial coverage, while This work has been supported by the H2020 EU GO
recurrent neural networks could expand temporal cover- GREEN ROUTES funded under the research and
innoage. vation program H2020- EU.3.5.2 grant agreement No</p>
      <p>For future research, we aim to investigate the use of 869764, and by the French National Research Agency
deep learning for interpolation. To generalize our ap- (ANR) project Polluscope, funded under the grant
agreeproach, we plan to collect more MPM data in the con- ment ANR-15-CE22-0018.
text of the GoGreen Routes project and seek out public
datasets with community-based data collection, e.g.,
openaq.org, aircasting.habitatmap.org, and, if available, data
2397–2423.
[13] X. Jurado, Atmospheric pollutant dispersion
estima[1] Air pollution, world health organization [on- tion at the scale of the neighborhood using sensors,
line]. available:https://www.who.int/health-topics/ numerical and deep learning models, Ph.D. thesis,
air-pollution (2023). Université de Strasbourg, 2021.
[2] P. Kumar, L. Morawska, C. Martani, G. Biskos, [14] T. Cassard, G. Jauvion, D. Lissmyr, High-resolution
M. Neophytou, S. Di Sabatino, M. Bell, L. Norford, air quality prediction using low-cost sensors, arXiv
R. Britter, The rise of low-cost sensing for managing preprint arXiv:2006.12092 (2020).
air pollution in cities, Environment international [15] Y. M. Idir, O. Orfila, V. Judalet, B. Sagot, P. Chatellier,
75 (2015) 199–205. Mapping urban air quality from mobile sensors
us[3] C. C. Lim, H. Kim, M. R. Vilcassim, G. D. Thurston, ing spatio-temporal geostatistics, Sensors 21 (2021)
T. Gordon, L.-C. Chen, K. Lee, M. Heimbinder, S.-Y. 4717.</p>
      <p>Kim, Mapping urban air quality using mobile sam- [16] K. Aberer, S. Sathe, D. Chakraborty, A. Martinoli,
pling with low-cost sensors and machine learning G. Barrenetxea, B. Faltings, L. Thiele, Opensense:
in seoul, south korea, Environment international open community driven sensing of environment, in:
131 (2019) 105022. Proceedings of the ACM SIGSPATIAL International
[4] A. Bekkar, B. Hssina, S. Douzi, K. Douzi, Air- Workshop on GeoStreaming, 2010, pp. 39–42.
pollution prediction in smart city, deep learning [17] R. Ma, N. Liu, X. Xu, Y. Wang, H. Y. Noh, P. Zhang,
approach, Journal of big Data 8 (2021) 1–21. L. Zhang, A deep autoencoder model for pollution
[5] M. Habermann, M. Billger, M. Haeger-Eugensson, map recovery with mobile sensing networks, in:
Land use regression as method to model air pollu- Adjunct Proceedings of the 2019 ACM International
tion. previous results for gothenburg/sweden, Pro- Joint Conference on Pervasive and Ubiquitous
Comcedia Engineering 115 (2015) 21–28. puting and Proceedings of the 2019 ACM
Interna[6] J. Song, M. E. Stettler, A novel multi-pollutant tional Symposium on Wearable Computers, 2019,
space-time learning network for air pollution infer- pp. 577–583.
ence, Science of The Total Environment 811 (2022) [18] D. Zhang, S. S. Woo, Real time localized air
qual152254. ity monitoring and prediction through mobile and
[7] R. Guo, Y. Qi, B. Zhao, Z. Pei, F. Wen, S. Wu, ifxed iot sensing network, IEEE Access 8 (2020)
Q. Zhang, High-resolution urban air quality map- 89584–89594.
ping for multiple pollutants based on dense mon- [19] B. Languille, V. Gros, N. Bonnaire, C. Pommier,
itoring data and machine learning, International C. Honoré, C. Debert, L. Gauvin, S. Srairi, I.
Annesijournal of environmental research and public health Maesano, B. Chaix, et al., A methodology for the
19 (2022) 8005. characterization of portable sensors for air quality
[8] R. Ma, X. Xu, H. Y. Noh, P. Zhang, L. Zhang, Gener- measure with the goal of deployment in citizen
sciative model based fine-grained air pollution infer- ence, Science of the Total Environment 708 (2020)
ence for mobile sensing systems, in: Proceedings of 134698.
the 16th ACM Conference on Embedded Networked [20] M. Abboud, H. El Hafyani, J. Zuo, K. Zeitouni,
Sensor Systems, 2018, pp. 426–427. Y. Taher, Micro-environment recognition in the
con[9] K. Hu, A. Rahman, H. Bhrugubanda, V. Sivaraman, text of environmental crowdsensing, in: Workshops
Hazeest: Machine learning based metropolitan air of the EDBT/ICDT Joint Conference,
EDBT/ICDTpollution estimation from fixed and mobile sensors, WS, 2021.</p>
      <p>IEEE Sensors Journal 17 (2017) 3517–3525. [21] H. El Hafyani, M. Abboud, J. Zuo, K. Zeitouni,
[10] J. Song, K. Han, M. E. Stettler, Deep-maps: Machine- Y. Taher, B. Chaix, L. Wang, Learning the
microlearning-based mobile air pollution sensing, IEEE environment from rich trajectories in the context
Internet of Things Journal 8 (2020) 7649–7660. of mobile crowd sensing, GeoInformatica (2022)
[11] B. Guo, Z. Wang, Z. Yu, Y. Wang, N. Y. Yen, R. Huang, 1–44.</p>
      <p>X. Zhou, Mobile crowd sensing and computing: [22] E. Bales, N. Nikzad, N. Quick, C. Ziftci, K. Patrick,
The review of an emerging human-powered sens- W. G. Griswold, Personal pollution monitoring:
ing paradigm, ACM computing surveys (CSUR) 48 mobile real-time air quality in daily life, Personal
(2015) 1–31. and Ubiquitous Computing 23 (2019) 309–328.
[12] S. Mailler, L. Menut, D. Khvorostyanov, M. Valari,</p>
      <p>F. Couvidat, G. Siour, S. Turquety, R. Briant, P.
Tuccella, B. Bessagnet, et al., Chimere-2017: From
urban to hemispheric chemistry-transport
modeling, Geoscientific Model Development 10 (2017)</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>