=Paper=
{{Paper
|id=Vol-2882/MediaEval_20_paper_39
|storemode=property
|title=Personal
Air Quality Index Prediction Using Inverse Distance Weighting Method
|pdfUrl=https://ceur-ws.org/Vol-2882/paper39.pdf
|volume=Vol-2882
|authors=Trung-Quan Nguyen,Dang-Hieu Nguyen,Loc Tai Tan
Nguyen
|dblpUrl=https://dblp.org/rec/conf/mediaeval/NguyenNN20
}}
==Personal
Air Quality Index Prediction Using Inverse Distance Weighting Method==
Personal Air Quality Index Prediction
Using Inverse Distance Weighting Method
Trung-Quan Nguyen1 , Dang-Hieu Nguyen2 , Loc Tai Tan Nguyen3
1, 2, 3 University of Information Technology, Ho Chi Minh City, Vietnam
1, 2, 3 Vietnam National University, Ho Chi Minh City, Vietnam
quannt.13@grad.uit.edu.vn,hieund.12@grad.uit.edu.vn,locntt.12@grad.uit.edu.vn
ABSTRACT
1
In this paper, we propose a method to predict the personal air π€π = (1)
quality index in an area by only using the levels of the following π (π₯, π₯π ) π
pollutants: PM2.5, NO2, O3. All of them are measured from the with π is the power value that is used to control the value of the
nearby weather stations of that area. Our approach uses one of weight. It should be noticed that the Haversine method is used to
the most well-known interpolation methods in spatial analysis, calculate the distance between the two coordinates.
the Inverse Distance Weighted (IDW) technique, to estimate the The value π¦ of an unknown point π₯ is calculated as:
missing air pollutant levels. After that, we can use those levels to Γπ
calculate the Air Quality Index (AQI). The results show that the π=1 π€π ππ
π¦ (π₯) = Γ π π€ (2)
proposed method is suitable for the prediction of those air pollutant π=1 π
levels. with π€π is the weight, ππ is the value of the known point ππ‘β .
3.1 Prediction
1 INTRODUCTION
At first, all possible time frame in hour-interval is listed by grouping
The need to know the personal air pollution data is vital because it the training data. Then, we start to loop through the training data
is better to provide each individual with regional air quality data, per time frame.
which seems to be more accurate than the global data measured In each loop, we get the coordinates of all unknown points that
from far away weather stations. The problem is finding a suitable need to be predicted. After that, we get the values of the known
method to predict air quality data in a local area from the global points and their respective coordinates from the public air pollution
data. This paper reports our solution to tackle this challenge. data provided by 26 weather stations surrounding the Tokyo area
To know more about this challenge and the dataset that we will also in that time frame.
use, you can refer to the overview paper of MediaEval 2020 - Insight With all the necessary data gathered, we can use the IDW for-
for Wellbeing: Multimodal personal health lifelog data analysis [1]. mula to make the prediction. Please note that the initial power
value π of the IDW formula is 2.
2 RELATED WORK After repeating those steps for each air pollutant data (PM2.5,
The inverse distance weighting method [4] is used commonly in NO2, O3), we have the final results.
spatial interpolation [3]. This paper will apply the basic form of
IDW without any modification. 3.2 Optimization
To have the best performance, we could find the optimal value
3 APPROACH of power value p by trying different values of π until the IDW
Due to the limited time available for experimenting with algorithms produces acceptable values of SMAPE/RMSE/MAE.
requiring more time to train data, such as neural network-related After evaluating the π-value ranges from 0 to 5, we find that
algorithms, we choose the IDW. Moreover, because there are no the best power values for PM2.5, NO2, and O3 are 1.5, 3.5, and 0,
statistical assumptions involved [2], it is simpler than Kriging or respectively.
other statistical interpolation methods. The way it works is easy to
understand. Based on the assumption that closer points will have 4 RESULTS AND ANALYSIS
similar values than further points, it will use the measured values The evaluation of PM2.5, NO2, O3, and AQI prediction provided by
surrounding the unknown point to predict the value. By giving MediaEval task organizers are shown in Table 1, Table 2, Table 3,
each known point a weight, the predicted value will be the average and Table 4, respectively.
of those points. In general, PM2.5 prediction is acceptable, but there is a big gap
The weight π€π for a known point π is the inverse of the distance in NO2 and O3 prediction results. It is mainly because the IDW
π from that point to the unknown point π₯, which is computed as: formula does not have any offset parameters to compensate for the
big difference between weather stationsβ public weather data and
Copyright 2020 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).
the one carried out by personal equipment used by volunteers. This
MediaEvalβ20, December 14-15 2020, Online could be because of some differences in methods and devices of
those two data providers.
MediaEvalβ20, December 14-15 2020, Online Quan N.T. et al.
Table 1: Evaluation of the PM2.5 prediction data, such as wind direction, wind speed, temperature, to improve
accuracy.
Sensor MAE RMSE SMAPE
REFERENCES
100001 5.190319201 8.732748788 0.45931373
[1] Zhao P. J Nguyen N.T. Nguyen T.B. Dang-Nguyen D. T. Gurrin C. Dao,
100002 3.720370835 5.511739014 0.406428735
M. S. 2020. Overview of MediaEval 2020: Insights for Wellbeing Task
100003 1.619832154 2.095919331 0.133032135 - Multimodal Personal Health Lifelog Data Analysis. In MediaEval
100005 2.874009812 4.055352722 0.35371517 Benchmarking Initiative for Multimedia Evaluation, CEUR Workshop
100006 3.233921439 4.341966928 0.468214919 Proceedings.
100007 1.695290448 1.707219278 0.625317245 [2] Leonardo Ramos Emmendorfer and GraΓ§aliz Pereira Dimuro. 2020. A
200003 6.465190052 9.724716828 0.444137991 Novel Formulation for Inverse Distance Weighting from Weighted
200004 4.815504659 7.436923815 0.400557289 Linear Regression. In Computational Science β ICCS 2020, Valeria V.
Krzhizhanovskaya, GΓ‘bor ZΓ‘vodszky, Michael H. Lees, Jack J. Don-
garra, Peter M. A. Sloot, SΓ©rgio Brissos, and JoΓ£o Teixeira (Eds.).
Table 2: Evaluation of the NO2 prediction Springer International Publishing, Cham, 576β589.
[3] Jin Li and Andrew D. Heap. 2011. A review of comparative studies of
Sensor MAE RMSE SMAPE spatial interpolation methods in environmental sciences: Performance
and impact factors. Ecological Informatics 6, 3 (2011), 228 β 241. https:
100001 30.15104 34.62797 0.729989 //doi.org/10.1016/j.ecoinf.2010.12.003
100002 13.80071 18.2614 0.399087 [4] Donald Shepard. 1968. A Two-Dimensional Interpolation Function for
100003 18.85267 20.40416 1.218212 Irregularly-Spaced Data. In Proceedings of the 1968 23rd ACM National
100005 12.69285 16.3694 0.411915 Conference (ACM β68). Association for Computing Machinery, New
100006 11.92978 14.12164 0.452494 York, NY, USA, 517β524. https://doi.org/10.1145/800186.810616
100007 14.99076 15.85102 0.562354
200003 12.27167 15.1809 0.364154
200004 7.664357 9.571268 0.257642
Table 3: Evaluation of the O3 prediction
Sensor MAE RMSE SMAPE
100001 11.14697072 16.74763774 0.474838877
β 100002 13.71316126 18.17918429 0.595873229
100003 12.15603603 14.13207772 0.554840783
100005 12.91552723 15.99672071 0.53328839
100006 15.72452576 19.40818331 0.728461886
100007 30.3013034 31.07255621 1.600495059
200003 14.62686484 18.79131409 0.490170718
200004 22.0919231 31.69232972 0.58440423
Table 4: Evaluation of the AQI prediction
Sensor MAE RMSE SMAPE
100001 18.21506046 34.20371647 0.496721967
100002 18.10474466 38.8695944 0.49921946
100003 30.32401094 78.4465017 0.311432437
100005 10.79848535 19.6665506 0.389208159
100006 14.29939129 34.48844262 0.44466795
100007 23.5094483 60.19537217 0.521219253
200003 16.31585216 22.42326978 0.4097449
200004 12.93598111 19.188617 0.378573048
5 DISCUSSION AND OUTLOOK
We intend to explore more advanced algorithms in our future work,
such as the advanced form of IDW [4], the combination of IDW
with multiple regression. Also, we plan to utilize more weather