<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Air Quality Estimation Using LSTM and An Approach for Data Processing Techniques</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>AISIA Research Lab</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Minh-Anh Ton-Thien</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Two first author have equal contribution</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Science</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>iLotusLand</institution>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper describes our approach for the MediaEval2021 “CrossData Analytics for (transboundary) Haze Prediction” subtask1. The objective of this subtask is to predict PM10 values at diferent locations in multiple countries using data only from each country itself. In addition, we have applied XGBoost to deal with missing PM10 values on the training dataset and Long Short-term Memory (LSTM) [2] models to predict air pollution.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Nowadays, air pollution leads to increasing cases of cardiovascular
and respiratory diseases. It also afects social and economic
activities. By using data from the last several days to predict air pollution
for upcoming days, we can plan appropriate activities to protect
our health.</p>
      <p>
        As given in the task description [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] of the MediaEval2021, subtask
1 provides time-series datasets collected from diferent air quality
and weather stations in Brunei, Singapore, and Thailand. Therefore,
we decided to use LSTM models to predict air pollution of the
next day from weather features and air quality of 10 previous days.
For Brunei, we built and compared diferent variants of the LSTM
model, i.e., the LSTM, Bidirectional LSTM, and Stacked LSTM. On
the other hand, for Singapore and Thailand datasets, because of
the lack of time and many PM10 values that need to be predicted
hourly, we only used Bidirectional LSTM.
      </p>
    </sec>
    <sec id="sec-2">
      <title>OUR APPROACH</title>
    </sec>
    <sec id="sec-3">
      <title>Missing Values Imputation</title>
      <p>
        Each data point in the training dataset represents the information
of a location of one day in the country. We observed many missing
values in the datasets; thus, we decided to employ two diferent
methods to impute the missing values according to the value type.
For data extracted from weather stations (i.e., temperature, rainfall,
humidity, and wind speed), we filled the missing values with the
mean values of its stations. For data related to air quality from
monitoring stations, i.e., PM10, we employed XGBoost [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to impute
the missing values from the weather features.
      </p>
      <p>First, the missing values of weather features on the training
dataset were filled using the first method. Next, we created a new
dataset from the original training dataset by dropping the rows
where PM10 values are missing. Then, the new dataset was used
to build the XGBoost model to predict missing PM10 values on the
original training data from weather features.</p>
      <p>It is worth noticing that all-weather features of Thailand
collected in 2015 are missing. Therefore, to avoid interference when
iflling missing values, we dropped all data points in that year.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Models</title>
      <p>
        Research studies have shown that LSTM is suitable for time series
data [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]; it is good at solving long-term memory problems,
especially predicting n-th samples using many time steps before. Thus,
we applied LSTM models in our study to predict air pollution.
      </p>
      <p>
        2.2.1 LSTM. One of the disadvantages of Recurrent Neural
Network (RNN) is it can not process long sequences; LSTM architecture
is proposed to solve that problem. An LSTM cell, as shown in
Figure 1, has a cell state C that allows the information to flow through
for long-term memory. It also includes three gates: forget
gatedecide what information should be kept or discarded by looking
at the previous state and current input. Input gate decides what
information is essential at the current step and how to add to the
cell state; output gate decides what the output should be.
2.2.2 Bi-LSTM. Bi-LSTM model, which was developed from
Bidirectional Recurrent Network [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], consists of two LSTM layers:
one taking the input in a forward direction, and the other in a
backward direction. The architecture of an unfolded Bi-LSTM, as
depicted in Figure 2, helps the network go through the input at the
same time so that it can recognize the pattern of our data better.
      </p>
      <p>
        2.2.3 Stacked LSTM. Stacked LSTM is a model that includes
multiple LSTM layers. By making the model deeper, it has proved
its efection in sequence data [
        <xref ref-type="bibr" rid="ref5 ref7">5, 7</xref>
        ]. In our study, we used 2-layers
Stacked LSTM architecture, as described in Figure 3.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Models description</title>
      <p>
        For each country, we did not develop multiple models for each
station but combined data from all the stations into one complete
dataset and build models to recognize the pattern form the dataset.
The first 80% of the data is used for training sets, and the last 20%
of the data is used for validation sets. Information of ten previous
days, which includes weather features and air quality, is used to
predict the PM10 values of the upcoming day. The Adam optimizer
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is employed for model training, and the number of epochs used
during training is 100 with a batch size of 512.
      </p>
      <p>For Brunei, we experimented with LSTM, Bidirectional LSTM
(Bi-LSTM), and Stacked LSTM. The organizers provided the PM10
values hourly for Singapore and Thailand, and we only tested all
three LSTM models for three first hours PM10 values (i.e.,  10_1,
 10_2, and  10_3). The results show that Bi-LSTM is slightly
better than LSTM and Stacked LSTM, so we chose to employ the
Bi-LSTM model for each hourly PM10 value for the final predictions.
3</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>We evaluate the proposed models using RMSE metric calculated as
follow:
√︄ Í</p>
      <p>=1 || () − ˆ () ||2
 =

where  is the number of data points,  () is the i-th measurement
and ˆ () is its ground truth.</p>
      <p>The experimental results on validation sets of Brunei, Singapore,
and Thailand datasets are shown in Table 1. In Brunei, Bi-LSTM
achieves a slightly better result than LSTM, with a score of 3.625.
For Singapore and Thailand, the average scores are 5.821 and 10.624.</p>
      <p>In this work, Bi-LSTM initially has shown some promising results
for Brunei and Singapore on the validation sets. However, it did
not perform as expected in the test sets. It might be because our
missing values imputation technique is not good enough and the
models can’t fully recognize the pattern of the datasets.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>Finally, we would like to send our thanks to AISIA Research Lab for
supporting our team; the Organization Board of MediaEval 2021
and the Task Organizer for providing us with an opportunity to
participate in the competition.
Insight for Wellbeing: Cross-Data Analytics for (transboundary)
Haze Prediction</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Tianqi</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>XGBoost: A Scalable Tree Boosting System</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16)</source>
          . ACM, New York, NY, USA,
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          . https://doi.org/10.1145/2939672.2939785
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9</source>
          ,
          <issue>8</issue>
          (
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Asem</given-names>
            <surname>Kasem</surname>
          </string-name>
          ,
          <string-name>
            <surname>Minh-Son</surname>
            <given-names>Dao</given-names>
          </string-name>
          , Efa Nabilla Aziz,
          <string-name>
            <surname>Duc-Tien</surname>
            <given-names>DangNguyen</given-names>
          </string-name>
          , Cathal Gurrin,
          <string-name>
            <surname>Minh-Triet</surname>
            <given-names>Tran</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thanh-Binh Nguyen</surname>
            , and
            <given-names>Wida</given-names>
          </string-name>
          <string-name>
            <surname>Suhaili</surname>
          </string-name>
          .
          <article-title>Overview of Insight for Wellbeing Task at MediaEval 2021: Cross-Data Analytics for Transboundary Haze Prediction</article-title>
          .
          <source>Proc. of the MediaEval 2021 Workshop</source>
          , Online,
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          December
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Diederik</surname>
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Kingma</surname>
            and
            <given-names>Jimmy</given-names>
          </string-name>
          <string-name>
            <surname>Ba</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Adam: A Method for Stochastic Optimization</article-title>
          . (
          <year>2014</year>
          ). http://arxiv.org/abs/1412.6980 cite arxiv:
          <volume>1412</volume>
          .6980Comment:
          <article-title>Published as a conference paper at the 3rd International Conference for Learning Representations</article-title>
          , San Diego,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Senior</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Beaufays</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Long short-term memory recurrent neural network architectures for large scale acoustic modeling</article-title>
          .
          <source>Proceedings of the Annual Conference of the International Speech Communication Association</source>
          , INTERSPEECH (01
          <year>2014</year>
          ),
          <fpage>338</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuster</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.K.</given-names>
            <surname>Paliwal</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>IEEE Transactions on Signal Processing</source>
          <volume>45</volume>
          ,
          <issue>11</issue>
          (
          <year>1997</year>
          ),
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          . https://doi.org/10.1109/78.650093
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , Oriol Vinyals, and
          <string-name>
            <given-names>Quoc</given-names>
            <surname>Le</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Sequence to Sequence Learning with Neural Networks</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          <volume>4</volume>
          (
          <issue>09</issue>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Yi-Ting</surname>
            <given-names>Tsai</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu-Ren Zeng</surname>
          </string-name>
          , and
          <string-name>
            <surname>Yue-Shan Chang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Air Pollution Forecasting Using RNN with LSTM</article-title>
          .
          <source>In 2018 IEEE 16th Intl Conf on Dependable, Autonomic and Secure Computing, 16th Intl Conf on Pervasive Intelligence and Computing, 4th Intl Conf on Big Data Intelligence and Computing and Cyber Science and Technology Congress</source>
          (DASC/PiCom/DataCom/CyberSciTech).
          <fpage>1074</fpage>
          -
          <lpage>1079</lpage>
          . https://doi. org/10.1109/DASC/PiCom/DataCom/CyberSciTec.
          <year>2018</year>
          .00178
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Thanongsak</given-names>
            <surname>Xayasouk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>HwaMin</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giyeol</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Air Pollution Prediction Using Long Short-Term Memory (LSTM) and Deep Autoencoder (DAE) Models</article-title>
          .
          <source>Sustainability</source>
          <volume>12</volume>
          (03
          <year>2020</year>
          ),
          <volume>2570</volume>
          . https://doi.org/10.3390/su12062570
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>