<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">34636</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2023.034636</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Hybrid Deep Learning Approach for PM2.5 Concentration Prediction in Smart Environmental Monitoring</article-title>
<alt-title alt-title-type="left-running-head">A Hybrid Deep Learning Approach for PM2.5 Concentration Prediction in Smart Environmental Monitoring</alt-title>
<alt-title alt-title-type="right-running-head">A Hybrid Deep Learning Approach for PM2.5 Concentration Prediction in Smart Environmental Monitoring</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Vo</surname><given-names>Minh Thanh</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Vo</surname><given-names>Anh H.</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Bui</surname><given-names>Huong</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Le</surname><given-names>Tuong</given-names></name><xref ref-type="aff" rid="aff-4">4</xref>
<xref ref-type="aff" rid="aff-5">5</xref><email>tuong.lecung@vlu.edu.vn</email></contrib>
<aff id="aff-1"><label>1</label><institution>Faculty of Information Technology, HUTECH University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-2"><label>2</label><institution>Natural Language Processing and Knowledge Discovery Laboratory, Faculty of Information Technology, Ton Duc Thang University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-3"><label>3</label><institution>Faculty of Computing Fundamentals, FPT University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-4"><label>4</label><institution>Laboratory for Artificial Intelligence, Institute for Computational Science and Artificial Intelligence, Van Lang University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-5"><label>5</label><institution>Faculty of Information Technology, Van Lang University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Tuong Le. Email: <email>tuong.lecung@vlu.edu.vn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>9</day><month>3</month><year>2023</year></pub-date>
<volume>36</volume>
<issue>3</issue>
<fpage>3029</fpage>
<lpage>3041</lpage>
<history>
<date date-type="received"><day>22</day><month>7</month><year>2022</year></date>
<date date-type="accepted"><day>22</day><month>12</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Vo et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Vo et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_34636.pdf"></self-uri>
<abstract><p>Nowadays, air pollution is a big environmental problem in developing countries. In this problem, particulate matter 2.5 (PM2.5) in the air is an air pollutant. When its concentration in the air is high in developing countries like Vietnam, it will harm everyone&#x2019;s health. Accurate prediction of PM2.5 concentrations can help to make the correct decision in protecting the health of the citizen. This study develops a hybrid deep learning approach named PM25-CBL model for PM2.5 concentration prediction in Ho Chi Minh City, Vietnam. Firstly, this study analyzes the effects of variables on PM2.5 concentrations in Air Quality HCMC dataset. Only variables that affect the results will be selected for PM2.5 concentration prediction. Secondly, an efficient PM25-CBL model that integrates a convolutional neural network (CNN) and Bidirectional Long Short-Term Memory (Bi-LSTM) is developed. This model consists of three following modules: CNN, Bi-LSTM, and Fully connected modules. Finally, this study conducts the experiment to compare the performance of our approach and several state-of-the-art deep learning models for time series prediction such as LSTM, Bi-LSTM, the combination of CNN and LSTM (CNN-LSTM), and ARIMA. The empirical results confirm that PM25-CBL model outperforms other methods for Air Quality HCMC dataset in terms of several metrics including Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE).</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Time series prediction</kwd>
<kwd>PM2.5 concentration prediction</kwd>
<kwd>CNN</kwd>
<kwd>Bi-LSTM network</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Machine learning and deep learning have been developing very rapidly and have been applied in many fields such as economics [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>], engineering [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>], and natural language processing [<xref ref-type="bibr" rid="ref-5">5</xref>]. Recently, there are a lot of research applying deep learning to analyzing and predicting time series. Kocheturov et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] developed a new method for mining complex temporal patterns in multivariate time series. This study proposed the Extended Vertical List structure to track positions of the first state of the pattern inside records. In addition, this structure can link them to appropriate positions of the prefix. This method will be used to extract the features of the multivariate time series. Next, Gu et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] proposed a new active multi-source transfer learning (MultiSrcTL) approach for time series prediction. The results on six benchmark datasets indicate that the applicability and effectiveness of MultiSrcTL algorithm. Then, Le et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] developed a new model, which is a combination of a convolutional neural network (CNN) and bidirectional long short-term memory (Bi-LSTM), for forecasting the electric energy consumption on various variations of the individual household electric power consumption dataset in various time spans including various timespan datasets such as real-time, short-term, medium-term, and long-term datasets. Next, Le et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] developed MEC-TLL framework for the multiple electric energy consumption forecasting utilizing transfer learning and LSTM. This study proposed a cluster-based strategy for transfer learning the LSTM models to reduce the processing time. The results confirmed that MEC-TLL saves processing time while keeping outstanding performance. Vo et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] proposed BOP-BL model for predicting oil prices that change over time. BOP-BL has two following modules: (i) Three Bi-LSTM layers are used in the first module. This module aims to learn useful information features in forward and backward directions. (ii) The second module has a fully connected layer to predict the oil price using extracted features. Zhang et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] developed a deep learning-based framework for time series prediction in finance. The proposed model integrates the advantages of CEEMD, PCA, and LSTM to improve the performance in predicting the stock indices. Therefore, deep learning has been widely used in time series predictions in many different fields. Recently, Vo et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] collects the monthly water consumption with 16,244 households in Can Tho city, Vietnam for three years from 2018 to 2020. Then, a new approach using Bi-LSTM for predicting the monthly household water consumption was developed.</p>
<p>In recent years, there have been many studies on air pollution using machine learning and deep learning [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>] when air pollution has become a severe problem in urban cities. Air pollution occurs in industrial cities or with high construction density. In such contexts, the surrounding air contains many things such as gases, dust, fumes, or odors in high enough quantities. This situation is harmful to the health of humans and animals. One of the biggest killers in this age is air pollution. In 2015, polluted air was the reason for 6.4 million deaths in the world including 2.8 million from household air pollution and 4.2 million from outdoor air pollution [<xref ref-type="bibr" rid="ref-15">15</xref>]. Therefore, it is necessary to develop models to accurately predict air pollution levels to have suitable strategies. Zhang et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] developed a hybrid approach, where multiple static sensors and mobile sensors are utilized to monitor the air quality. In addition, this study also builds a machine learning model which utilizes the collected data for training and provides the predicted information about air quality. Zeinalnezhad et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed two models of semi-experimental nonlinear regression and adaptive neuro-fuzzy inference system (ANFIS). This system aims to predict the concentration of four important pollutants including CO, SO2, O3, and NO2. Sch&#x00FC;rholz et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed a framework that includes context-aware computing concepts to merge LSTM with other information. This model was evaluated by an application in the Melbourne Urban Area (Victoria, Australia) which provides high values of the precision metric. Ma et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] introduced the problem of predicting air quality for the new stations that lack training data. They then developed a transfer learning-based stacked Bi-LSTM model (TLS-BLSTM) to predict air quality in this situation. The results indicate that TLS-BLSTM achieved 35.21&#x0025; lower in terms of RMSE for the experimented three pollutants in new stations. Chang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] utilized an Aggregated LSTM model that combines three kinds of stations: local air quality monitoring stations, industrial areas stations, and external pollution sources stations. This approach was verified by the dataset collected by Taiwan Environmental Protection Agency between 2012 and 2018. This model obtained the best experimental results compared with SVM-based Regression, Gradient Boosted Tree Regression, LSTM, etc. The above forecasting models are applied to predict air pollution in smart cities. The current state of air pollution is having a very serious impact on human life, causing 5&#x0025; of mortality from tracheal, bronchial, and lung cancers, 3&#x0025; of mortality from cardiopulmonary disease, and about 1&#x0025; mortality from acute respiratory infections in children under 5 years of age [<xref ref-type="bibr" rid="ref-21">21</xref>]. The prevention and mitigation of consequences of air pollution are very urgent issues today. The accurate forecast of air pollution helps to warn people when the pollution index exceeds the permissible threshold through public announcements or an automatic messaging system. People will plan to carry out their daily activities in places with less air pollution or limit going out and take necessary precautions. This problem of air pollution forecasting is very important in contributing to the protection of public health.</p>
<p>Among all the particulate matters of air pollution problem, PM2.5 is of particular concern. Therefore, there is also a lot of research on PM2.5 concentration prediction. Feng et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] used Multi-layer Perceptron, Ma et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] utilized XGBoost classifier, Wang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] made use of the artificial neural network, and Pak et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] employed Deep learning for PM2.5 concentration prediction in several cities in China. Xu et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] developed a spatial ensemble-based approach for hourly PM2.5 concentrations prediction around Beijing railway station in China while Sun et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a novel stacking-driven ensemble approach to predict hourly PM2.5 concentration in the winter of the Beijing-Tianjin-Hebei, China. Liu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a novel factory-aware attention mechanism by a LSTM neural network (FAA-LSTM) to forecast the PM2.5 concentrations. Next, Ma et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] developed a Lag-FLSTM (Lag layer-LSTM-Fully Connected network) model based on Bayesian Optimization (BO) for multivariate PM2.5 concentration forecasting at the Wayne County in Michigan, the U.S. Xu et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] developed a temporal-spatial-regression-tree model to predict the distribution of PM2.5. However, the above models cannot be applied to another location which has dataset with many different features. As well as, with the importance of PM2.5 predictions, this study proposes a hybrid deep learning approach for PM2.5 concentration prediction in Ho Chi Minh City, Vietnam. From that result, smart systems in smart city can warn to protect the health of their citizens as a short-term strategy. In addition, managers will have appropriate long-term strategies to protect the environment.</p>
<p>The major contributions of this study are summarized as follows.
<list list-type="bullet">
<list-item><p>This study analyzes the effects of variables on PM2.5 concentrations in the Air Quality HCMC dataset.</p></list-item>
<list-item><p>PM25-CBL model that integrates CNN and Bidirectional Long Short-Term Memory (Bi-LSTM) is developed.</p></list-item>
<list-item><p>This study conducts the experiment to evaluate the predictability of the proposed approach and LSTM, Bi-LSTM, CNN, CNN-LSTM, CNN-Bi-LSTM, and ARIMA models. The results indicate that the PM25-CBL model outperforms other experimental methods for the Air Quality HCMC dataset in terms of MSE, RMSE, MAE, and MAPE metrics.</p></list-item>
</list></p>
<p>The remainder of this article is organized as follows. The detail of the Air Quality HCMC dataset and PM25-CBL model that integrates CNN and Bi-LSTM for PM2.5 concentration prediction, are presented in Section 2. The experiments are conducted in Section 3. Finally, Section 4 gives the conclusion of this study. Several future works are introduced in this section.</p>
</sec>
<sec id="s2"><label>2</label><title>Material and Method</title>
<sec id="s2_1"><label>2.1</label><title>The Air Quality HCMC Dataset</title>
<p>The Vietnam Air Quality Data Series [<xref ref-type="bibr" rid="ref-31">31</xref>] provides daily air quality values for several cities in Vietnam. This study focuses on PM2.5 concentration prediction in Ho Chi Minh City, Vietnam. Therefore, only air quality values in Ho Chi Minh City are selected. Seven variables include the date, temperature, humidity, wind speed, PM2.5 concentrations&#x2005;(&#x03BC;g/m&#x00B3;), dew, and pressure in Air Quality HCMC dataset attached with the descriptions and basic statistics are shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>The variables of the air quality HCMC dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">&#x0023;No</th>
<th align="left">Variable</th>
<th align="left">Description</th>
<th align="left">Min</th>
<th align="left">Max</th>
<th align="left">Median</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">Date</td>
<td align="left">Date time</td>
<td align="left">2020-01-03</td>
<td align="left">2021-01-20</td>
<td align="left">2020-07-20</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Temperature</td>
<td align="left">Median of temperature</td>
<td align="left">23</td>
<td align="left">31</td>
<td align="left">27.5</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Humidity</td>
<td align="left">Median of humidity</td>
<td align="left">47</td>
<td align="left">100</td>
<td align="left">78</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Wind speed</td>
<td align="left">Median of wind speed</td>
<td align="left">0.5</td>
<td align="left">5.4</td>
<td align="left">2.3</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">PM25</td>
<td align="left">Median of PM2.5 concentrations</td>
<td align="left">5</td>
<td align="left">171</td>
<td align="left">65</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Dew</td>
<td align="left">Median of dew</td>
<td align="left">14.5</td>
<td align="left">26.5</td>
<td align="left">24</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">Pressure</td>
<td align="left">Median of pressure</td>
<td align="left">1003</td>
<td align="left">1014</td>
<td align="left">1009</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Before further analysis, this study preprocesses the data first. To avoid the effects of dimensional difference and improve processing time, six variables including temperature, humidity, wind speed, PM2.5 concentrations, dew, and pressure are normalized by the Min-Max Scaler as the following equation:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">scaled</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Then, the correlations between temperature, humidity, wind speed, dew, and pressure with PM2.5 concentrations in Air Quality HCMC dataset are shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The last column of the scatter plot matrix indicates that temperature, dew, humidity, and wind speed have a low negative correlation with PM2.5 concentrations while pressure has a low positive correlation with PM2.5 concentrations. Therefore, this study will use all variables including temperature, dew, humidity, wind speed, and pressure in the proposed model.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>A scatter plot matrix for air quality HCMC dataset</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_34636-fig-1.tif"/></fig>
</sec>
<sec id="s2_2"><label>2.2</label><title>PM25-CBL: A Hybrid Deep Learning Approach for PM2.5 Concentration Prediction</title>
<p>This section presents the overall architecture of a hybrid deep learning approach named the PM25-CBL model for PM2.5 concentration prediction in the Air Quality HCMC dataset. This model shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> is the combination of CNN and Bi-LSTM, the latest deep learning models to solve the time series prediction problem. The proposed model consists of two following phases. In the training phase, six input variables from Air Quality HCMC dataset are put into the CNN module for extracting features. The extracted features will be passed to the Bi-LSTM module. Then, PM25-CBL can predict the PM2.5 concentrations by the fully connected module. In the testing phase, only five features including temperature, humidity, wind speed, dew, and pressure are passed to the trained model to predict PM2.5 concentrations in the testing set.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>PM25-CBL model</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_34636-fig-2.tif"/></fig>
<p><bold>CNN module.</bold> This module utilizes two one-dimensional (1D) convolution layers for analyzing the input variables in Air Quality HCMC dataset. Each 1D convolution layer is followed by an 1D max-pooling layer to reduce the computational complexity. Specifically, this CNN module aids to eliminate the feature redundancies by applying the filters on the input time series data.</p>
<p>For detail, the left of <xref ref-type="fig" rid="fig-3">Fig. 3</xref> is the input time series data. At a data point in that time series, there are six variables. The red and yellow windows in <xref ref-type="fig" rid="fig-3">Fig. 3</xref> represent the filters. In the PM25-CBL model, we opt for the size of filters with 64 and kernel size with 2 to apply for our model. The number of the extracted feature dimensions is N&#x2009;&#x00D7;&#x2009;1 after convolution with the red filter, where N is based on the number of input data dimensions, the size of the filter, and the convolution step length. The yellow window indicates another filter, which can be followed by other filters. Suppose that we have M filters, then the extracted feature dimension will be N&#x2009;&#x00D7;&#x2009;M. Finally, the output of this process will push into the activation function (ReLU), allowing the PM25-CBL model to learn faster and perform better.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>The process of 1D convolution layer for multivariate time series data [<xref ref-type="bibr" rid="ref-32">32</xref>]</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_34636-fig-3.tif"/></fig>
<p><bold>Bi-LSTM module.</bold> To understand Bi-LSTM, this study first briefs the LSTM network as follows. This network, which first introduced by Graves and Schmidhuber [<xref ref-type="bibr" rid="ref-33">33</xref>], aims to improve performances of the traditional recurrent neural networks (RNNs). An LSTM-based model uses a unique set of memory cells instead of the hidden layer neurons in RNNs model. LSTM networks filter information through the gate structure to maintain and update the state of memory cells. The input, forget, and output gates are three main types of gate structure in a memory cell. Each memory cell in an LSTM-based model has two types of nonlinear activation function: <italic>sigmoid</italic> and <italic>tanh</italic> functions. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the diagram for a memory cell at the time step <italic>t</italic>.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>A memory cell at the time step <italic>t</italic></title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_34636-fig-4.tif"/></fig>
<p>The forget gate in a memory cell provides which cell state information will be discarded. In <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the memory cell takes <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as inputs where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the output of the previous step, and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the external information at the current step. The forget gate uses the <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> to combine two above variables <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> into a long vector through <italic>sigmoid</italic> function denoted by <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> as follows.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the weight matrix and bias vector of the forget gate. The forget gate aims to record how much the cell state of the previous step (<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>) is reserved to the cell state of the current step (<inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>). The output of the forget gate is a value in [0, 1]. Where value of 1 indicates the complete reservation and value of 0 shows the full discernment.</p>
<p>The input gate identifies how much of the current moment input <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is used to reserve into the cell state of the current step (<inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>). This gate aims to prevent useless information going to the memory cells. This gate has two following functions. <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> is to find the state of the cell, that must be updated, by the <italic>sigmoid</italic> function while <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> aims to update the information to the cell state. A new candidate vector denoted by <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>^{'}</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> is created through the <italic>tanh</italic> function to control how much new information will be added. Finally, this model uses <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> to update the cell state of the memory cells.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>The output gate determines how much of the current cell state will be discarded. The output information, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mi mathvariant="bold">o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, is first calculated by the <italic>sigmoid</italic> function as follows.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Then the cell state (<inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) is determined by <italic>tanh</italic> function and multiplied by the output information, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mi mathvariant="bold">o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, to obtain the final value. The following equation is used to find the cell state.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>LSTM network only analysis one directional of a sequence which leads to reduce its effectiveness. Meanwhile, both forward and backward directional information on the sequence may contain interesting patterns. Therefore, Bi-LSTM which considers both forward and backward directions in the sequence [<xref ref-type="bibr" rid="ref-34">34</xref>] was introduced. This model&#x2019;s fundamental idea is that it looks at a particular sequence from the front-to-back and back-to-front. In which an LSTM layer is for forwarding processing while the remaining layer is for backward processing. In this way, the network could capture the change of PM2.5 concentrations in both its past and future.</p>
<p><bold>Fully connected module.</bold> Finally, the outputs of Bi-LSTM are fed to the fully connected module which consists of a fully connected layer to generate the PM2.5 concentrations, which helps to reduce the learning parameters in the model instead of following many fully connected layers. Finally, the configurations of the PM25-CBL model are presented in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>The configurations of PM25-CBL model</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">&#x0023;No</th>
<th align="left">Layer type</th>
<th align="left">Neurons</th>
<th align="left">Parameters</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">1D Convolution</td>
<td align="left">(None, None, 5, 64)</td>
<td align="left">192</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">1D Max pooling</td>
<td align="left">(None, None, 5, 64)</td>
<td align="left">0</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">1D Convolution</td>
<td align="left">(None, None, 4, 64)</td>
<td align="left">8256</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">1D Max pooling</td>
<td align="left">(None, None, 4, 64)</td>
<td align="left">0</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Flatten</td>
<td align="left">(None, None, 256)</td>
<td align="left">0</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Bi-LSTM</td>
<td align="left">(None, None, 128)</td>
<td align="left">164,352</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">Dropout</td>
<td align="left">(None, 128)</td>
<td align="left">0</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">Bi-LSTM</td>
<td align="left">(None, 64)</td>
<td align="left">41,216</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">Fully connected layer</td>
<td align="left">(None, 1)</td>
<td align="left">65</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Experiments</title>
<sec id="s3_1"><label>3.1</label><title>Experiment Setting</title>
<p>This section evaluates our approach and several state-of-the-art models for time series prediction including LSTM, Bi-LSTM, CNN, CNN-Bi-LSTM, CNN-LSTM and ARIMA for Air Quality HCMC dataset. The above methods are implemented in the Keras framework and executed in the Ubuntu computer with an Intel Core i7-4790&#x2005;K (4.0&#x2005;GHz&#x2009;&#x00D7;&#x2009;8 cores), 32&#x2005;GB of RAM, and GeForce GTX 1080 Ti. To demonstrate the effectiveness of PM25-CBL model, the Air Quality HCMC dataset is divided into 5-folds with 20&#x0025; for testing set and the remaining of this dataset for training set. This study utilizes four following common performance metrics to compare the experimental methods.</p>
<p>The first metric, MSE, is the average squared difference between the predicted values given by the machine learning model and the actual values. Meanwhile, the second metric namely RMSE is the square root of the MSE. They are determined by the following equations.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:math></disp-formula></p>
<p>Next, MAE gives the average magnitude of the prediction errors and ignores their directions by the absolute operator. Meanwhile, MAPE provides prediction accuracy in the percentage of a forecasting method. The following equations can obtain them.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>100</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>|</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow></mml:mfrac><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_2"><label>3.2</label><title>Experiment Results</title>
<p>Firstly, the loss values during training and testing phases were tracked, which are shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The results confirm that the loss of training and testing phases are stable after 20 epochs. Therefore, our model will be trained in 20 epochs. Besides, this study utilizes batch size at 30 and Adam optimizer for PM25-CBL model. The Adam optimizer has an initial learning rate of 0.001.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>The loss values of PM25-CBL model in training and testing phases</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_34636-fig-5.tif"/></fig>
<p>Secondly, this section reports the performances of the experimental methods, including LSTM, Bi-LSTM, CNN, CNN-LSTM, CNN-Bi-LSTM, ARIMA and the proposed approach for the Air Quality HCMC dataset in terms of MSE, RMSE, MAE, and MAPE. The experimental results in <xref ref-type="table" rid="table-3">Table 3</xref> show that our approach achieves the best values of MSE, RMSE, MAE, and MAPE at 1.23, 1.09, 0.89, and 3.19, respectively. For the MSE metric, the proposed approach improves 12.7&#x0025; compared with CNN-LSTM, improves 6.1&#x0025; compared with Bi-LSTM, improves 16.9&#x0025; compared with LSTM, improves 8.2&#x0025; compared with CNN, improves 18.5&#x0025; compared with CNN-Bi-LSTM, and improves 37.6&#x0025; compared with ARIMA. Likewise, our approach improves from 3.5&#x0025; to 22&#x0025; in the RMSE metric, from 5.3&#x0025; to 12.7&#x0025; in the MAE metric, and from 5&#x0025; to 80&#x0025; in the MAPE metric compared with the remaining experimental methods. Therefore, the proposed method outperforms LSTM, Bi-LSTM, CNN, CNN-LSTM, CNN-Bi-LSTM, and ARIMA models in all performance metrics.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Results of four experimental methods for Air Quality HCMC dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">&#x0023;No</th>
<th align="left">Model</th>
<th align="left">MSE</th>
<th align="left">RMSE</th>
<th align="left">MAE</th>
<th align="left">MAPE</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">LSTM</td>
<td align="left">1.48&#x2009;&#x00B1;&#x2009;0.34</td>
<td align="left">1.21&#x2009;&#x00B1;&#x2009;0.13</td>
<td align="left">1.02&#x2009;&#x00B1;&#x2009;0.09</td>
<td align="left">3.63&#x2009;&#x00B1;&#x2009;0.27</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Bi-LSTM</td>
<td align="left">1.31&#x2009;&#x00B1;&#x2009;0.64</td>
<td align="left">1.13&#x2009;&#x00B1;&#x2009;0.18</td>
<td align="left">0.94&#x2009;&#x00B1;&#x2009;0.22</td>
<td align="left">3.36&#x2009;&#x00B1;&#x2009;0.4</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CNN</td>
<td align="left">1.34&#x2009;&#x00B1;&#x2009;0.36</td>
<td align="left">1.15&#x2009;&#x00B1;&#x2009;0.16</td>
<td align="left">0.94&#x2009;&#x00B1;&#x2009;0.12</td>
<td align="left">3.37&#x2009;&#x00B1;&#x2009;0.4</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CNN-LSTM</td>
<td align="left">1.41&#x2009;&#x00B1;&#x2009;0.4</td>
<td align="left">1.17&#x2009;&#x00B1;&#x2009;0.22</td>
<td align="left">0.97&#x2009;&#x00B1;&#x2009;0.13</td>
<td align="left">3.46&#x2009;&#x00B1;&#x2009;0.4</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">CNN-Bi-LSTM</td>
<td align="left">1.51&#x2009;&#x00B1;&#x2009;0.5</td>
<td align="left">1.2&#x2009;&#x00B1;&#x2009;0.25</td>
<td align="left">1.0&#x2009;&#x00B1;&#x2009;0.21</td>
<td align="left">3.57&#x2009;&#x00B1;&#x2009;0.7</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">PM25-CBL</td>
<td align="left"><bold>1.23</bold>&#x2009;&#x00B1;&#x2009;<bold>0.4</bold></td>
<td align="left"><bold>1.09</bold>&#x2009;&#x00B1;&#x2009;<bold>0.21</bold></td>
<td align="left"><bold>0.89</bold>&#x2009;&#x00B1;&#x2009;<bold>0.14</bold></td>
<td align="left"><bold>3.19</bold>&#x2009;&#x00B1;&#x2009;<bold>0.45</bold></td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">ARIMA</td>
<td align="left">1.97&#x2009;&#x00B1;&#x2009;0.2</td>
<td align="left">1.40&#x2009;&#x00B1;&#x2009;0.07</td>
<td align="left">0.99&#x2009;&#x00B1;&#x2009;0.05</td>
<td align="left">16.19&#x2009;&#x00B1;&#x2009;0.002</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, this study evaluates the processing time of the experimental methods for the Air Quality HCMC dataset. <xref ref-type="table" rid="table-4">Table 4</xref> indicates that CNN is the best method in both training time and predicting time with 1.31 and 0.027&#x2005;s, respectively. CNN-LSTM reached 2nd place with 3.255 and 0.107&#x2005;s for training time and predicting time, respectively. LSTM achieved 3rd place with 3.302 and 0.124&#x2005;s for training time and predicting time. The proposed method is 4th place with insignificant time gaps compared with the two above methods. Specifically, PM25-CBL is 0.149 s slower than the CNN-LSTM method in training time and 0.063 s in testing time, but it is faster CNN-Bi-LSTM with 1.496&#x2005;s for training time. Bi-LSTM is the worst method in terms of processing time, with 5.679 and 0.234 for training time and testing time, respectively.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Processing time of four experimental methods for air quality HCMC dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">&#x0023;No</th>
<th align="left">Model</th>
<th align="left">Training phase (s)</th>
<th align="left">Predicting phase (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">LSTM</td>
<td align="left">3.302</td>
<td align="left">0.124</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Bi-LSTM</td>
<td align="left">5.679</td>
<td align="left">0.234</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CNN</td>
<td align="left"><bold>1.31</bold></td>
<td align="left"><bold>0.027</bold></td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CNN-LSTM</td>
<td align="left">3.255</td>
<td align="left">0.107</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">CNN-Bi-LSTM</td>
<td align="left">4.9</td>
<td align="left">0.17</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">PM25-CBL</td>
<td align="left">3.404</td>
<td align="left">0.17</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">ARIMA</td>
<td align="left">4.39</td>
<td align="left">0.112</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3"><label>3.3</label><title>Discussions</title>
<p>The experimental results on processing time indicate that the proposed method has not achieved the best training and predicting time. It has twice the training time as the best model (CNN). However, training time of the proposed method only takes 3.4 s. This is not significant when hardware has evolved dramatically. In addition, the proposed method achieves the best results in accuracy with MSE, RMSE, MAE, and MAPE metrics. Obviously, with the improvements in terms of performance, the processing time of the proposed method is acceptable compared with CNN, LSTM, Bi-LSTM, CNN-LSTM, and ARIMA models. With the above analysis, the advantage of the proposed model is about predictability while the limitation of the model is the training time. However, the time difference between the methods is not significant., therefore, the proposed method is recommended to be integrated in smart environmental monitoring to automatically provide forecasts for citizens when PM2.5 concentrations reach dangerous thresholds in smart city.</p>
<p>Currently, this study developed a deep learning model for prediction PM2.5 concentration prediction applied in Ho Chi Minh City, Vietnam. Therefore, it is necessary to conduct the study to interact with this model to the system to be able to use it in practice. In addition, the collection of features is still very limited, leading to low accuracy. Therefore, it is necessary to collect more information regarding the environment in the future.</p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Conclusion</title>
<p>This study developed a hybrid deep learning approach named PM25<bold>-</bold>CBL model for PM2.5 concentration prediction in Ho Chi Minh City, Vietnam. This study first analyzed the effects of variables on PM2.5 concentrations in the Air Quality HCMC dataset which helps to choose the best variables. Secondly, the PM25<bold>-</bold>CBL model that integrates CNN and Bi-LSTM is developed. The PM25<bold>-</bold>CBL model consists of three following modules: CNN, Bi-LSTM, and Fully connected modules. Finally, this study conducts the experiment to compare the performance of the proposed approach and three state-of-the-art frameworks for time series prediction including LSTM, Bi-LSTM, CNN, CNN-LSTM, CNN-Bi-LSTM, and ARIMA. The empirical results confirm that PM25<bold>-</bold>CBL model outperforms other experimental methods for Air Quality HCMC dataset in terms of several common metrics including MSE, RMSE, MAE, and MAPE.</p>
<p>For future work, we focus on improving the PM2.5 concentration prediction model&#x2019;s performance by applying several advanced techniques such as evolutionary algorithms to the proposed approach. In addition, the Air Quality datasets in several cities in Vietnam are collected to verify the proposed model.</p>
</sec>
</body>
<back>
<sec><title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Fujita</surname></string-name>, <string-name><given-names>N. T.</given-names> <surname>Nguyen</surname></string-name> and <string-name><given-names>S. W.</given-names> <surname>Baik</surname></string-name></person-group>, &#x201C;<article-title>A fast and accurate approach for bankruptcy forecasting using squared logistics loss with GPU-based extreme gradient boosting</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>494</volume>, pp. <fpage>294</fpage>&#x2013;<lpage>310</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Le</surname></string-name></person-group>, &#x201C;<article-title>A comprehensive survey of imbalanced learning methods for bankruptcy prediction</article-title>,&#x201D; <source>IET Communications</source>, vol. <volume>16</volume>, no. <issue>5</issue>, pp. <fpage>433</fpage>&#x2013;<lpage>441</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q. H.</given-names> <surname>Doan</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Le and D</surname></string-name></person-group>. <person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Thai</surname></string-name></person-group>, &#x201C;<article-title>Optimization strategies of neural networks for impact damage classification of RC panels in a small dataset</article-title>,&#x201D; <source>Applied Soft Computing</source>, vol. <volume>102</volume>, pp. <fpage>107100</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. A.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>H. S.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>M. T.</given-names> <surname>Vo</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Le</surname></string-name></person-group>, &#x201C;<article-title>A novel framework for trash classification using deep transfer learning</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, no. <issue>1</issue>, pp. <fpage>178631</fpage>&#x2013;<lpage>178639</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Sharma</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>H. S.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>B. T.</given-names> <surname>Pham</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Crime rate detection using social media of different crime locations and Twitter part-of-speech tagger with Brown clustering</article-title>,&#x201D; <source>Journal of Intelligent and Fuzzy Systems</source>, vol. <volume>38</volume>, no. <issue>4</issue>, pp. <fpage>4287</fpage>&#x2013;<lpage>4299</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Kocheturov</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Momcilovic</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Bihorac</surname></string-name> and <string-name><given-names>P. M.</given-names> <surname>Pardalos</surname></string-name></person-group>, &#x201C;<article-title>Extended vertical lists for temporal pattern mining from multivariate time series</article-title>,&#x201D; <source>Expert Systems</source>, vol. <volume>36</volume>, no. <issue>5</issue>, pp. <fpage>e12448</fpage>, <year>2019</year>; <pub-id pub-id-type="pmid">33162636</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Gu</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Dai</surname></string-name></person-group>, &#x201C;<article-title>A novel active multi-source transfer learning algorithm for time series forecasting</article-title>,&#x201D; <source>Applied Intelligence</source>, vol. <volume>51</volume>, no. <issue>3</issue>, pp. <fpage>1326</fpage>&#x2013;<lpage>1350</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>M. T.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Hwang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Rho</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Improving electric energy consumption prediction using CNN and Bi-LSTM</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>9</volume>, no. <issue>20</issue>, pp. <fpage>4237</fpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>M. T.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Kieu</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Hwang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Rho</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Multiple electric energy consumption forecasting using a cluster-based strategy for transfer learning in smart building</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>20</volume>, no. <issue>9</issue>, pp. <fpage>2668</fpage>, <year>2020</year>; <pub-id pub-id-type="pmid">32392858</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. A.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Nguyen</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Le</surname></string-name></person-group>, &#x201C;<article-title>Brent oil price prediction using Bi-LSTM network</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>26</volume>, no. <issue>6</issue>, pp. <fpage>1307</fpage>&#x2013;<lpage>1317</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Yan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Memon</surname></string-name></person-group>, &#x201C;<article-title>A novel deep learning framework: Prediction and analysis of financial time series using CEEMD and LSTM</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>159</volume>, pp. <fpage>113609</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. T.</given-names> <surname>Vo</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Vu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Nguyen</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Bui</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Le</surname></string-name></person-group>, &#x201C;<article-title>Predicting monthly household water consumption</article-title>,&#x201D; in <conf-name>Proc. of Int. Conf. on Computing and Communication Technologies</conf-name>, <conf-loc>Ho Chi Minh City, Vietnam</conf-loc>, pp. <fpage>720</fpage>&#x2013;<lpage>724</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P. Y.</given-names> <surname>Wong</surname></string-name>, <string-name><given-names>H. Y.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>Y. C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>Y. R.</given-names> <surname>Chern</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Using a land use regression model with machine learning to estimate ground level PM2.5</article-title>,&#x201D; <source>Environmental Pollution</source>, vol. <volume>277</volume>, pp. <fpage>116846</fpage>, <year>2021</year>; <pub-id pub-id-type="pmid">33735646</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Bozda&#x011F;</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Dokuz</surname></string-name> and <string-name><given-names>O. B.</given-names> <surname>G&#x00F6;k&#x00E7;ek</surname></string-name></person-group>, &#x201C;<article-title>Spatial prediction of PM10 concentration using machine learning algorithms in Ankara, Turkey</article-title>,&#x201D; <source>Environmental Pollution</source>, vol. <volume>263</volume>, pp. <fpage>114635</fpage>, <year>2020</year>; <pub-id pub-id-type="pmid">33618491</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="web"><collab>T. Global Burden of Diseases, Injuries, and Risk Factors Study 2015</collab>. <ext-link ext-link-type="uri" xlink:href="https://publichealth.wustl.edu/global-burden-diseases-injuries-risk-factors-study-2015">https://publichealth.wustl.edu/global-burden-diseases-injuries-risk-factors-study-2015</ext-link>, <comment>accessed on November 18, 2022</comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>S. S.</given-names> <surname>Woo</surname></string-name></person-group>, &#x201C;<article-title>Real time localized air quality monitoring and prediction through mobile and fixed IoT sensing network</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>89584</fpage>&#x2013;<lpage>89594</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Zeinalnezhad</surname></string-name>, <string-name><given-names>A. G.</given-names> <surname>Chofreh</surname></string-name>, <string-name><given-names>F. A.</given-names> <surname>Goni</surname></string-name> and <string-name><given-names>J. J.</given-names> <surname>Kleme&#x0161;</surname></string-name></person-group>, &#x201C;<article-title>Air pollution prediction using semi-experimental regression model and Adaptive Neuro-Fuzzy Inference System</article-title>,&#x201D; <source>Journal of Cleaner Production</source>, vol. <volume>261</volume>, pp. <fpage>121218</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Sch&#x00FC;rholz</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kubler</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zaslavsky</surname></string-name></person-group>, &#x201C;<article-title>Artificial intelligence-enabled context-aware air quality prediction for smart cities</article-title>,&#x201D; <source>Journal of Cleaner Production</source>, vol. <volume>271</volume>, pp. <fpage>121941</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J. C.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Lin</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Air quality prediction at new stations using spatially transferred bi-directional long short-term memory network</article-title>,&#x201D; <source>Science of the Total Environment</source>, vol. <volume>705</volume>, pp. <fpage>135771</fpage>, <year>2020</year>; <pub-id pub-id-type="pmid">31972931</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. S.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>H. T.</given-names> <surname>Chiao</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Abimannan</surname></string-name>, <string-name><given-names>Y. P.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Tsai</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>An LSTM-based aggregated model for air pollution forecasting</article-title>,&#x201D; <source>Atmospheric Pollution Research</source>, vol. <volume>11</volume>, no. <issue>8</issue>, pp. <fpage>1451</fpage>&#x2013;<lpage>1463</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. J.</given-names> <surname>Cohen</surname></string-name>, <string-name><given-names>H. R.</given-names> <surname>Anderson</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Ostro</surname></string-name>, <string-name><given-names>K. D.</given-names> <surname>Pandey</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Krzyzanowski</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>The global burden of disease due to outdoor air pollution</article-title>,&#x201D; <source>Journal of Toxicology and Environmental Health, Part A</source>, vol. <volume>68</volume>, no. <issue>13&#x2013;14</issue>, pp. <fpage>1301</fpage>&#x2013;<lpage>1307</lpage>, <year>2005</year>; <pub-id pub-id-type="pmid">16024504</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Luo</surname></string-name> and <string-name><given-names>J. R.</given-names> <surname>Fan</surname></string-name></person-group>, &#x201C;<article-title>Analysis and accurate prediction of ambient PM2. 5 in China using multi-layer perceptron</article-title>,&#x201D; <source>Atmospheric Environment</source>, vol. <volume>232</volume>, pp. <fpage>117534</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Qu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Cao</surname></string-name></person-group>, &#x201C;<article-title>Application of the XGBoost machine learning method in PM2. 5 prediction: A case study of Shanghai</article-title>,&#x201D; <source>Aerosol and Air Quality Research</source>, vol. <volume>20</volume>, no. <issue>1</issue>, pp. <fpage>128</fpage>&#x2013;<lpage>138</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yuan</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Prediction and analysis of PM2. 5 in Fuling District of Chongqing by artificial neural network</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>33</volume>, pp. <fpage>517</fpage>&#x2013;<lpage>524</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>U.</given-names> <surname>Pak</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Ryu</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ryom</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Juhyok</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Deep learning-based PM2. 5 prediction considering the spatiotemporal correlations: A case study of Beijing, China</article-title>,&#x201D; <source>Science of the Total Environment</source>, vol. <volume>699</volume>, pp. <fpage>133561</fpage>, <year>2020</year>; <pub-id pub-id-type="pmid">31689669</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Spatial ensemble prediction of hourly PM2. 5 concentrations around Beijing railway station in China</article-title>,&#x201D; <source>Air Quality, Atmosphere &#x0026; Health</source>, vol. <volume>13</volume>, no. <issue>5</issue>, pp. <fpage>563</fpage>&#x2013;<lpage>573</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Sun</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Hourly PM2. 5 concentration forecasting based on feature extraction and stacking-driven ensemble model for the winter of the Beijing-Tianjin-Hebei area</article-title>,&#x201D; <source>Atmospheric Pollution Research</source>, vol. <volume>11</volume>, no. <issue>6</issue>, pp. <fpage>110</fpage>&#x2013;<lpage>121</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Y. K.</given-names> <surname>Hsu</surname></string-name>, <string-name><given-names>H. Y.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>H. J.</given-names> <surname>Jau</surname></string-name></person-group>, &#x201C;<article-title>Air pollution prediction based on factory-aware attentional LSTM neural network</article-title>,&#x201D; <source>Computing</source>, vol. <volume>103</volume>, pp. <fpage>75</fpage>&#x2013;<lpage>98</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>J. C.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>V. J.</given-names> <surname>Gan</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A Lag-FLSTM deep learning network based on Bayesian optimization for multi-sequential-variant PM2. 5 prediction</article-title>,&#x201D; <source>Sustainable Cities and Society</source>, vol. <volume>60</volume>, pp. <fpage>102237</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tong</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Meng</surname></string-name></person-group>, &#x201C;<article-title>Fine-grained prediction of PM2. 5 concentration based on multisource data and deep learning</article-title>,&#x201D; <source>Atmospheric Pollution Research</source>, vol. <volume>11</volume>, no. <issue>10</issue>, pp. <fpage>1728</fpage>&#x2013;<lpage>1737</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="web"><collab>Dataset on Air Quality in Vietnam in 2020</collab>. <ext-link ext-link-type="uri" xlink:href="https://data.opendevelopmentmekong.net/dataset/timelines-dataset-on-air-quality-in-vietnam">https://data.opendevelopmentmekong.net/dataset/timelines-dataset-on-air-quality-in-vietnam</ext-link>, <comment>accessed on February 18, 2021</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>W. C.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Modeling long- and short-term temporal patterns with deep neural networks</article-title>,&#x201D; in <conf-name>Proc. of Int. ACM SIGIR Conf. on Research &#x0026; Development in Information Retrieval</conf-name>, <conf-loc>Ann Arbor, MI, USA</conf-loc>, pp. <fpage>95</fpage>&#x2013;<lpage>104</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Hochreiter</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Long short-term memory</article-title>,&#x201D; <source>Neural Computation</source>, vol. <volume>9</volume>, no. <issue>8</issue>, pp. <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>, <year>1997</year>; <pub-id pub-id-type="pmid">9377276</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Graves</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Framewise phoneme classification with bidirectional LSTM and other neural network architectures</article-title>,&#x201D; <source>Neural Networks</source>, vol. <volume>18</volume>, no. <issue>5&#x2013;6</issue>, pp. <fpage>602</fpage>&#x2013;<lpage>610</lpage>, <year>2005</year>; <pub-id pub-id-type="pmid">16112549</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>