<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">38162</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.038162</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Research on PM<sub>2.5</sub> Concentration Prediction Algorithm Based on Temporal and Spatial Features</article-title>
<alt-title alt-title-type="left-running-head">Research on PM<sub>2.5</sub> Concentration Prediction Algorithm Based on Temporal and Spatial Features</alt-title>
<alt-title alt-title-type="right-running-head">Research on PM<sub>2.5</sub> Concentration Prediction Algorithm Based on Temporal and Spatial Features</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yu</surname><given-names>Song</given-names></name><email>ys@csu.edu.cn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Chen</given-names></name></contrib>
<aff><institution>School of Computer Science and Engineering, Central South University</institution>, <addr-line>Changsha, 410000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Song Yu. Email: <email>ys@csu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>1</day><month>5</month><year>2023</year></pub-date>
<volume>75</volume>
<issue>3</issue>
<fpage>5555</fpage>
<lpage>5571</lpage>
<history>
<date date-type="received"><day>29</day><month>11</month><year>2022</year>
</date>
<date date-type="accepted"><day>15</day><month>3</month><year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Yu and Wang</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Yu and Wang</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_38162.pdf"></self-uri>
<abstract>
<p>PM<sub>2.5</sub> has a non-negligible impact on visibility and air quality as an important component of haze and can affect cloud formation and rainfall and thus change the climate, and it is an evaluation indicator of air pollution level. Achieving PM<sub>2.5</sub> concentration prediction based on relevant historical data mining can effectively improve air pollution forecasting ability and guide air pollution prevention and control. The past methods neglected the impact caused by PM<sub>2.5</sub> flow between cities when analyzing the impact of inter-city PM<sub>2.5</sub> concentrations, making it difficult to further improve the prediction accuracy. However, factors including geographical information such as altitude and distance and meteorological information such as wind speed and wind direction affect the flow of PM<sub>2.5</sub> between cities, leading to the change of PM<sub>2.5</sub> concentration in cities. So a PM<sub>2.5</sub> directed flow graph is constructed in this paper. Geographic and meteorological data is introduced into the graph structure to simulate the spatial PM<sub>2.5</sub> flow transmission relationship between cities. The introduction of meteorological factors like wind direction depicts the unequal flow relationship of PM<sub>2.5</sub> between cities. Based on this, a PM<sub>2.5</sub> concentration prediction method integrating spatial-temporal factors is proposed in this paper. A spatial feature extraction method based on weight aggregation graph attention network (WGAT) is proposed to extract the spatial correlation features of PM<sub>2.5</sub> in the flow graph, and a multi-step PM<sub>2.5</sub> prediction method based on attention gate control loop unit (AGRU) is proposed. The PM<sub>2.5</sub> concentration prediction model WGAT-AGRU with fused spatiotemporal features is constructed by combining the two methods to achieve multi-step PM<sub>2.5</sub> concentration prediction. Finally, accuracy and validity experiments are conducted on the KnowAir dataset, and the results show that the WGAT-AGRU model proposed in the paper has good performance in terms of prediction accuracy and validates the effectiveness of the model.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Spatiotemporal fusion</kwd>
<kwd>PM<sub>2.5</sub> concentration prediction</kwd>
<kwd>graph neural network</kwd>
<kwd>recurrent neural network</kwd>
<kwd>attention mechanism</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Central South University Research Programme of Advanced Interdisciplinary Studies</funding-source>
<award-id>2023QYJC041</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>PM<sub>2.5</sub> is the airborne particulate matter with an equivalent diameter less than or equal to 2.5 &#x03BC;m, also known as fine particulate matter, which is an important component of haze and has a non-negligible impact on visibility and air quality, and is an evaluation indicator of air pollution level. PM<sub>2.5</sub> can affect cloud formation and rainfall and thus change the climate. Using historical PM<sub>2.5</sub> concentration data released by monitoring stations to accurately predict future PM<sub>2.5</sub> concentrations can help meteorological experts predict the situation of haze and help relevant departments grasp air quality information in time, which is beneficial for the state to make decisions on air pollution in advance and provide help for urban air pollution management and related policy formulation.</p>
<p>Air pollutants such as PM<sub>2.5</sub> stay and accumulate in the atmosphere and flow between cities influenced by altitude, distance, and geographic and meteorological factors. So the concentrations show obvious spatial and temporal correlations. Currently, Recurrent neural networks (RNNs) and their variants become the mainstream of time series prediction, and Graph Convolution Networks (GCN) models are mostly used to extract spatial features. In recent studies, some scholars have considered using PM<sub>2.5</sub> concentration-related data to construct spatially correlated graphs and combining the graph structure with neural networks to mine spatiotemporal features for PM<sub>2.5</sub> concentration prediction. Lin et al. [<xref ref-type="bibr" rid="ref-1">1</xref>] proposed a GC-DCRNN model to calculate similarities by geographic features of neighborhoods and construct undirected graphs based on these similarities. And then captures the spatial correlation of PM<sub>2.5</sub> by expanding convolutional RNNs [<xref ref-type="bibr" rid="ref-2">2</xref>] and captures the temporal dependence using sequences. Qi et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed a GC-LSTM prediction method that takes monitoring stations as graph vertices and extracted spatial correlation features using GCN and then combined them with Long Short-Term Memory (LSTM) to extract PM<sub>2.5</sub> temporal correlation. The above method uses geographic features such as altitude distance between buildings or cities in a small area and meteorological features such as wind and temperature as influencing factors of PM<sub>2.5</sub> concentration to construct an undirected graph about geographical and meteorological features. However, factors such as altitude, mountain range separation, and wind direction between cities can also affect the PM<sub>2.5</sub> flow, resulting in unequal PM<sub>2.5</sub> influence between cities, and the undirected structure map cannot fit this unequal PM<sub>2.5</sub> flow relationship between cities well.</p>
<p>Therefore, a PM<sub>2.5</sub> directed flow graph is constructed in this paper to simulate the inter-city PM<sub>2.5</sub> directed flow process. The spatial features in the flow graph are extracted by the graph neural network. The recurrent neural network can effectively extract the temporal features in the historical data, and combine the spatial features to realize the multi-step prediction of PM<sub>2.5</sub> concentration by fusing the temporal and spatial features.</p>
<p>Contributions can be summarized as follows:
<list list-type="bullet">
<list-item>
<p>To reveal the inter-city PM<sub>2.5</sub> spatiotemporal flow relationship, the inter-city PM<sub>2.5</sub> directed flow graph is constructed by combining relevant geographical features such as distance, altitude, mountain range, and meteorological features such as wind and wind direction.</p></list-item>
<list-item>
<p>A PM<sub>2.5</sub> concentration prediction method incorporating spatiotemporal features is proposed. First, this paper proposes WGAT, which updates the feature representation of the central vertex by aggregating the graph vertices through a message-passing paradigm and uses the graph attention layer to weigh the similarity of city spatial correlation features to extract deep spatial features. Then the time-dependent features of PM<sub>2.5</sub> are captured by Gated Recurrent Unit (GRU) and focused on the historical time-step information highly correlated with the current prediction time-step by a time-series attention mechanism. The two modules are combined to construct WGAT-AGRU, a PM<sub>2.5</sub> concentration prediction model incorporating spatiotemporal features.</p></list-item>
<list-item>
<p>We conduct Experiments on the KnowAir dataset to verify the validity and reasonableness of the model component setup.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Traditional Concentration Forecasting Methods</title>
<p>Air pollutant concentrations were first predicted by numerical simulations and statistical models [<xref ref-type="bibr" rid="ref-4">4</xref>]. The statistical model is mainly based on the Autoregressive Integrated Moving Average (ARIMA), which uses the historical series of PM<sub>2.5</sub> concentrations as model inputs to predict the PM<sub>2.5</sub> concentration values at the next moment. Zhang et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] compared PM<sub>2.5</sub> concentrations with other pollutants and with meteorological parameters and applied the ARIMA model to forecast PM<sub>2.5</sub> concentrations. Venkataraman et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] analyzed the factors influencing PM<sub>2.5</sub> in Mumbai by wavelet and regression analysis. Tai et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] predicted PM<sub>2.5</sub> concentrations in the United States by incorporating information on the characteristics of air pollutants related to PM<sub>2.5</sub>, such as CO, NO<sub>2,</sub> and SO<sub>2</sub> into the model through multiple regression. Traditional air pollutant concentration prediction models assume that PM<sub>2.5</sub> concentration has a linear relationship, cannot use a large amount of PM<sub>2.5</sub> concentration data for prediction, and cannot effectively mine historical data feature information.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Machine Learning Concentration Prediction Methods</title>
<p>To capture the nonlinear characteristics of PM<sub>2.5</sub> concentration, machine learning prediction algorithms emerge, mainly Random Forest (RF), Support Vector Machine (SVM), and Artificial Neural Network (ANN) algorithms. Shamsoddini et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] used RF for feature selection to improve the prediction performance of the PM<sub>2.5</sub> concentration. Dong et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] combined the latent Dirichlet allocation, points of interest, and wavelet decomposition based on SVM to improve the PM<sub>2.5</sub> concentration prediction accuracy. Wang et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] combined ARIMA with SVM to capture linear relationships by ARIMA and used SVM to model nonlinearities. Asadollahfardi et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] used historical data such as air quality and humidity in Tehran to train ANNs to predict PM<sub>2.5</sub> concentrations. Mao et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] used backpropagation multilayer perceptron to predict PM<sub>2.5</sub> concentrations and the model had good generalization ability. McKendry [<xref ref-type="bibr" rid="ref-13">13</xref>] confirmed that the accuracy of PM<sub>2.5</sub> concentration prediction of Multilayer Perceptron-based (MLP-based) ANN is not significantly improved compared with traditional statistical models. Machine learning algorithms are difficult to dig deep into the deep feature information contained in a large amount of data, which limits the accuracy of PM<sub>2.5</sub> concentration prediction.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Deep Learning Concentration Prediction Methods</title>
<p>Deep neural networks mine the deep spatial and temporal features contained in the large number of PM<sub>2.5</sub> concentration data to improve the prediction accuracy of PM<sub>2.5</sub> concentration. RNNs and their variants, LSTM, capture temporal dependencies in data sequences. Ong et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] used deep RNN to predict PM<sub>2.5</sub> concentrations in Japan, which is much more accurate than traditional models. Authors in [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>] used LSTMs to predict PM<sub>2.5</sub> concentrations. GRU is a streamlined variant of LSTM with fewer parameters and a simpler structure for faster convergence. Chi et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] regarded dissolved oxygen concentration as a time-series data and achieved a better fit with GRU by wavelet transformation. The spatial correlation characteristics of PM<sub>2.5</sub> can be extracted by combining image or graph structure with the neural network. Authors in [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>] used Convolutional Neural Networks (CNN) to capture the spatial correlation of PM<sub>2.5</sub> in the image grid for prediction, and the prediction accuracy was further improved compared with RNN-based prediction methods. Cheng et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed the EAT-GCN gas concentration prediction method, using GRU to capture time dependence and graph convolutional neural network to capture spatial features, and obtained high prediction accuracy. These methods ignore the unequal influence of geographic meteorological factors on the inter-city flow of PM<sub>2.5</sub>.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>The WGAT-AGRT architecture is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. In general, it consists of three parts: (1) PM<sub>2.5</sub> directed flow graph construction, (2) WGAT-based spatial features, and (3) AGRU-based spatiotemporal fusion multi-step prediction. The detailed design of these three phases is described as follows.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The architecture of the WGAT-AGRU</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>PM<sub>2.5</sub> Directed Flow Graph Construction</title>
<p>The PM<sub>2.5</sub> directed flow graph is constructed to fit the inter-city flow of PM<sub>2.5</sub> based on geographic and meteorological data, with the city as the vertex, PM<sub>2.5</sub> concentration data, meteorological,latitude, and longitude and altitude factors as the vertex attributes. The edge information of <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>G</mml:mi></mml:math></inline-formula> is constructed and refined in two steps in turn by constructing the adjacency matrix based on the altitude distance information and calculating the edge weights based on the geographic and meteorological data. After analysis, eight meteorological features are selected in this paper as the influencing factors of city vertices in <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>G</mml:mi></mml:math></inline-formula>, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Weather characteristics of city vertices</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Weather characteristics</th>
<th>Unit</th>
<th>Description</th>
<th>Relationship between weather characteristics and PM<sub>2.5</sub> concentration</th>
</tr>
</thead>
<tbody>
<tr>
<td>relative_humidity&#x002B;950</td>
<td><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></td>
<td>Relative humidity</td>
<td>High humidity promotes PM<sub>2.5</sub> formation</td>
</tr>
<tr>
<td>2m_temperature</td>
<td><inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>K</mml:mi></mml:math></inline-formula></td>
<td>2m temperature</td>
<td>Correlation</td>
</tr>
<tr>
<td>boundary_layer_height</td>
<td><inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>m</mml:mi></mml:math></inline-formula></td>
<td>Boundary layer height</td>
<td>Negative correlation</td>
</tr>
<tr>
<td>k_index</td>
<td><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>K</mml:mi></mml:math></inline-formula></td>
<td>K index</td>
<td>Negative correlation</td>
</tr>
<tr>
<td>surface_pressure</td>
<td><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:math></inline-formula></td>
<td>Surface pressure</td>
<td>Negative correlation</td>
</tr>
<tr>
<td>total_precipitation</td>
<td><inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>m</mml:mi></mml:math></inline-formula></td>
<td>Total precipitation</td>
<td>Negative correlation</td>
</tr>
<tr>
<td>h_component_of_wind&#x002B;950</td>
<td><inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>m</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>s</mml:mi></mml:math></inline-formula></td>
<td>Horizontal component of wind speed</td>
<td>Negative correlation</td>
</tr>
<tr>
<td>v_component_of_wind&#x002B;950</td>
<td><inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>m</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>s</mml:mi></mml:math></inline-formula></td>
<td>Vertical component of wind speed</td>
<td>Negative correlation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Construct adjacency matrix:</bold> The adjacency matrix is a matrix used to describe the connection relationship of graph vertices in the graph structure. And in the PM<sub>2.5</sub> flow graph, the vertices represent cities. The adjacency matrix is constructed in this paper based on the altitude and distance information. The two city vertices whose distance and altitude between cities are less than the distance threshold <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the altitude threshold <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can construct the connection, respectively.</p>
<p>The earth is approximated as a sphere, so the Haversine formula is introduced to calculate the geodesic distance between two points on the sphere [<xref ref-type="bibr" rid="ref-21">21</xref>], and the shortest straight-line distance between city <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>i</mml:mi></mml:math></inline-formula> and city <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>j</mml:mi></mml:math></inline-formula> on the earth&#x2019;s sphere is calculated as shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.</p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mi>r</mml:mi><mml:mi>arcsin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:msup><mml:mi>sin</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>lat</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>lat</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>lat</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>lat</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>sin</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>lon</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>lon</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>r</mml:mi></mml:math></inline-formula> is the radius of the earth, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> correspond to the longitude and latitude of cities, respectively, and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the shortest straight-line distance between the two cities on the earth&#x2019;s sphere.</p>
<p>The altitude difference and mountain range blockage between the cities&#x2019; fixed point links is calculated by Bresenham linear interpolation algorithm. And the highest altitude difference <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> between cities is shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>.</p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">sup</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>Bresenham</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> correspond to the pixels in the altitude map obtained by the latitude and longitude mapping of cities <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>j</mml:mi></mml:math></inline-formula>, respectively. <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the mean altitude of the <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>&#x03C1;</mml:mi></mml:math></inline-formula> pixel area. And the Bresenham linear algorithm outputs the altitude of the pixel area interpolated by a straight line.</p>
<p>The values of the elements in the adjacency matrix corresponding to cities <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>j</mml:mi></mml:math></inline-formula>, are given as <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.</p>
<p><disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the element value of column <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:math></inline-formula> in row <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:math></inline-formula> of the adjacency matrix, representing the connection relationship between the city vertices <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>j</mml:mi></mml:math></inline-formula> in the PM<sub>2.5</sub> spatial correlation graph. <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the Heaviside function, which has a function value of 0 when the input value is less than 0 and a function value of 1 when it is greater than or equal to 0.</p>
<p><bold>Calculate edge weights:</bold> The edge weights are calculated based on geographic and meteorological data. The wind force and the altitude difference between cities affect the inter-city PM<sub>2.5</sub> flow. This paper simplifies the pollutant dispersion equation to simulate the magnitude of the impact of spatial transport of planar pollutants under the effect of wind and introduces the city altitude information. Then the impact w of PM<sub>2.5</sub> concentration in the source city <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>a</mml:mi></mml:math></inline-formula> on the polluted city <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>b</mml:mi></mml:math></inline-formula> can be calculated by <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>.</p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mi>d</mml:mi></mml:mfrac><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03C9;</mml:mi></mml:math></inline-formula> is the city altitude factor when the flow direction is from a high-altitude city to a low-altitude city, &#x03C9; is equal to 1, otherwise, <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>&#x03C9;</mml:mi></mml:math></inline-formula> is equal to the ratio of the altitude of the two places. <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>d</mml:mi></mml:math></inline-formula> represents the distance between cities <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>a</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>b</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> represents the wind speed of city <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>a</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is the angle between the wind direction of city <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>a</mml:mi></mml:math></inline-formula> and the direction from city <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>a</mml:mi></mml:math></inline-formula> to city <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>b</mml:mi></mml:math></inline-formula>. ReLU is the linear rectification function, and the function value is 0 in the negative semi-axis, so when the angle <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is greater than 90 degrees, PM<sub>2.5</sub> cannot flow to city <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>b</mml:mi></mml:math></inline-formula> through wind action, and the corresponding PM<sub>2.5</sub> concentration impact value <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>w</mml:mi></mml:math></inline-formula> is 0. The update of the edge weights is completed by updating the attribute information of the city vertices in the flow graph. And then the simulation of the estimation of PM<sub>2.5</sub> concentration impact between cities is realized.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>WGAT-AGRU</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Spatial Feature Extraction</title>
<p>The model extracts spatial features from the PM<sub>2.5</sub> directed flow graph. Given the PM<sub>2.5</sub> flow graph <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>G</mml:mi></mml:math></inline-formula>, the feature representation of the central city vertex <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>i</mml:mi></mml:math></inline-formula> is updated by aggregating the adjacency graph vertices through the message-passing paradigm as shown in the following equations.</p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></disp-formula></p>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi>E</mml:mi></mml:math></disp-formula></p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mi>&#x03B6;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>V</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>E</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>P</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>Q</mml:mi></mml:math></inline-formula> are the city vertex set and the edge set of <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>G</mml:mi></mml:math></inline-formula>, the PM<sub>2.5</sub> concentration data, meteorological features of city vertex <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>i</mml:mi></mml:math></inline-formula>, and edge attribute features, respectively. <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi mathvariant="normal">&#x03A8;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:math></inline-formula> are linear transformations. <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the feature representation of city vertex <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>i</mml:mi></mml:math></inline-formula> at time <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>t</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are the PM<sub>2.5</sub> inflow impact and outflow impact of city vertex <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>i</mml:mi></mml:math></inline-formula> at time <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>t</mml:mi></mml:math></inline-formula>, respectively.</p>
<p>The deeper vertex spatial feature representation is extracted by graph attention weighting, and the graph attention layer operation is shown in the following equations.</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B6;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></disp-formula></p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mi>W</mml:mi><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>W</mml:mi><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:munder><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the feature representation of city vertex <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>i</mml:mi></mml:math></inline-formula> at time <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>t</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>W</mml:mi></mml:math></inline-formula> is the linear transformation matrix; <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> is a splicing operation. The feedforward neural network <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> maps the features to real numbers and obtains the similarity degree <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the feature representation of city vertex <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>i</mml:mi></mml:math></inline-formula> and its neighbor vertex <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>j</mml:mi></mml:math></inline-formula> through LeakyReLU activation. Then the Softmax function calculates the attention weight <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of city vertex <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>i</mml:mi></mml:math></inline-formula> and its neighbor vertex <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>j</mml:mi></mml:math></inline-formula>. Finally, the attention-weighted feature representation <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of city vertex <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>i</mml:mi></mml:math></inline-formula> at time <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>t</mml:mi></mml:math></inline-formula> is obtained by weighting all neighboring vertex feature representations at time <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>t</mml:mi></mml:math></inline-formula>, so the spatial correlation feature extraction of city vertices in the PM<sub>2.5</sub> flow graph is realized in this paper.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Spatiotemporal Fusion Prediction</title>
<p>GRU is similar to LSTM and is also able to capture the long-term dependence of urban PM<sub>2.5</sub> concentration data. Combined with the time-series attention mechanism focusing on the highly correlated historical time step information of the current prediction time step, a long-term prediction of PM<sub>2.5</sub> concentration can be achieved. For spatiotemporal features fusion, the city spatial features extracted by WGAT are used as external features and input to a single prediction unit GRU together with the corresponding historical PM<sub>2.5</sub> concentration data of cities and meteorological features.</p>
<p>First, the loop structure of the GRU outputs a multi-step prediction of the hidden state, as shown in the following equations.</p>
<p><disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msubsup><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2217;</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2217;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the input of the spatiotemporal fusion prediction unit of the city vertex <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>i</mml:mi></mml:math></inline-formula> at the prediction time step <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>t</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>W</mml:mi></mml:math></inline-formula> are parameters that can be trained and learned in GRU. <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the Sigmoid activation function. <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msubsup><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are reset gate and update gate structures in GRU, respectively. <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is the hidden state output of the prediction unit of the previous prediction time step, <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the candidate hidden state of the prediction unit of the prediction time step <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>t</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the hidden state output of the prediction unit of the prediction time step <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>t</mml:mi></mml:math></inline-formula>.</p>
<p>Then the hidden state of the prediction unit output is weighted by time series attention, so that the current forecast time step focuses on the key historical time step information, as shown in the following equations.</p>
<p><disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>W</mml:mi><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munderover><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the attention score of the hidden state <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of the city vertex <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>i</mml:mi></mml:math></inline-formula> at time step <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>t</mml:mi></mml:math></inline-formula> output through the fully connected network. <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the attention weight of the city vertex <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>i</mml:mi></mml:math></inline-formula> at the current time step <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>t</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the spatiotemporal fusion feature hidden state.</p>
<p>Finally, <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> output from AGRU is linearly transformed through the fully connected layer to obtain the predicted PM<sub>2.5</sub> concentration value <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msubsup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> for city vertex i at prediction time step <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>t</mml:mi></mml:math></inline-formula>.</p>
<p>The current PM<sub>2.5</sub> concentration prediction combined with the meteorological feature at the next moment predicts the output at the next moment. And at the next moment, the predicted value can be used as the PM<sub>2.5</sub> concentration at that time and combined with the corresponding meteorological feature to predict the PM<sub>2.5</sub> concentration. The multi-step prediction of PM<sub>2.5</sub> concentration is achieved through an iterative process.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>PM<sub>2.5</sub> Concentration Prediction Based on Temporal and Spatial Features</title>
<p>The prediction of PM<sub>2.5</sub> concentration is generally regarded as the prediction of spatiotemporal series. Assume that the PM<sub>2.5</sub> concentration of city vertices at time <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>t</mml:mi></mml:math></inline-formula> is <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi>N</mml:mi></mml:math></inline-formula> represents the number of vertices in the PM<sub>2.5</sub> directed flow graph <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>G</mml:mi></mml:math></inline-formula>, that is, the total number of cities studied in the paper. <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. represent the vertex feature matrix and edge attribute matrix of all city vertices in PM<sub>2.5</sub> directed flow graph <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>G</mml:mi></mml:math></inline-formula> at time t, separately, and M is the number of edges in <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>G</mml:mi></mml:math></inline-formula>. The vertex feature P represents the meteorological features such as temperature and humidity of the city, and the edge attribute matrix <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>Q</mml:mi></mml:math></inline-formula> is the meteorological characteristics such as wind direction and wind speed, and geographical features such as distance and altitude required to calculate the edge weight, that is, the features related to the impact of PM<sub>2.5</sub> concentration between cities.</p>
<p>The multi-step of PM<sub>2.5</sub> concentration prediction in the paper is realized by iteration. In order to predict PM<sub>2.5</sub> concentration for a period of time in the future, the vertex feature matrix <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, edge attribute matrix <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and PM<sub>2.5</sub> directed flow graph structure <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>G</mml:mi></mml:math></inline-formula> of the future <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mi>T</mml:mi></mml:math></inline-formula> time step are used as the input of the prediction model. And the multi-step prediction of PM<sub>2.5</sub> concentration can be defined by <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>.</p>
<p><disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>G</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mover><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The iteration multi-step prediction process of the model is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The PM<sub>2.5</sub> prediction model uses the predicted PM<sub>2.5</sub> concentration <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> in the previous prediction time step as the input of the model to predict the next time step. And the PM<sub>2.5</sub> concentration of the city in the future <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>T</mml:mi></mml:math></inline-formula> time step is predicted through continuous iterative calculation. The iterative process is shown in <xref ref-type="disp-formula" rid="eqn-21">Eq. (21)</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The iteration multi-step prediction process of the proposed model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-2.tif"/>
</fig>
<p><disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The prediction operation <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of PM<sub>2.5</sub> concentration in the future <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>T</mml:mi></mml:math></inline-formula> time step is achieved by iterating the <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>T</mml:mi></mml:math></inline-formula> times prediction model <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The predicted value <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of PM<sub>2.5</sub> concentration at time step <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> in the future can be expressed by <xref ref-type="disp-formula" rid="eqn-22">Eq. (22)</xref>.</p>
<p><disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>&#x03C4;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is the predicted value of the PM<sub>2.5</sub> concentration of the city output by the prediction model at the last time. <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the vertex feature corresponding to the city vertex of the current prediction time step. <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the edge attribute matrix used to calculate the edge weight between city vertices in the current prediction time step. And <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mi>G</mml:mi></mml:math></inline-formula> is the constructed PM<sub>2.5</sub> directed flow graph structure. <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:math></inline-formula> represents the corresponding parameters of the neural network in the prediction model. <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:msup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the predicted value of PM<sub>2.5</sub> concentration in the current prediction time step output by the prediction model.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>We select 184 cities in China (103&#x00B0;E&#x2013;120&#x00B0;E, 28&#x00B0;N&#x2013;42&#x00B0;N) sampled at 3-h intervals for a total of 4 years (January 1, 2015, to December 31, 2018) from the KnowAir dataset. The heating measures are taken in northern cities of China from early November to late February every year, so the PM<sub>2.5</sub> concentration in northern cities will change dramatically during this period. And these cities are mainly dominated by northwest and north winds. So, the data set is divided into two subsets, which are the full data set and the heating season data set, and divided into the training set (67%) and the test set (33%). During the heating season, more coal is burned in northern cities in China and the prevailing north wind leads the inter-city PM<sub>2.5</sub> concentration more affected by wind transmission. We considered the influence of wind speed and direction when constructing the PM<sub>2.5</sub> directed flow graph. Through the comparative experiments on the heating data set, the fitting effect of PM<sub>2.5</sub> directed flow graph on PM<sub>2.5</sub> flow between cities and the ability of the WGAT-AGRU model proposed in this paper to capture PM<sub>2.5</sub> flow in the graph under the influence of wind can be analyzed.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Setup</title>
<p><bold>Evaluation Metrics.</bold> In this paper, Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and R Squared (<inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) are used to evaluate the prediction accuracy, and Critical Success Index (CSI), Probability Of Detection (POD), and False Alarm Rate (FAR) are used to evaluate the pollution forecasting capability of the model. The smaller the RMSE and MAE represent the higher the prediction accuracy of the model. The coefficient of determination <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is used to evaluate the degree of fit of the prediction model, and the better fit is the larger value of <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> ([0, 1]). Generally speaking, the value of <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is larger than 0.5, which indicates that the fit of the prediction model is excellent. The larger value of CSI ([0, 1]) and POD ([0, 1]) and the smaller value of FAR ([0, 1]) means stronger air pollution forecasting ability of the prediction model.</p>
<p><bold>Parameter Settings.</bold> The model proposed in this paper is implemented by a computer with an 8-core, 16-thread Intel i9-9900KF CPU and a GDDR6 8 GB NVIDIA GeForce RTX 2080ti graphics card, trained by PyTorch and PyTorch Geometric. In this paper, the loss function is MSELoss and the optimization is AdamW. The batch size is set to 64, the epochs are set to 100, and fix the learning rate is set to 0.0005.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental Results and Analysis</title>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Accuracy Analysis</title>
<p>To analyze the accuracy of the proposed WGAT-AGRU, a comparison experiment is conducted with several PM<sub>2.5</sub> concentration prediction models on the overall data set and the heating season data set.</p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> (prediction time step &#x003D; 24 h) demonstrates the prediction results of each model on the full data set, and the prediction step of each model was uniformly set to 24 h. In the STA-ResCNN [<xref ref-type="bibr" rid="ref-22">22</xref>] model, correlation analysis technology is used to screen the spatial information of pollution and meteorology and combine it with the time series to complete the prediction of PM<sub>2.5</sub> concentration in the city. WGAT-AGRU achieves a minimum RMSE of 16.61 (9.3% lower than MLP) and a minimum MAE of 13.47 (8.8% lower than MLP). And WGAT-AGRU shows the best fitting ability with a maximum <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> of 66.71%.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of the proposed model and various models on the full data set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Methods</th>
<th>RMSE</th>
<th>MAE</th>
<th>R<sup>2</sup></th>
<th>CSI</th>
<th>POD</th>
<th>FAR</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td>MLP</td>
<td>18.32</td>
<td>14.77</td>
<td>59.03%</td>
<td>53.05%</td>
<td>66.18%</td>
<td>27.23%</td>
</tr>
<tr>
<td></td>
<td>GRU</td>
<td>17.31</td>
<td>13.86</td>
<td>63.29%</td>
<td>54.19%</td>
<td>65.69%</td>
<td>24.73%</td>
</tr>
<tr>
<td>Prediction time step &#x003D; </td>
<td>GC-LSTM</td>
<td>16.94</td>
<td>13.54</td>
<td>64.78%</td>
<td>54.63%</td>
<td>66.13%</td>
<td>24.13%</td>
</tr>
<tr>
<td>24 h</td>
<td>STA-ResCNN</td>
<td>16.85</td>
<td>13.51</td>
<td>64.82%</td>
<td>54.62%</td>
<td>66.21%</td>
<td>24.15%</td>
</tr>
<tr>
<td/>
<td>EAT-GCN</td>
<td>16.75</td>
<td>13.48</td>
<td>65.12%</td>
<td>55.51%</td>
<td>67.05%</td>
<td>23.67%</td>
</tr>
<tr>
<td/>
<td>WGAT-AGRU</td>
<td><bold>16.61</bold></td>
<td><bold>13.47</bold></td>
<td><bold>66.71%</bold></td>
<td><bold>56.01%</bold></td>
<td><bold>67.26%</bold></td>
<td><bold>23.34%</bold></td>
</tr>
<tr>
<td></td>
<td>MLP</td>
<td>20.10</td>
<td>16.22</td>
<td>53.11%</td>
<td>48.52%</td>
<td>60.68%</td>
<td>29.24%</td>
</tr>
<tr>
<td>Prediction time step &#x003D; </td>
<td>GRU</td>
<td>19.04</td>
<td>15.26</td>
<td>57.80%</td>
<td>49.78%</td>
<td>61.12%</td>
<td>27.15%</td>
</tr>
<tr>
<td>36 h</td>
<td>GC-LSTM</td>
<td>18.70</td>
<td>14.95</td>
<td>59.28%</td>
<td>50.23%</td>
<td>61.47%</td>
<td>26.70%</td>
</tr>
<tr>
<td/>
<td>WGAT-AGRU</td>
<td><bold>18.32</bold></td>
<td><bold>14.63</bold></td>
<td><bold>61.11%</bold></td>
<td><bold>53.12%</bold></td>
<td><bold>64.02%</bold></td>
<td><bold>26.28%</bold></td>
</tr>
<tr>
<td></td>
<td>MLP</td>
<td>21.53</td>
<td>17.38</td>
<td>48.42%</td>
<td>45.68%</td>
<td>58.79%</td>
<td>32.79%</td>
</tr>
<tr>
<td>Prediction time step &#x003D; </td>
<td>GRU</td>
<td>20.30</td>
<td>16.25</td>
<td>53.92%</td>
<td>47.17%</td>
<td>59.06%</td>
<td>29.90%</td>
</tr>
<tr>
<td>48 h</td>
<td>GC-LSTM</td>
<td>19.91</td>
<td>15.89</td>
<td>55.63%</td>
<td>47.70%</td>
<td>59.45%</td>
<td>29.29%</td>
</tr>
<tr>
<td/>
<td>WGAT-AGRU</td>
<td><bold>19.48</bold></td>
<td><bold>15.52</bold></td>
<td><bold>58.42%</bold></td>
<td><bold>49.37%</bold></td>
<td><bold>61.52%</bold></td>
<td><bold>28.71%</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the multi-step PM<sub>2.5</sub> concentration prediction results of the comparison models on the full data set, taking Beijing as an example. In each sub-figure, the continuous purple dash is the real value of PM<sub>2.5</sub> concentration in Beijing over a period of time, and the other color dashes are the multi-step prediction sequences output by the comparison model. And the multi-step prediction results are displayed every four prediction time steps. The overall fitting ability of the multi-step prediction model can be measured by observing how well the multi-step prediction sequence fits the true values. We can see that the PM<sub>2.5</sub> concentration in Beijing is extremely high in the mid-term. It can be seen that the prediction results of GC-LSTM and WGAT-AGRU are good and better than GRU and MLP.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Comparison of performance evaluation metrics between different model models on the full data set (time step of 24 h, taking Beijing as an example)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-3.tif"/>
</fig>
<p><xref ref-type="table" rid="table-3">Table 3</xref> demonstrates the prediction results of each model on the heating season data set, and the experimental setup is the same as on the full data set. WGAT-AGRU achieves optimal results on the heating season dataset with a minimum RMSE of 26.71, a minimum MAE of 21.70, and a maximum <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> of 59.23%, which shows the best prediction accuracy. And WGAT-AGRU also shows the best fitting ability with the maximum CSI of 61.02%, the maximum POD of 75.11%, and the minimum FAR of 23.71%.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of the proposed model and various models on the heating season data set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>RMSE</th>
<th>MAE</th>
<th>R2</th>
<th>CSI</th>
<th>POD</th>
<th>FAR</th>
</tr>
</thead>
<tbody>
<tr>
<td>MLP</td>
<td>30.15</td>
<td>24.87</td>
<td>47.32%</td>
<td>58.86%</td>
<td>72.82%</td>
<td>24.57%</td>
</tr>
<tr>
<td>GRU</td>
<td>27.90</td>
<td>22.81</td>
<td>55.03%</td>
<td>60.06%</td>
<td>74.48%</td>
<td>24.38%</td>
</tr>
<tr>
<td>GC-LSTM</td>
<td>27.36</td>
<td>22.31</td>
<td>57.19%</td>
<td>60.20%</td>
<td>74.40%</td>
<td>24.08%</td>
</tr>
<tr>
<td>WGAT-AGRU</td>
<td><bold>26.71</bold></td>
<td><bold>21.70</bold></td>
<td><bold>59.23%</bold></td>
<td><bold>61.02%</bold></td>
<td><bold>75.11%</bold></td>
<td><bold>23.71%</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows more visually the multi-step PM<sub>2.5</sub> concentration prediction capability of the comparison models on the heating season data set, taking Xi&#x2019;an as an example. The prediction result of the GC-LSTM is similar to the prediction result of the WGAT-AGRU. They both achieve better prediction results than the other two models, GRU and MLP. However, comparing the prediction results of WGAT-AGRU and GC-LSTM, it can be found that the concentration curve predicted by WGAT-AGRU fits the true value of the change better. For example, WGAT-AGRU performs better for the slightly decreasing trend of PM<sub>2.5</sub> concentration at the 10-time step, 30-time step, and 40-time step.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison of performance evaluation metrics between different model models on the heating season data set (time step of 24 h, taking Xi&#x2019;an as an example)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-4.tif"/>
</fig>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Long-term (72 h) predictions of PM<sub>2.5</sub> concentrations from contrasting model models (taking Beijing as an example)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-5.tif"/>
</fig>
<p>To evaluate the long-term prediction performance of the prediction models, experiments were conducted at different prediction time steps, including 24, 36, and 48 h in this paper. The long-term PM<sub>2.5</sub> concentration prediction performance of the models was compared by analyzing the degree of decay of the PM<sub>2.5</sub> concentration prediction accuracy of each model with the increase of the prediction step. The experimental results are shown in <xref ref-type="table" rid="table-2">Table 2</xref>. And we can see that the experimental metrics of the proposed model are optimal at each prediction step.</p>

<p>The prediction results of PM<sub>2.5</sub> concentration for the next 72 h and the RMSE variation with the prediction step from each model with prediction steps are shown in <xref ref-type="fig" rid="fig-4">Figs. 5</xref> and <xref ref-type="fig" rid="fig-6">6</xref>. As we can see that the WGAT-AGRU model still maintains an excellent fitting ability for PM<sub>2.5</sub> concentration prediction in a long prediction time, and can accurately capture the trend of PM<sub>2.5</sub> concentration compared with other prediction models.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparison model RMSE variation with prediction time step (taking Beijing as an example)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_38162-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Validity Analysis</title>
<p>In order to verify the effectiveness of each module of our prediction model, ablation experiments are conducted on the full dataset. The validity of the PM<sub>2.5</sub> directed flow graph structure, spatial correlation feature extraction module, and spatiotemporal fusion prediction module on the model are analyzed in this paper, respectively.</p>
<p>The traditional PM<sub>2.5</sub> spatial correlation graph DG-INPUT constructs an adjacency matrix by determining whether the distance between city vertices exceeds the threshold and whether the edge weight is the reciprocal of the distance between cities. DG-INPUT and the flow graph FG-INPUT proposed in this paper are used as inputs for experiments on the full data set, respectively, with the same prediction time step set to 24 h. The results are shown in <xref ref-type="table" rid="table-4">Table 4</xref> (different graph structure input modules). We can see from the result that FG-INPUT improves the model prediction accuracy with the minimum RMSE and MAE because it simulates the impact of inter-city PM<sub>2.5</sub> flow transmission well.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of prediction performance of different graph structure input modules and spatial correlation feature extraction module</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Methods</th>
<th>RMSE</th>
<th>MAE</th>
<th>R<sup>2</sup></th>
<th>CSI</th>
<th>POD</th>
<th>FAR</th>
</tr>
</thead>
<tbody>
<tr>
<td>Different graph structure input </td>
<td>WGAT-AGRU<break/>(DG-INPUT)</td>
<td>16.73</td>
<td>13.50</td>
<td>64.91%</td>
<td>54.63%</td>
<td>66.13%</td>
<td>24.13%</td>
</tr>
<tr>
<td>module</td>
<td>WGAT-AGRU<break/>(FG-INPUT)</td>
<td><bold>16.61</bold></td>
<td><bold>13.47</bold></td>
<td><bold>66.71%</bold></td>
<td><bold>56.01%</bold></td>
<td><bold>67.26%</bold></td>
<td><bold>23.34%</bold></td>
</tr>
<tr>
<td>Different spatial </td>
<td>FC-AGRU</td>
<td>17.47</td>
<td>14.12</td>
<td>57.55%</td>
<td>57.39%</td>
<td>66.70%</td>
<td>26.25%</td>
</tr>
<tr>
<td>extraction </td>
<td>GAT-AGRU</td>
<td>16.62</td>
<td>13.49</td>
<td>66.21%</td>
<td>55.37%</td>
<td>67.01%</td>
<td>23.41%</td>
</tr>
<tr>
<td>correlation feature </td>
<td>GNN- AGRU</td>
<td>16.63</td>
<td><bold>13.26</bold></td>
<td>66.15%</td>
<td>55.31%</td>
<td>66.65%</td>
<td>23.53%</td>
</tr>
<tr>
<td>module</td>
<td>WGAT-AGRU</td>
<td><bold>16.61</bold></td>
<td>13.47</td>
<td><bold>66.71%</bold></td>
<td><bold>56.01%</bold></td>
<td><bold>67.26%</bold></td>
<td><bold>23.34%</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To verify the effectiveness of the spatial correlation feature extraction module in the model, WGAT is replaced with the fully connected network (FC), and ablation weight aggregation component (GAT), ablation graph attention component (GNN), respectively. And they are combined with AGRU. The experiments on the full data set are conducted and the prediction step size is uniformly set to 24 h. The results are shown in <xref ref-type="table" rid="table-4">Table 4</xref> (different spatial correlation feature extraction modules). We can see that the combination of WGAT can effectively improve the prediction result of the model.</p>

<p>To analyze the effectiveness of the spatiotemporal fusion prediction module, WGAT is combined with the fully connected network (FC), ablation temporal attention component (GRU), and the spatiotemporal fusion prediction module (AGRU) of this paper, respectively. Three prediction steps of 24, 36, and 48 h are set, to analyze the effects of each component above on the multi-step prediction of PM<sub>2.5</sub> concentration. The experimental results are shown in <xref ref-type="table" rid="table-5">Table 5</xref>. The proposed model WGAT-AGRU achieves the best results in most cases.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of spatiotemporal fusion prediction performance</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Methods</th>
<th>RMSE</th>
<th>MAE</th>
<th>R2</th>
<th>CSI</th>
<th>POD</th>
<th>FAR</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td rowspan="3">Prediction time step &#x003D; 24 h</td>
<td>WGAT-FC</td>
<td>17.03</td>
<td>13.69</td>
<td>65.48%</td>
<td>55.65%</td>
<td>69.58%</td>
<td><bold>22.14%</bold></td>
</tr>
<tr>
<td>WGAT-GRU</td>
<td><bold>16.60</bold></td>
<td>13.47</td>
<td>66.56%</td>
<td>55.87%</td>
<td>66.91%</td>
<td>23.56%</td>
</tr>
<tr>
<td>WGAT-AGRU</td>
<td>16.61</td>
<td><bold>13.47</bold></td>
<td><bold>66.71%</bold></td>
<td><bold>56.01%</bold></td>
<td><bold>67.26%</bold></td>
<td>23.34%</td>
</tr>
<tr align="center">
<td rowspan="3">Prediction time step &#x003D; 36 h</td>
<td>WGAT-FC</td>
<td>19.19</td>
<td>15.23</td>
<td>59.93%</td>
<td>52.08%</td>
<td>63.81%</td>
<td>27.82%</td>
</tr>
<tr>
<td>WGAT-GRU</td>
<td>18.13</td>
<td>15.06</td>
<td>60.43%</td>
<td>52.77%</td>
<td>63.87%</td>
<td>27.26%</td>
</tr>
<tr>
<td>WGAT-AGRU</td>
<td><bold>18.13</bold></td>
<td><bold>14.63</bold></td>
<td><bold>61.11%</bold></td>
<td><bold>53.12%</bold></td>
<td><bold>64.02%</bold></td>
<td><bold>26.28%</bold></td>
</tr>
<tr align="center">
<td rowspan="3">Prediction time step &#x003D; 48 h</td>
<td>WGAT-FC</td>
<td>20.85</td>
<td>16.93</td>
<td>57.03%</td>
<td>47.18%</td>
<td>58.21%</td>
<td>33.01%</td>
</tr>
<tr>
<td>WGAT-GRU</td>
<td>19.92</td>
<td>15.73</td>
<td>57.53%</td>
<td>48.21%</td>
<td>59.85%</td>
<td>30.88%</td>
</tr>
<tr>
<td>WGAT-AGRU</td>
<td><bold>19.48</bold></td>
<td><bold>15.52</bold></td>
<td><bold>58.42%</bold></td>
<td><bold>49.37%</bold></td>
<td><bold>61.52%</bold></td>
<td><bold>28.71%</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Discussion</title>
<p>In this paper, the proposed model WGAT-LSTM with fused spatiotemporal features achieves the best result in most cases. It can be seen from <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, WGAT-AGRU performs better on the full data set compared with MLP. The MLP model cannot capture the time-dependent relationship, and it cannot predict such drastic changes, and is a poor fit for the peak PM<sub>2.5</sub> concentration. While the model proposed in this paper can cope with this situation well and predict the concentration change trend more accurately. And WGAT-AGRU can capture the directional PM<sub>2.5</sub> flow through the PM<sub>2.5</sub> directed flow graph. GC-LSTM is a graph convolutional neural network that extracts the spatial features between city vertices in the graph structure and achieves spatiotemporal prediction by combining it with LSTM. Thus the prediction result of GC-LSTM is similar to the prediction result of the WGAT-AGRU model. However, during the heating season, when more coal is burned in northern cities in China and the prevailing north wind makes the inter-city PM<sub>2.5</sub> concentration more affected by wind propagation, WGAT-AGRU achieves better results as shown in <xref ref-type="fig" rid="fig-4">Figs. 4c</xref> and <xref ref-type="fig" rid="fig-4">4d</xref>. It can be inferred that extracts the PM<sub>2.5</sub> spatial flow feature between city vertices in the PM<sub>2.5</sub> flow map through WGAT. By analyzing the multi-step prediction ability of the model, it can be found that WGAT-AGRU has better prediction ability when the model is forecasting for a long time. FC cannot capture long-term time dependencies leading to poorer prediction results. When the predicted step size is 24 h, which is the single step length, our multi-step prediction model does not fully reflect the advantages as shown in <xref ref-type="table" rid="table-5">Table 5</xref> (prediction time step &#x003D; 24 h).</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<sec id="s5_1">
<label>5.1</label>
<title>Conclusion</title>
<p>The existing PM<sub>2.5</sub> concentration prediction methods ignore the directionality of PM<sub>2.5</sub> flow between cities when constructing spatial correlation maps, and the accuracy of multi-step prediction is low. In this paper, a PM<sub>2.5</sub> concentration prediction model WGAT-AGRU integrating spatial and temporal features is constructed and geographic and meteorological data is introduced into the graph structure to construct a PM<sub>2.5</sub> directional flow graph to realize the simulation of inter-city PM<sub>2.5</sub> flow transmission. Prediction can focus on the key time step information, thus improving the accuracy of the PM<sub>2.5</sub> prediction model in multi-step prediction. The comparison experiments with other models on the KnowAir dataset are conducted. The experimental findings indicated that the WGAT-AGRU model is superior to other models in predicting PM<sub>2.5</sub> concentration.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Limitations and Future Work</title>
<p>There are still some limitations in this research. In this paper, the PM<sub>2.5</sub> spatial correlation graph is mainly constructed based on the longitude, latitude, altitude, and wind direction data of the cities, and the edge weight is calculated by a simple diffusion transfer model. However, the spatial flow of PM<sub>2.5</sub> is more complex. In addition, the PM<sub>2.5</sub> concentration prediction model in the paper realizes multi-step prediction of future PM<sub>2.5</sub> concentration through continuous iteration. This recursive multi-step prediction strategy makes the prediction error accumulate with the increase of the prediction time step, so it is impossible to predict the long-term PM<sub>2.5</sub> concentration.</p>
<p>In the future, we will consider introducing more external features to further optimize the construction of the adjacency matrix of the PM<sub>2.5</sub> spatial correlation graph and the calculation of edge weight, to more accurately simulate the flow and transmission of PM<sub>2.5</sub> at the spatial level. In addition, optimize the prediction model structure for the multi-step prediction task of PM<sub>2.5</sub> concentration, and explore the use of a structure such as the Seq2Seq model to solve the multi-step prediction problem.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their gratitude to Central South University for the financial support through Central South University Research Programme of Advanced Interdisciplinary Studies (2023QYJC041).</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by Central South University Research Programme of Advanced Interdisciplinary Studies (2023QYJC041).</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. J.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mago</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Y. G.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y. Y.</given-names> <surname>Chiang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Exploiting spatiotemporal patterns for accurate air quality forecasting using deep learning</article-title>,&#x201D; in <conf-name>Proc. of the 26th ACM SIGSPATIAL Int. Conf. on Advances in Geographic Information Systems</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, pp. <fpage>359</fpage>&#x2013;<lpage>368</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. G.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Shahabi</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Diffusion convolutional recurrent neural network: Data-driven traffic forecasting</article-title>,&#x201D; in <conf-name>Proc. of the 6th Int. Conf. on Learning Representations</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. L.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Karimiana</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>A hybrid model for spatiotemporal forecasting of PM<sub>2.5</sub> based on graph convolutional neural network and long short-term memory</article-title>,&#x201D; <source>Science of the Total Environment</source>, vol. <volume>664</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2019</year>; <pub-id pub-id-type="pmid">30743109</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X. B.</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>N. X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>X. Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>T. L.</given-names> <surname>Su</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Integrated predictor based on decomposition mechanism for PM<sub>2.5</sub> long-term prediction</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>9</volume>, no. <issue>21</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L. Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>R. Z.</given-names> <surname>Qiu</surname></string-name>, <string-name><given-names>X. S.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>H. H.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Trend analysis and forecast of PM<sub>2.5</sub> in Fuzhou, China using the ARIMA model</article-title>,&#x201D; <source>Ecological Indicators</source>, vol. <volume>95</volume>, pp. <fpage>702</fpage>&#x2013;<lpage>710</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Venkataraman</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Usmanulla</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sonnappa</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Sadashiv</surname></string-name>, <string-name><given-names>S. S.</given-names> <surname>Mohammed</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Wavelet and multiple linear regression analysis for identifying factors affecting particulate matter PM<sub>2.5</sub> in Mumbai City, India</article-title>,&#x201D; <source>International Journal of Quality &#x0026; Reliability Management</source>, vol. <volume>36</volume>, no. <issue>10</issue>, pp. <fpage>1750</fpage>&#x2013;<lpage>1783</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Tai</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Mickley</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Jacob</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Leibensperger</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Meteorological modes of variability for fine particulate matter (PM<sub>2.5</sub>) air quality in the United States: Implications for PM<sub>2.5</sub> sensitivity to climate change</article-title>,&#x201D; <source>Atmospheric Chemistry And Physics</source>, vol. <volume>12</volume>, no. <issue>6</issue>, pp. <fpage>3131</fpage>&#x2013;<lpage>3145</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Shamsoddini</surname></string-name>, <string-name><given-names>M. R.</given-names> <surname>Aboodi</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Karami</surname></string-name></person-group>, &#x201C;<article-title>Tehran air pollutants prediction based on random forest feature selection method</article-title>,&#x201D; <source>International Archives of the Photogrammetry, Remote Sensing &#x0026; Spatial Information Sciences</source>, vol. <volume>42</volume>, pp. <fpage>483</fpage>&#x2013;<lpage>488</lpage>, <year>2017</year>; <pub-id pub-id-type="pmid">25642100</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. H.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>An improved model for PM2.5 inference based on support vector machine</article-title>,&#x201D; in <conf-name>2016 17th IEEE/ACIS Int. Conf. on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing (SNPD)</conf-name>, <publisher-loc>Shanghai, China</publisher-loc>, pp. <fpage>27</fpage>&#x2013;<lpage>31</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z. D.</given-names> <surname>Qin</surname></string-name> and <string-name><given-names>G. S.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>A novel hybrid-Garch model based on ARIMA and SVM for PM<sub>2.5</sub> concentrations forecasting</article-title>,&#x201D; <source>Atmospheric Pollution Research</source>, vol. <volume>8</volume>, no. <issue>5</issue>, pp. <fpage>850</fpage>&#x2013;<lpage>860</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Asadollahfardi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Madinejad</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Aria</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Motamadi</surname></string-name></person-group>, &#x201C;<article-title>Predicting particulate matter (PM<sub>2.5</sub>) concentrations in the air of Shahr-e Ray City, Iran, by using an artificial neural network</article-title>,&#x201D; <source>Environmental Quality Management</source>, vol. <volume>25</volume>, no. <issue>4</issue>, pp. <fpage>71</fpage>&#x2013;<lpage>83</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Shen</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Feng</surname></string-name></person-group>, &#x201C;<article-title>Prediction of hourly ground-level PM<sub>2.5</sub> concentrations 3 days in advance using neural networks with satellite data in eastern China</article-title>,&#x201D; <source>Atmospheric Pollution Research</source>, vol. <volume>8</volume>, no. <issue>6</issue>, pp. <fpage>1005</fpage>&#x2013;<lpage>1015</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I. G.</given-names> <surname>McKendry</surname></string-name></person-group>, &#x201C;<article-title>Evaluation of artificial neural networks for fine particulate pollution (PM<sub>10</sub> and PM<sub>2.5</sub>) forecasting</article-title>,&#x201D; <source>Journal of the Air &#x0026; Waste Management Association</source>, vol. <volume>52</volume>, no. <issue>9</issue>, pp. <fpage>1096</fpage>&#x2013;<lpage>1101</lpage>, <year>2002</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B. T.</given-names> <surname>Ong</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Sugiura</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Zettsu</surname></string-name></person-group>, &#x201C;<article-title>Dynamic pre-training of deep recurrent neural networks for predicting environmental monitoring data</article-title>,&#x201D; in <conf-name>2014 IEEE Int. Conf. on Big Data (Big Data)</conf-name>, <publisher-loc>Washington, DC, USA</publisher-loc>, pp. <fpage>760</fpage>&#x2013;<lpage>765</lpage>, <year>2014</year>. </mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. T.</given-names> <surname>Tsai</surname></string-name>, <string-name><given-names>Y. R.</given-names> <surname>Zeng</surname></string-name> and <string-name><given-names>Y. S.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<article-title>Air pollution forecasting using RNN with LSTM</article-title>,&#x201D; in <conf-name>2018 IEEE 16th Int. Conf. on Dependable, Autonomic and Secure Computing, 16th Int. Conf. on Pervasive Intelligence and Computing, 4th Int. Conf. on Big Data Intelligence and Computing and Cyber Science and Technology Congress (DASC/PiCom/DataCom/CyberSciTech)</conf-name>, <publisher-loc>Athens, Greece</publisher-loc>, pp. <fpage>1074</fpage>&#x2013;<lpage>1079</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Krishan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Jha</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Das</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>M. K.</given-names> <surname>Goyal</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Air quality modelling using long short-term memory (LSTM) over NCT-Delhi, India</article-title>,&#x201D; <source>Air Quality, Atmosphere &#x0026; Health</source>, vol. <volume>12</volume>, no. <issue>8</issue>, pp. <fpage>899</fpage>&#x2013;<lpage>908</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Chi</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Huang</surname></string-name> and <string-name><given-names>L. Z.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Dissolved oxygen concentration prediction model based on WT-MIC-GRU&#x2014;a case study in dish-shaped lakes of Poyang Lake</article-title>,&#x201D; <source>Entropy</source>, vol. <volume>24</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>18</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Chakma</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Vizena</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lin</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Image-based air quality analysis using deep convolutional neural network</article-title>,&#x201D; in <conf-name>2017 IEEE Int. Conf. on Image Processing (ICIP)</conf-name>, <publisher-loc>Beijing, China</publisher-loc>, pp. <fpage>3949</fpage>&#x2013;<lpage>3952</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y. H.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>P. F.</given-names> <surname>Du</surname></string-name> and <string-name><given-names>J. Y.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Image-based PM2.5 estimation and its application on depth estimation</article-title>,&#x201D; in <conf-name>2018 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>, <publisher-loc>Calgary, AB, Canada</publisher-loc>, pp. <fpage>1857</fpage>&#x2013;<lpage>1861</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ran</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Prediction of gas concentration evolution with evolutionary attention-based temporal graph convolutional network</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>200</volume>, no. <issue>7</issue>, pp. <fpage>116944</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. F. F.</given-names> <surname>Karney</surname></string-name></person-group>, &#x201C;<article-title>Algorithms for geodesics</article-title>,&#x201D; <source>Journal of Geodesy</source>, vol. <volume>87</volume>, no. <issue>1</issue>, pp. <fpage>43</fpage>&#x2013;<lpage>55</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. F.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X. L.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Th&#x00E9;</surname></string-name>, <string-name><given-names>Z. C.</given-names> <surname>Tan</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Multi-step forecast of PM<sub>2.5</sub> and PM<sub>10</sub> concentrations using convolutional neural network integrated with spatial&#x2013;temporal attention and residual learning</article-title>,&#x201D; <source>Environment International</source>, vol. <volume>171</volume>, pp. <fpage>107691</fpage>, <year>2023</year>; <pub-id pub-id-type="pmid">36516675</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>







