<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">55442</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2024.055442</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Self-Attention Spatio-Temporal Deep Collaborative Network for Robust FDIA Detection in Smart Grids</article-title>
<alt-title alt-title-type="left-running-head">Self-Attention Spatio-Temporal Deep Collaborative Network for Robust FDIA Detection in Smart Grids</alt-title>
<alt-title alt-title-type="right-running-head">Self-Attention Spatio-Temporal Deep Collaborative Network for Robust FDIA Detection in Smart Grids</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Zu</surname><given-names>Tong</given-names></name>
</contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Fengyong</given-names></name>
<email>fyli@shiep.edu.cn</email></contrib>
<aff>
<institution>College of Computer Science and Technology, Shanghai University of Electric Power</institution>, <addr-line>Shanghai, 201306</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Fengyong Li. Email: <email>fyli@shiep.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>27</day><month>9</month><year>2024</year></pub-date>
<volume>141</volume>
<issue>2</issue>
<fpage>1395</fpage>
<lpage>1417</lpage>
<history>
<date date-type="received">
<day>27</day>
<month>6</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>8</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_55442.pdf"></self-uri>
<abstract>
<p>False data injection attack (FDIA) can affect the state estimation of the power grid by tampering with the measured value of the power grid data, and then destroying the stable operation of the smart grid. Existing work usually trains a detection model by fusing the data-driven features from diverse power data streams. Data-driven features, however, cannot effectively capture the differences between noisy data and attack samples. As a result, slight noise disturbances in the power grid may cause a large number of false detections for FDIA attacks. To address this problem, this paper designs a deep collaborative self-attention network to achieve robust FDIA detection, in which the spatio-temporal features of cascaded FDIA attacks are fully integrated. Firstly, a high-order Chebyshev polynomials-based graph convolution module is designed to effectively aggregate the spatio information between grid nodes, and the spatial self-attention mechanism is involved to dynamically assign attention weights to each node, which guides the network to pay more attention to the node information that is conducive to FDIA detection. Furthermore, the bi-directional Long Short-Term Memory (LSTM) network is introduced to conduct time series modeling and long-term dependence analysis for power grid data and utilizes the temporal self-attention mechanism to describe the time correlation of data and assign different weights to different time steps. Our designed deep collaborative network can effectively mine subtle perturbations from spatiotemporal feature information, efficiently distinguish power grid noise from FDIA attacks, and adapt to diverse attack intensities. Extensive experiments demonstrate that our method can obtain an efficient detection performance over actual load data from New York Independent System Operator (NYISO) in IEEE 14, IEEE 39, and IEEE 118 bus systems, and outperforms state-of-the-art FDIA detection schemes in terms of detection accuracy and robustness.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>False data injection attacks</kwd>
<kwd>smart grid</kwd>
<kwd>deep learning</kwd>
<kwd>self-attention mechanism</kwd>
<kwd>spatio-temporal fusion</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Guangxi Key Lab of Multi-Source Information Mining</funding-source>
<award-id>MIMS21-M-02</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the advancement of smart sensor and wireless communication technologies, traditional power systems are gradually transitioning to intelligent grid cyber-physical systems to enhance energy utilization efficiency and system stability [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. The advantages of smart grids are becoming increasingly apparent, as the cyber-physical systems in smart grids offer end-to-end bi-directional power flow, which allows the users to feedback energy into the grid, thus improving system scalability, efficiency, and stability [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. However, the openness and diversity of cyber-physical systems make the smart grid more susceptible to diverse malicious network attacks [<xref ref-type="bibr" rid="ref-5">5</xref>], especially at the boundary of interaction between the public Internet and the power system private network. For example, denial of service attack [<xref ref-type="bibr" rid="ref-6">6</xref>], man-in-the-middle attack [<xref ref-type="bibr" rid="ref-7">7</xref>], network topology attack [<xref ref-type="bibr" rid="ref-8">8</xref>], channel measurement attack [<xref ref-type="bibr" rid="ref-9">9</xref>] and false data injection attack [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>False Data Injection Attacks (FDIAs) have garnered increasing attention due to their destructive and covert nature [<xref ref-type="bibr" rid="ref-11">11</xref>]. Generally, in the cyber-physical systems of smart grids, crucial operations such as emergency analysis, Volt-VAR optimization, real-time pricing, etc., are always conducted using power system state estimators. Sensor measurements are utilized as inputs to these estimators, which then generate corresponding outputs including a range of voltage magnitudes and phase angles. Control signals are dispatched by Supervisory Control and Data Acquisition (SCADA) systems based on the output of state estimation, which is crucial for the efficient operation of smart grids [<xref ref-type="bibr" rid="ref-12">12</xref>]. Although the cyber-physical system can mitigate the information interference of the public Internet by deploying Bad Data Detection (BDD) algorithms in the state estimator, FDIAs can circumvent these checks by injecting carefully crafted attack vectors, thereby tampering with measurement data from SCADA systems and impacting the accuracy of state estimation. This could lead to severe consequences for grid computation, scheduling operations, and the overall stability of the system [<xref ref-type="bibr" rid="ref-13">13</xref>]. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents a complete schematic diagram of an FDIA attack on a cyber-physical system.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Typical scenarios of FDIA in Smart Grid</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-1.tif"/>
</fig>
<p>In general, FDIA attack detection techniques mainly contain traditional machine learning model-based methods and data-driven methods [<xref ref-type="bibr" rid="ref-14">14</xref>]. Model-based detection methods rely on system modeling to compare with expected behavior to detect attacks. For example, Wei et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed a method using Forecast-Aided State Estimation (FASE) and Square Root Unscented Kalman Filter (SR-UKF) for FDIA detection, which involved Generalized Likelihood Ratio Test (GLRT). Qu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] suggested an FDIA detection approach that utilizes the Hellinger distance to track measurement value changes and determine attack presence. Shen et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] introduced a detection method based on random matrix theory, utilizing random variables from load short-term forecasts to construct random matrices for FDIA detection.</p>
<p>Data-driven detection methods usually leverage large amounts of historical data and employ statistical analysis or machine learning techniques to find unusual patterns to detect potential attack behavior. These methods require no prior knowledge and can adapt to various complex attack scenarios. For instance, James et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] utilized discrete wavelet transform (DWT) and deep neural networks (DNN) to analyze system states continuously over time, effectively capturing FDIAs. Habibi et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed a neural network based on time series analysis and the Nonlinear Autoregressive Exogenous model (NARX) for detecting network attacks using estimated errors. Lu et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] introduced a convolutional neural network called representation learning CNN (RL-CNN) to capture local data features, exhibiting superior performance in locating network attacks as a multi-label classifier. With the gradual increase of time series data in the smart grid, the detection effect of FDIAs can be improved more effectively by using the measurement of time series dependence in time series data. For example, Wang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] introduced a two-level learner-based scheme for detecting FDIAs, integrating linear and nonlinear time series data from the power grid and employing a combination of Kalman filter and Recurrent Neural Network (KFRNN). To effectively capture long-term dependencies in the data. Ayad et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a deep learning (DL) method based on Long Short-Term Memory (LSTM) framework to detect FDIAs. However, the above methods focus on the time characteristics of the measured data, but ignore the topology of the grid, resulting in the loss of many key information in the learning process. Furthermore, the researchers try to explore spatial correlations in grid topology to further address the FDIA detection problem. Boyaci et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an FDIA identification method based on Graph Neural Network (GNN) using the graph topology of the power system and the spatial correlation of the measurement data. In order to better extract the spatial characteristics of power grid topology information and operation data, Li et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] designed a detection method of FDIAs based on Gated Graph Neural Network (GGNN) to improve the detection accuracy under the change of power grid topology. Su et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed a Dual-Attention Multi-head Graph Attention Network (DAMGAT) for FDIAs detection to improve the interpretability of graph neural network model and the representation ability of power topology nodes. Considering the spatio-temporal dependence of power grid data, Zhang et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] analyzed temporal correlation and spatial correlation by volumetric Kalman filter and Gaussian process regression to capture the dynamic characteristics of the state vector to evaluate and localize FDIAs. Han et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] designed a multi-graph mechanism and a time correlation layer to mine the correlation features of power data, and constructed a graph topology for the detection of FDIAs.</p>
<p>Overall, existing FDIAs detection methods can already utilize data and power grid topology to achieve efficient attack detection. However, they rarely consider the spatio-temporal dependencies of power grid data, which limits the detection performance of FDIAs. <italic>First</italic>, in terms of spatial correlation, most methods only consider the influence between nodes connected by the power grid topology, ignoring the potential impacts from other nodes, significantly weakening the understanding of the overall correlation of the power grid topology. <italic>Second</italic>, in terms of temporal correlation, current works often overlook the long-term correlation of sequences and the correlation of data features with different time steps. <italic>Third</italic>, most methods hardly consider the subtle perturbations in spatiotemporal feature information, making it difficult for existing methods to distinguish between power grid noise and FDIA attacks, resulting in lower robustness of detection models.</p>
<p>Facing the aforementioned problems, we are thus motivated to design an efficient deep collaborative self-attention network in the context of robust FDIAs detection, which makes the following novel contributions:
<list list-type="bullet">
<list-item>
<p>We design a deep collaborative self-attention network to achieve effective robust FDIA detection. Our proposed collaborative network model can effectively capture subtle perturbations from spatiotemporal feature information, efficiently distinguish power grid noise from FDIA attacks, and adapt to diverse attack intensities.</p></list-item>
<list-item>
<p>We design a graph convolution module based on Chebyshev polynomials, which utilizes the characteristics of the graph convolution network to aggregate the node information in the power grid and introduces the spatial self-attention mechanism to adjust the degree of attention to different nodes. The proposed module can better adapt to different grid structures and characteristics, thus improving the robustness of the model to potential changes and anomalies in grid data.</p></list-item>
<list-item>
<p>Bi-directional LSTM network with a self-attention mechanism is introduced to conduct time series modeling and long-term dependence analysis on power grid data and utilizes the temporal self-attention module to describe the time correlation of data and assign different weights to different time steps.</p></list-item>
<list-item>
<p>Comprehensive experiments are performed over three standard datasets and demonstrate that our method can outperform state-of-the-art FDIA detection schemes in terms of detection accuracy and robustness.</p></list-item>
</list></p>
<p>The rest of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews the preliminary state estimation of smart gird and false data injection attacks. In <xref ref-type="sec" rid="s3">Section 3</xref>, we propose a robust FDIA detection scheme by designing an efficient deep collaborative self-attention network. Extensive experiments are performed to evaluate the overall performance of the proposed scheme, and the corresponding results and discussions are presented in <xref ref-type="sec" rid="s4">Section 4</xref>. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Preliminary</title>
<sec id="s2_1">
<label>2.1</label>
<title>State Estimation and Bad Data Detection</title>
<p>In general, the control center estimates the state through the measurement information in the monitoring and data acquisition system to ensure the safety and stability of the power grid [<xref ref-type="bibr" rid="ref-15">15</xref>]. State estimation primarily uses the redundancy of measurement data to estimate the operating state of the grid, including bus voltage, transmission line power flow, and bus power [<xref ref-type="bibr" rid="ref-24">24</xref>]. In energy management systems, state estimation can enable the functions of power flow calculation and load forecasting and the DC power model is usually used to ensure the convergence of the state estimation. The unit voltage of each node in the system is assumed to be 1, and the effect of line resistance and ground branch is ignored, the active power between bus <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>a</mml:mi></mml:math></inline-formula> and bus <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>b</mml:mi></mml:math></inline-formula> in the test system can be accordingly represented by the following model:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>e</mml:mi><mml:mtext>&#xA0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> represent voltage amplitude and voltage phase of bus <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>a</mml:mi></mml:math></inline-formula> and bus <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>b</mml:mi></mml:math></inline-formula>, respectively. <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the branch impedance and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>e</mml:mi></mml:math></inline-formula> is the measurement error. The active power injection on busbar <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>a</mml:mi></mml:math></inline-formula> can be expressed as the sum of the active power flow of each adjacent branch, which can be calculated as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> represents active injection of bus <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>a</mml:mi></mml:math></inline-formula> and <italic>N</italic> is the set of branches adjacent to bus <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>a</mml:mi></mml:math></inline-formula>. Accordingly, the generalization can be expressed as:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>z</mml:mi></mml:math></inline-formula> is the measurement data obtained in the SCADA system. <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:mi mathvariant="bold">H</mml:mi></mml:mrow></mml:math></inline-formula> is the Jacobian matrix. <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>x</mml:mi></mml:math></inline-formula> is the system state vector. <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>e</mml:mi></mml:math></inline-formula> is the measurement error vector. The covariance matrix is diagonal when the measurement errors are assumed to be uncorrelated. Correspondingly, according to the principle of residual and least squares, the minimization objective function <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> can be calculated as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mi>x</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:math></inline-formula> is the weight matrix of direction-finding quantity, and the state estimator variable <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>x</mml:mi></mml:math></inline-formula> can be obtained from the objective function as:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mi>z</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In order to grasp the real-time operation status of the smart power grid, the SCADA system uses intelligent terminals and other devices to collect measurement data [<xref ref-type="bibr" rid="ref-25">25</xref>]. Since these measurements are easily affected by traditional power system faults, such as equipment aging, communication failure, and noise interference, the SCADA system introduces the bad data detection (BDD) module to identify and eliminate such independent and accidental natural faults, where the constructed residual vector <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>r</mml:mi></mml:math></inline-formula> in BDD module is the difference between the real direction-finding vector <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:mover><mml:mi>z</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and the estimated theoretical vector, which can be expressed as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>z</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the process of FDIA bad data detection, the Euclidean norm of the residual <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>r</mml:mi></mml:math></inline-formula> is first compared with the threshold <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>. If <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>r</mml:mi><mml:mo>&gt;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, the presence of bad data is determined and further processed by identification and correction operations. The aforementioned bad data detection process is repeated until all residual vectors meet the pre-determined criteria.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>False Data Injection Attack</title>
<p>During FDIAs, false data vector <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> is usually injected into the vector measurement data <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, which inevitably cause the input vector <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>z</mml:mi><mml:mo>+</mml:mo><mml:mi>a</mml:mi></mml:math></inline-formula> of the state estimate to be biased from the real case [<xref ref-type="bibr" rid="ref-28">28</xref>]. Accordingly, the state variable <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>x</mml:mi></mml:math></inline-formula> may be also offset, resulting in <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>c</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> is the deviation of the state variable before and after the attack. If the attacker can invade the configuration information of the power system and obtain the Jacobian matrix <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mrow><mml:mi mathvariant="bold">H</mml:mi></mml:mrow></mml:math></inline-formula> from the system, the attacker could construct a false data attack vector <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mi>c</mml:mi></mml:math></inline-formula> so that bad data detection module cannot recognize it. Correspondingly, the attacked measurement vector <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and residual <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> can be represented as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mi>a</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>+</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow></mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>r</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Meanwhile, since the bad data detection module may fail to detect FDIAs, the control center is likely to make wrong decision instructions according to the estimated state <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>x</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-29">29</xref>]. It is worth noting that the attack effect of the FDIA method described above is related to the attacker&#x2019;s mastery of the grid information. When the attacker can use more power grid topology information and electrical parameters, it is entirely possible for them to construct a Jacobian matrix <bold>H</bold> that is closer to the topology information of the power grid, and the constructed false data vectors are very similar to the weak fluctuation noise of the power grid. Correspondingly, the bad data detection module is highly likely to identify these FDIA attacks as general noise from the power grid, allowing them to easily bypass the detection model for effective attacks.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Method</title>
<sec id="s3_1">
<label>3.1</label>
<title>Overview of Proposed Detection Scheme</title>
<p>Our designed deep collaborative network mainly consists of two network branches. One branch is a graph convolution network based on Chebyshev polynomials, which aggregates the node information in the power grid, and adaptively adjusts the degree of attention to different nodes by introducing the spatio self-attention mechanism. Another branch is the bi-directional LSTM network, which can conduct time series modeling and long-term dependence analysis for power grid data describe the time correlation of data, and assign different weights to different time steps by introducing the temporal self-attention mechanism. The overall detection process can be described as follows. The operational data of sensors is first collected through the SCADA system, including the bus injection active power <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and reactive power <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, branch power flow <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, to form the grid measurement data <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>Z</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. Subsequently, after obtaining the power grid information, the attacker performs an FDIA attack to evade the BDD module. Furthermore, the operational data that is not detected by the BDD module is then fed into the deep collaborative network. Our deep collaborative network extracts temporal and spatial self-attention features through graph convolutional networks and bidirectional LSTM networks, respectively. Finally, spatio-temporal self-attention features are fully integrated by the deep collaborative network to make a final judgment decision. The overview of the proposed network architecture is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The overall architecture of proposed robust FDIA detection scheme</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>High-Order Chebyshev Graph Convolution Network</title>
<p>Graph convolution network (GCN) is usually used in node classification, graph classification, link prediction, and other tasks [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. In general, a traditional graph convolution network only calculates the result of Chebyshev graph convolution at <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>. Nevertheless, since the <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>k</mml:mi></mml:math></inline-formula>-order convolution operator of Chebyshev graph convolution can cover the <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>-order adjacent nodes, the general GCN can only extract the spatial correlation of the first-order adjacent nodes while ignoring the effective spatial information of high-order adjacent nodes. In order to better represent the topology characteristics of the power grid, we re-construct graph convolution network by introducing <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>k</mml:mi></mml:math></inline-formula>-order Chebyshev polynomial (<inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>k</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>), which can effectively capture the spatial relationship between higher-order adjacent nodes, and realize the information dissemination and feature aggregation between nodes through the adjacency matrix. Correspondingly, the connection mode and interaction between nodes in the power grid data can be better understood.</p>
<p>To be specific, we model the power grid topology as an undirected graph <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mrow><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mrow><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> represents the set of <italic>N</italic> nodes in the power grid, and <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mrow><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is the set of branches connected between nodes. Through the connection relationship between nodes in the power grid, the adjacency matrix <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> of the graph can be obtained. If two bus nodes are directly connected, the corresponding element of two nodes in matrix <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow></mml:math></inline-formula> is 1; that is, <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, otherwise <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. The complete procedure for using Chebyshev graph convolution to update node information can be defined as follows:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>x</mml:mi></mml:math></inline-formula> is the input data. <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>k</mml:mi></mml:math></inline-formula> is the order of Chebyshev polynomial. <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the parameter matrix of <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>k</mml:mi></mml:math></inline-formula>-order Chebyshev polynomial. <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is a Chebyshev polynomial function, which can be defined as follows:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>I</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>2</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>I</italic> is the identity matrix and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:math></inline-formula> stands for the symmetric normalized Laplace matrix, which can maintain the structural information, sparsity and positive definiteness of graph, and can be expressed as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula> is the largest characteristic of Laplacian matrix and <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> is the normalized Laplace matrix, which can be calculated as:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the degree matrix recording the degree of each node. Since Laplacian matrix is a positive semi-definite matrix and eigenvalues are greater than or equal to 0, each eigenvalue of <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mfrac><mml:mrow><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> is accordingly within [0, 1], <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> eigenvalues is within [0, 2], and the eigenvalue of <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> is within [&#x2212;1, 1]. This result can ensure that the range of independent variables of Chebyshev polynomials is limited to [&#x2212;1, 1] interval, and guarantee the stability and convergence of numerical calculation in graph convolution network, thereby adapting to the properties of Chebyshev polynomials.</p>
<p>Although high-order Chebyshev GCN can effectively propagate information and control the propagation range, all nodes in each order polynomial only share one parameter, which makes it impossible to adaptively allocate neighbor node weights according to node differences. Therefore, we introduce the spatio self-attention mechanism to further improve its feature representation capability. The spatio self-attention mechanism can dynamically adjust the attention weight between nodes. Accordingly, the network model can accurately learn the importance of each node relationship and then pay more attention to the nodes that are crucial to FDIA detection task. The attention scoring for node <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>i</mml:mi></mml:math></inline-formula> and its adjacent node <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>j</mml:mi></mml:math></inline-formula> is performed through dot multiplication, and the formula is as follows:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msup><mml:mi>K</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>q</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>v</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> are the queries, keys, and values obtained from the linear transformation of node feature <italic>X</italic>, respectively. <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>q</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>v</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> are learnable parameter matrices. <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> stands for the feature dimension used to prevent large dot multiplication results. <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mrow><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:math></inline-formula> is the auxiliary transformation matrix to facilitate the calculation of dimension matching.</p>
<p>Furthermore, the attention scores of all adjacent nodes of node i are normalized by using the softmax function to obtain the attention coefficient <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. Finally, the value matrix <italic>V</italic> can be weighted by the calculated node attention coefficient <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> to obtain the updated feature representation of each node.
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>When the network aggregates the node information according to <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>k</mml:mi></mml:math></inline-formula>-order Chebyshev polynomials, we can multiply the elements of the attention weighting matrix and the polynomial <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to adjust the weight distribution. And then focus on those vulnerable nodes:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>k</mml:mi><mml:mi>y</mml:mi><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Finally, after all node information is aggregated, each node can dynamically adjust the feature representation according to its influence in the power grid to better reflect its role in the whole detection model.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>BiLSTM Network with Temporal Self-Attention Mechanism</title>
<p>Bidirectional Long Short-Term Memory (BiLSTM) is a variant of recurrent neural networks (RNNs). Compared to traditional unidirectional LSTMs, BiLSTM can simultaneously consider both past and future information in a sequence, thus better capturing the long-term dependencies and contextual information of sequences [<xref ref-type="bibr" rid="ref-32">32</xref>]. A BiLSTM consists of two opposing LSTM layers, combining forward and backward information flow to enhance the capture of sequence data features [<xref ref-type="bibr" rid="ref-33">33</xref>]. In forward propagation, the BiLSTM can process the entire time sequence to capture past-to-present temporal information.
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mover><mml:mrow><mml:mrow><mml:mrow><mml:mtext>LSTM</mml:mtext></mml:mrow></mml:mrow></mml:mrow><mml:mo>&#x2192;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>while in backward propagation, it can process from the time sequence end and mainly focus on future-to-present temporal information. BiLSTM allows each time point to access contextual information, comprehensively understanding temporal dependencies.
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mover><mml:mrow><mml:mrow><mml:mrow><mml:mtext>LSTM</mml:mtext></mml:mrow></mml:mrow></mml:mrow><mml:mo>&#x2190;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the input of the current sample. <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:math></inline-formula> represent the backward and forward cell states of the current sample. <italic>T</italic> is the length of the input time series. After processing the entire sequence, BiLSTM can ensure the output contains important features from both the beginning and end of the sequence by merging the hidden states from forward and backward propagation, which can significantly enhance the understanding and predictive accuracy of time series data.
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo></mml:mover></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> are the weight ratios of forward and backward hidden states. In general, a complete LSTM unit is composed of input gate, forgetting gate, output gate, and cell state [<xref ref-type="bibr" rid="ref-32">32</xref>]. The input gate that controlled by the sigmoid function can regulate the impact of the current input on the cell state. The forget gate determines the extent to which the cell state from the previous time step is forgotten, while the output gate regulates the influence of the cell state on the hidden state. The cell state is responsible for transmitting and storing long-term dependency information, and gate mechanisms control the flow and storage of information, which enable LSTM to capture long-term dependencies in sequences while maintaining gradient stability.</p>
<p>In view of the dynamic and complexity of the time series data involved in the FDIAs detection task, we further introduce the temporal self-attention mechanism to enhance the performance and robustness of the BiLSTM model. Our goal is to improve the model&#x2019;s perception of the dynamic characteristics of time series data, so as to more accurately find the possible abnormal data, and further improve the model&#x2019;s robustness. Specifically:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>B</mml:mi><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is temporal self-attention weight matrix. <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msubsup><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the forward and backward hidden state extracted by the current BiLSTM, and <italic>B</italic> is an auxiliary transformation matrix. Then, we combine the attention weight matrix of BiLSTM for weighted output to update the representation of temporal features. Through the temporal self-attention mechanism, the weight of the power grid time series data at different time points can be efficiently allocated, which helps in processing the long series data more effectively, taking into account the context dependency of the time series and the importance of different time points. Notably, a dropout layer is added at the end of the network to prevent its overfitting and improve the generalization capability of the model.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Spatio-Temporal Feature Fusion</title>
<p>Considering the spatial correlation and temporal dependence, the features can be further fused to improve the representation capability of the features by combining self-attention and temporal self-attention features. To be specific, we use the pre-processed measurement data as the input <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> of the network model. In the processing of spatial correlation, the topology of the power grid can be firstly obtained to construct the corresponding adjacency matrix <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mrow><mml:mi mathvariant="bold">A</mml:mi></mml:mrow></mml:math></inline-formula>, and the corresponding Laplace matrix <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is sequentially calculated. Then, we perform Chebyshev graph convolution by setting appropriate hyperparameters to effectively aggregate the information of adjacent nodes.
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mtext>CGCN</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where LeakyReLU is nonlinear activation function, <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is parameter matrix, and <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the deviation. After obtaining the spatial characteristics, we further consider the temporal characteristics of the grid measurement data.
<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mtext>BiLSTM</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Finally, the spatio correlation information and temporal dependence information can be effectively fused to improve the representation capability of the features.
<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mi>T</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are two weighted matrices. After performing spatio-temporal feature fusion, the fused features can be fed to the full connection layer, and then output the results through the activation function to determine the normal and abnormal probability of the grid measurement data.
<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>u</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are weighted matrices, <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the deviation, and <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is sigmoid activation function. By combining the high-order Chebyshev graph convolution network and BiLSTM network with temporal self-attention mechanism, an efficient deep collaborative network can be integrated, which is called as CGCN-BiLSTM(Chebyshev Graph Convolution-Bidirectional Long Short-Term Memory). The complete network architecture is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Structure diagram of a deep collaborative network with spatiotemporal self-attention mechanism</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-3.tif"/>
</fig>
<p>In addition, it is necessary to construct the calculation error and back-propagation error of the loss function to guide the updating of network parameters during the training process of the neural network. In our network model, the cross entropy loss function is used to calculate the loss of the model, and the cosine annealing learning rate scheduler is also involved to dynamically adjust the learning rate and optimize the training process of the deep learning model.
<disp-formula id="eqn-27"><label>(27)</label><mml:math id="mml-eqn-27" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>N</italic> is the input sequence length, <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> stands for the real label (&#x2013;1 or 1) of the <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>i</mml:mi></mml:math></inline-formula>-th sample, and <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the prediction probability that the model belongs to the positive class for the <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>i</mml:mi></mml:math></inline-formula>-th sample.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Robust FDIA Detection Procedure</title>
<p>Based on the above-mentioned the spatiotemporal feature fusion model, we can build the detailed FDIA detection procedure, which can be described detailedly as follows, e.g., <bold>Algorithm 1</bold>, which can be described detailedly as follows:
<list list-type="bullet">
<list-item>
<p><bold>Step 1:</bold> We conduct power flow calculations on the data collected from the power grid to obtain measurement data, which is then pre-processed to get input features <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item>
<p><bold>Step 2:</bold> We further compute the adjacency matrix <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and Laplacian matrix <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> corresponding to the bus system, and then normalize the standardized <italic>X</italic>, <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mrow><mml:mover><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> as the inputs to the CGCN-BiLSTM model.</p></list-item>
<list-item>
<p><bold>Step 3:</bold> The information of various nodes in the power grid is sequentially aggregated by using CGCN to get the feature representation <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>. Subsequently, spatial information can be further integrated to obtain feature representation <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> by combining spatial self-attention <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> with Chebyshev polynomials, <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and the Laplacian matrix.</p></list-item>
<list-item>
<p><bold>Step 4:</bold> Input feature <italic>X</italic> into bidirectional LSTM units to capture its long-term dependencies. Then, the forward and backward hidden states <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mrow><mml:msubsup><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> can be generated. Subsequently, the state feature <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is further updated by using temporal attention and obtaining the temporal feature representation <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>T</mml:mi><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> through activation functions.</p></list-item>
<list-item>
<p><bold>Step 5:</bold> Finally, the spatio-temporal features are fully fused to output the model decision <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mrow><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> through fully connected layers. Notably, the parameters can be gradually optimized through the loss function.</p></list-item>
</list></p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results and Discussions</title>
<p>In this section, we first introduced the experimental setup and evaluation metrics in <xref ref-type="sec" rid="s4_1">Sections 4.1</xref> and <xref ref-type="sec" rid="s4_2">4.2</xref>, respectively. Then, a series of comparisons and discussions aiming at overall detection performance were sequentially performed in <xref ref-type="sec" rid="s4_3">Section 4.3</xref>. Furthermore, we discussed the robustness between our scheme and several state-of-the-art schemes in <xref ref-type="sec" rid="s4_4">Section 4.4</xref>. Last but not least, the computation complexity of different detection schemes was tested in <xref ref-type="sec" rid="s4_5">Section 4.5</xref>.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<p>In our experiment, the load information of different regions of NYISO was utilized as the basic power gird data to generate normal measurement data. We simulated the measured data <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mrow><mml:mi>&#x1D4B5;</mml:mi></mml:mrow></mml:math></inline-formula> of the power system by power flow calculation based on this load information. Matpower toolkit was used to obtain relevant information, such as Jacobian matrix <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mrow><mml:mi mathvariant="bold">H</mml:mi></mml:mrow></mml:math></inline-formula> and power grid topology, and FDIA sample sets were constructed to maximize the avoidance of BDD detection.</p>
<p>In order to ensure the authenticity of the experiment, we added the standard Gaussian distribution noise with a variance of 0.01 and mean value of 0, which was applied to different datasets, IEEE 14 bus system, IEEE 39 bus system, and IEEE 118 bus system, to generate experimental data with varying degrees of noise interference. Our goal is to verify that our detection model can still efficiently and accurately identify FDIA attacks under different levels of noise interference. In addition, for each bus system, we generated 15,000 groups of measurement data, including 7500 groups of FDIA data and 7500 groups of normal data. We labeled FDIA data as 1 and normal data as &#x2212;1 to facilitate model detection. Correspondingly, all data samples were divided the data set into <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mn>75</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> training set and <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mn>25</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> test set for each experiment.</p>
<p>Moreover, in the process of model training, our network model can gradually update the weight parameters by calculating the loss value and gradient. In order to speed up the convergence speed of model training and prevent gradient explosion, we normalized the input data by the <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:math></inline-formula> function, and then adopted the exit strategy to temporarily stop the work of some neurons with a certain probability <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi>p</mml:mi></mml:math></inline-formula> to avoid the over-fitting of the model. All experiments were performed over Pytorch 1.7 framework with the Intel Core i9-12900hx CPU, 16GB RAM and NVIDIA GeForce RTX 4060 GPU. The power flow calculation and state estimation of power grid data were simulated using Matpower toolkit on MATLAB.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Evaluation Metrics</title>
<p>To provide sufficient performance comparison in terms of detection performance, we introduced four evaluation metrics, i.e., accuracy, precision, recall and <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score, to give the experimental results. In general, the definition of statistical variables involved in the evaluation indicators were true negative (TN), false positive (FP), true positive (TP), and false negative (TN), respectively, which are as follows: 1) TN represents the number of correctly identified normal measurement data of the power grid as normal data, denoted as <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. 2) FP refers to the number of errors in identifying normal measurement data of the power grid as FDIA, denoted as <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. 3) TP stands for the number of FDIA data correctly identified as FDIA in the power grid, denoted as <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. 4) TN means the number of times FDIA data in the power grid is incorrectly identified as normal data, denoted as <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. Correspondingly, the calculation expressions for the four evaluation indicators are given as follows:
<disp-formula id="eqn-28"><label>(28)</label><mml:math id="mml-eqn-28" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi></mml:math></inline-formula> is the ratio of all correctly judged FDIA samples. The higher the value of <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi></mml:math></inline-formula>, the better the overall performance of the detection model.
<disp-formula id="eqn-29"><label>(29)</label><mml:math id="mml-eqn-29" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> is the ratio of real FDIA samples in the predicted FDIA samples. The larger the value of <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula>, the lower the false alarm rate of the detection model and the better the detection effect.
<disp-formula id="eqn-30"><label>(30)</label><mml:math id="mml-eqn-30" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:math></inline-formula> is the ratio of correctly predicted FDIA samples in real FDIA samples. The higher the recall rate, the lower the missed detection rate of the detection model.
<disp-formula id="eqn-31"><label>(31)</label><mml:math id="mml-eqn-31" display="block"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the harmonic average of precision and recall. The higher the <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> score, the better the overall performance of the detection model.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparison with the State of the Arts</title>
<p>In this section, we conducted a series of experiments to compare the proposed scheme with existing state-of-the-art FDIA detection schemes, CNN [<xref ref-type="bibr" rid="ref-20">20</xref>], LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>], GCN [<xref ref-type="bibr" rid="ref-24">24</xref>], CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>] and DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]. In order to verify the effectiveness and reliability of the proposed FDIAs detection method, four different evaluation metrics, Accuracy, Precision, Recall, <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score, were sequentially tested. All experiments were carried out over three standard data sets, IEEE 14 bus system, IEEE 39 bus system, and IEEE 118 bus system, to provide the experimental results. In each experiment, the input batch of all models was set to 64, and the number of training iterations was 80. The initial learning rate was set to 0.01 and adjusted using the Adam optimizer. The parameter probability of the dropout layer was initially set to 0.4. The order of Chebyshev polynomial of the proposed model is set to 3, the number of layers of LSTM is 2, and the number of layers of self-attention is 2.</p>
<p>The corresponding experimental results are shown in <xref ref-type="table" rid="table-1">Table 1</xref>. As can be observed from this table, our scheme can obtain the best overall detection performance compared to other state-of-the-art schemes, no matter which dataset is used. To be specific, on IEEE 14 bus system, compared with the current best scheme DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>], our scheme can obtain a significant performance gain with 3.25<inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for accuracy, 4.37<inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for precision, 1.81<inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for recall, and 3.02<inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score. For IEEE 39 bus system, our solution can achieve approximate improvements with 3.60<inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for accuracy, 3.67<inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for precision, 3.25<inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for recall, and 3.47<inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score. Similarly, for the larger scale data from the IEEE 118 bus system, our scheme can still achieve an approximate performance advantage, that is, the performance gain with 3.14<inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for accuracy, 2.92<inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for precision, 3.24<inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for recall, and 3.08<inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score. In addition, we conducted the experimental comparison of FPR (False Positive Rate) on IEEE 14, 39 and 118 bus systems, respectively, and the results are shown in <xref ref-type="table" rid="table-2">Table 2</xref>. It can be seen from the data in the table that the FPR value of our model is significantly lower than that of other advanced models, indicating that the proposed method performs better in distinguishing between true and false samples.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Overall performance comparison for six detection methods, CNN, LSTM, GCN, CNN-LSTM, DAMGAT, CGCN-BiLSTM. Four evaluation metrics, accuracy, precision, recall, <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score, were used to provide the results. All experiments were performed on IEEE 14, IEEE 39 and IEEE 118 bus systems</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="4">IEEE 14 bus system</th>
<th align="center" colspan="4">IEEE 39 bus system</th>
<th align="center" colspan="4">IEEE 118 bus system</th>
</tr>
<tr>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula></th>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula></th>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.8151</td>
<td>0.8522</td>
<td>0.7651</td>
<td>0.7963</td>
<td>0.8109</td>
<td>0.8170</td>
<td>0.7360</td>
<td>0.7634</td>
<td>0.8030</td>
<td>0.8069</td>
<td>0.7260</td>
<td>0.7526</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.8662</td>
<td>0.8331</td>
<td>0.8864</td>
<td>0.8721</td>
<td>0.8553</td>
<td>0.8354</td>
<td>0.8845</td>
<td>0.8593</td>
<td>0.8457</td>
<td>0.9015</td>
<td>0.7757</td>
<td>0.8339</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.9157</td>
<td>0.9299</td>
<td>0.9039</td>
<td>0.9167</td>
<td>0.9067</td>
<td>0.8754</td>
<td>0.9539</td>
<td>0.9130</td>
<td>0.8883</td>
<td>0.8855</td>
<td>0.8987</td>
<td>0.8920</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.9223</td>
<td>0.9062</td>
<td>0.9419</td>
<td>0.9237</td>
<td>0.9110</td>
<td>0.9168</td>
<td>0.9091</td>
<td>0.9129</td>
<td>0.8983</td>
<td>0.8792</td>
<td>0.9232</td>
<td>0.9007</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.9493</td>
<td>0.9460</td>
<td>0.9558</td>
<td>0.9509</td>
<td>0.9360</td>
<td>0.9282</td>
<td>0.9487</td>
<td>0.9383</td>
<td>0.9263</td>
<td>0.9347</td>
<td>0.9208</td>
<td>0.9277</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.9818</bold></td>
<td><bold>0.9897</bold></td>
<td><bold>0.9739</bold></td>
<td><bold>0.9811</bold></td>
<td><bold>0.9720</bold></td>
<td><bold>0.9649</bold></td>
<td><bold>0.9812</bold></td>
<td><bold>0.9730</bold></td>
<td><bold>0.9577</bold></td>
<td><bold>0.9639</bold></td>
<td><bold>0.9532</bold></td>
<td><bold>0.9585</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the maximum in current column.</p></fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Overall performance comparison for six detection methods, CNN, LSTM, GCN, CNN-LSTM, DAMGAT, CGCN-BiLSTM. Evaluation metrics false positive rate (FPR) was used to provide the results. All experiments were performed on IEEE 14, IEEE 39 and IEEE 118 bus systems</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left" />
<col align="left" />
<col align="left" />
<col align="left" />
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="3">Different datasets</th>
</tr>
<tr>
<th>IEEE 14 bus system</th>
<th>IEEE 39 bus system</th>
<th>IEEE 118 bus system</th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.1986</td>
<td>0.2028</td>
<td>0.2364</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.1412</td>
<td>0.1571</td>
<td>0.1849</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.0970</td>
<td>0.1069</td>
<td>0.1269</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.0850</td>
<td>0.0976</td>
<td>0.1158</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.0591</td>
<td>0.0731</td>
<td>0.0812</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.0193</bold></td>
<td><bold>0.0314</bold></td>
<td><bold>0.0489</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the minimum in current column.</p></fn></table-wrap-foot>
</table-wrap>
<p>The above experimental results demonstrated that our scheme can achieve significant performance improvements on both small-scale and large-scale datasets. In fact, this phenomenon can be easily explained by the following two reasons. Firstly, our proposed deep collaborative network combined the characteristics of the graph convolution network and Bi-LSTM to optimize the network structure, which can better adapt to different power grid structures and effectively capture the difference between the slight changes and actual FDIA attacks in the power grid, thus improving the robustness of the model to potential FDIA attacks. Secondly, the self-attention mechanism was introduced in spatiotemporal feature construction, which can further enhance the representation capability of spatiotemporal features, thereby efficiently guiding detection features to pay more attention to the differences between FDIA attack samples and normal data, resulting in a significant improvement of detection accuracy.</p>
<p>Furthermore, we can observe an interesting phenomenon from <xref ref-type="table" rid="table-1">Table 1</xref>, that is, the overall detection performance on the IEEE 118 bus system is slightly lower than that on the IEEE 39 and IEEE 14 bus systems. To be specific, for accuracy measurement, the average reductions for IEEE 118 bus system were approximately 2.41<inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for IEEE 14 bus system, 1.43<inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for IEEE 39 bus system. For <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score, the average reductions for IEEE 118 bus system were approximately 2.26<inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for IEEE 14 bus system, 1.45<inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for IEEE 39 bus system. This is mainly because if a bus system has a larger scale and more complex topology, it means that more system nodes and more complex power grid structures need to be processed in the FDIA attack detection process. This is inevitably more susceptible to external interference on the topology of the power grid, which can introduce more complex noise interference and lead to attack misjudgment by the detection model. Correspondingly, our proposed network model can deeply analyze and understand the behavior and interaction of each power node, which can efficiently process the spatial-temporal complexity of the power grid and make the detection model more sensitive to identify and respond to abnormal data in various power grid operations, thereby enhancing the adaptability and detection accuracy in the changing complex power grid environment.</p>

</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Robustness Analysis</title>
<p>In order to gain more insight, we further verified the applicability and robustness of the proposed detection scheme in the real power grid environment. In our experiments, a series of comparative experiments, including different attack intensities, different noise environments and different node attacks, were conducted over standard datasets from three bus systems, IEEE 14, IEEE 39 and IEEE 118 bus systems.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Attack Intensity</title>
<p>Attack intensity refers to the degree of deviation between the false data injected by the attacker and the actual measured data. A higher attack intensity means that there is a large deviation between the injected false data and the actual data, while a lower attack strength means that there is only a slight difference. Intuitively, FDIA attackers always hope that the attack data can bypass the detection model as much as possible while completing an effective attack, which forces the detection model to adapt to attack levels of different densities as much as possible. To give a more valid observation, we divided the generated FDIAs samples into the strong attack, medium attack, and weak attack [<xref ref-type="bibr" rid="ref-35">35</xref>] and observed the robustness of our proposed detection model under different levels of attack intensity, where strong attack is that the ratio of the average injection power deviation to the actual measurement is greater than 30<inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, medium attack is that the ratio ranges between 10<inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and 30<inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, while weak attack is that the ratio is less than 10<inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>.</p>
<p><xref ref-type="fig" rid="fig-4">Figs. 4</xref>&#x2013;<xref ref-type="fig" rid="fig-6">6</xref> present the experimental comparison of different attack intensities on IEEE 14, IEEE 39 and IEEE 118 bus systems, respectively. The abscissa in these figures shows the accuracy rate, accuracy rate, recall rate and <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score index, while the ordinate is the overall detection performance of each metric. From the results, we can easily observe that the proposed scheme can almost always obtain superior detection performance values compared with most of the existing schemes, no matter which data set from IEEE bus system is used. To be specific, for a high-intensity attack on IEEE 118 bus system, the proposed scheme can obtain <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score about 3.23<inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> higher than DAMGAT scheme [<xref ref-type="bibr" rid="ref-25">25</xref>] and <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score about 10.33<inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> higher than CNN-LSTM scheme [<xref ref-type="bibr" rid="ref-34">34</xref>]. Even if for low-intensity attack on IEEE 118 bus system, our scheme still showed an obvious advantage compared with other schemes, e.g., the <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score gains of our scheme were 2.95<inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for DAMGAT scheme, 12.09<inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for CNN-LSTM scheme, 16.05<inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> for GCN scheme, respectively. We believe that this mainly benefits from our proposed deep collaborative network and the introduction of a spatio-temporal self-attention mechanism. On the one hand, the deep collaborative network can better adapt to different power grid structures and effectively capture the difference between the slight changes and actual FDIA attacks in the power grid. On the other hand, the introduction of self-attention mechanism further enhances the representation capability of spatio-temporal features, resulting in the improvement of detection performance.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Performance comparison with four evaluation metrics for different attack intensities on IEEE 14 bus system. Three types of intensity attacks, Low (9<inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), Medium (20<inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), High (35<inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), were tested in each experiment</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Performance comparison with four evaluation metrics for different attack intensities on IEEE 39 bus system. Three types of intensity attacks, Low (9<inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), Medium (20<inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), High (35<inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), were tested in each experiment</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Performance comparison with four evaluation metrics for different attack intensities on IEEE 118 bus system. Three types of intensity attacks, Low (9<inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), Medium (20<inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), High (35<inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), were tested in each experiment</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Noise Interference</title>
<p>To give more insight, we implemented a series of experiments to further verify the robustness of the proposed method against measurement noise interference. Standard Gaussian noise with different variances were introduced into power grid measurement to simulate the real power grid environment, and performed specific experiments over IEEE 14, IEEE 39 and IEEE 118 bus systems.</p>
<p>The experimental results were shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, where the abscissa represents the noise variance, while the ordinate is <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> score, which is used to measure the detection performance. Moreover, we compared the accuracy rate with other advanced detection models at different noise levels, and the experimental results are shown in <xref ref-type="table" rid="table-3">Table 3</xref>, where the noise error is between 0.5<inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and 2<inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. As can be seen from the figure and table, our scheme always outperformed other state-of-the-art detection methods in various noise levels, demonstrating the superior detection capability of our scheme. In addition, we can also observe that with the noise level increasing, the overall detection performance of all schemes decreases. This is because high noise levels may obscure the characteristics of attack data, making the detection model difficult to distinguish between attack and normal data. However, our scheme exhibited the smallest performance decreasing under identical noise interference, suggesting that the model&#x2019;s self-attention mechanism dynamically adjusts its focus on different parts of the input data, effectively reducing the interference of noise on the decision-making process.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Performance comparison under different noise environments. All experiments were performed over IEEE 14, IEEE 39 and IEEE 118 bus systems and standard Gaussian noises were used to simulate the real power grid environment</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_55442-fig-7.tif"/>
</fig><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of accuracy of six detection methods: CNN, LSTM, GCN, CNN-LSTM, DAMGAT and CGCN-BiLSTM under different ambient noise. All experiments were performed on IEEE 14, IEEE 39, and IEEE 118 bus systems</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="4">IEEE 14 bus system</th>
<th align="center" colspan="4">IEEE 39 bus system</th>
<th align="center" colspan="4">IEEE 118 bus system</th>
</tr>
<tr>
<th>0.5<inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.0<inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.5<inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>2.0<inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>0.5<inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.0<inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.5<inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>2.0<inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>0.5<inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.0<inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>1.5<inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
<th>2.0<inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.8198</td>
<td>0.7963</td>
<td>0.7789</td>
<td>0.7549</td>
<td>0.8067</td>
<td>0.7738</td>
<td>0.7598</td>
<td>0.7249</td>
<td>0.7876</td>
<td>0.7526</td>
<td>0.7398</td>
<td>0.7032</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.9022</td>
<td>0.8711</td>
<td>0.8672</td>
<td>0.8385</td>
<td>0.8731</td>
<td>0.8583</td>
<td>0.8322</td>
<td>0.8175</td>
<td>0.8621</td>
<td>0.8349</td>
<td>0.8142</td>
<td>0.8005</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.9277</td>
<td>0.9167</td>
<td>0.9051</td>
<td>0.8902</td>
<td>0.9177</td>
<td>0.9130</td>
<td>0.9070</td>
<td>0.8852</td>
<td>0.9018</td>
<td>0.8931</td>
<td>0.8779</td>
<td>0.8516</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.9329</td>
<td>0.9196</td>
<td>0.9041</td>
<td>0.9012</td>
<td>0.9244</td>
<td>0.9128</td>
<td>0.9071</td>
<td>0.8879</td>
<td>0.9055</td>
<td>0.8962</td>
<td>0.8813</td>
<td>0.8702</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.9651</td>
<td>0.9509</td>
<td>0.9419</td>
<td>0.9324</td>
<td>0.9427</td>
<td>0.9308</td>
<td>0.9211</td>
<td>0.9094</td>
<td>0.9298</td>
<td>0.9197</td>
<td>0.9019</td>
<td>0.8902</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.9897</bold></td>
<td><bold>0.9812</bold></td>
<td><bold>0.9782</bold></td>
<td><bold>0.9709</bold></td>
<td><bold>0.9769</bold></td>
<td><bold>0.9719</bold></td>
<td><bold>0.9694</bold></td>
<td><bold>0.9563</bold></td>
<td><bold>0.9639</bold></td>
<td><bold>0.9587</bold></td>
<td><bold>0.9502</bold></td>
<td><bold>0.9411</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the maximum in current column.</p></fn></table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_4_3">
<label>4.4.3</label>
<title>Node Number Discussions</title>
<p>To accurately simulate real-world conditions in power grids, we further study the comparative experiments on single-node and multi-node attacks. In general, for FDIA attackers, single-node attacks can significantly reduce attack costs when the resources are limited. However, in some specific scenarios, the attackers may simultaneously implement FDIA attacks on multiple nodes to achieve maximum attack effectiveness.</p>
<p>We tested the detailed comparative experiments with different numbers of attack nodes on the IEEE 14, IEEE 39, and IEEE 118 test systems. The corresponding experimental results were shown in <xref ref-type="table" rid="table-4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="table-6">6</xref>, respectively. From these tables, the proposed method is significantly superior to the existing several detection methods on three bus systems, whether the single node or multiple nodes scenario was tested. Furthermore, the experimental results indicated that in single-node attack scenarios, the attack data are often limited to a very small number of data points, greatly increasing the difficulty of identifying FDIAs within a vast dataset of normal operations. In contrast, for multi-node attacks, due to strong inter-node relationships and the wide distribution of data, the model can obtain a richer set of data points for analysis, thereby generally achieving higher detection performance than single-node attacks. This observation further validates the effectiveness and high adaptability of the proposed scheme in complex attack scenarios.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance comparison with four metrics for single-node FDIAs detection and multi-node FDIAs detection on IEEE 14 bus system</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="4">Single-node detection</th>
<th align="center" colspan="4">Multi-node detection</th>
</tr>
<tr>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.7820</td>
<td>0.8074</td>
<td>0.7728</td>
<td>0.7755</td>
<td>0.8175</td>
<td>0.8412</td>
<td>0.7863</td>
<td>0.7925</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.8539</td>
<td>0.8564</td>
<td>0.9067</td>
<td>0.8699</td>
<td>0.8710</td>
<td>0.8330</td>
<td>0.9364</td>
<td>0.8817</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.9097</td>
<td>0.8953</td>
<td>0.9331</td>
<td>0.9138</td>
<td>0.9173</td>
<td>0.9017</td>
<td>0.9416</td>
<td>0.9212</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.9093</td>
<td>0.8783</td>
<td>0.9558</td>
<td>0.9154</td>
<td>0.9210</td>
<td>0.8941</td>
<td>0.9597</td>
<td>0.9258</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.9417</td>
<td>0.9756</td>
<td>0.9091</td>
<td>0.9412</td>
<td>0.9510</td>
<td>0.9462</td>
<td>0.9591</td>
<td>0.9526</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.9733</bold></td>
<td><bold>0.9662</bold></td>
<td><bold>0.9825</bold></td>
<td><bold>0.9742</bold></td>
<td><bold>0.9871</bold></td>
<td><bold>0.9945</bold></td>
<td><bold>0.9796</bold></td>
<td><bold>0.9865</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the maximum in current column.</p></fn></table-wrap-foot>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance comparison with four metrics for single-node FDIAs detection and multi-node FDIAs detection on IEEE 39 bus system</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="4">Single-node detection</th>
<th align="center" colspan="4">Multi-node detection</th>
</tr>
<tr>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.7743</td>
<td>0.8058</td>
<td>0.7313</td>
<td>0.7405</td>
<td>0.8127</td>
<td>0.8363</td>
<td>0.7813</td>
<td>0.7869</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.8381</td>
<td>0.8054</td>
<td>0.8889</td>
<td>0.8397</td>
<td>0.8575</td>
<td>0.8638</td>
<td>0.9010</td>
<td>0.8715</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.8909</td>
<td>0.8969</td>
<td>0.9052</td>
<td>0.8957</td>
<td>0.9093</td>
<td>0.8783</td>
<td>0.9558</td>
<td>0.9154</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.8959</td>
<td>0.8995</td>
<td>0.9094</td>
<td>0.8998</td>
<td>0.9173</td>
<td>0.9017</td>
<td>0.9416</td>
<td>0.9212</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.9143</td>
<td>0.9160</td>
<td>0.9105</td>
<td>0.9116</td>
<td>0.9280</td>
<td>0.9066</td>
<td>0.9584</td>
<td>0.9318</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.9649</bold></td>
<td><bold>0.9627</bold></td>
<td><bold>0.9681</bold></td>
<td><bold>0.9643</bold></td>
<td><bold>0.9727</bold></td>
<td><bold>0.9754</bold></td>
<td><bold>0.9708</bold></td>
<td><bold>0.9725</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the maximum in current column.</p></fn></table-wrap-foot>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Performance comparison with four metrics for single-node FDIAs detection and multi-node FDIAs detection on IEEE 118 bus system</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="4">Single-node detection</th>
<th align="center" colspan="4">Multi-node detection</th>
</tr>
<tr>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
<th>Acc</th>
<th>Pre</th>
<th>Rec</th>
<th><inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.7634</td>
<td>0.7916</td>
<td>0.7535</td>
<td>0.7561</td>
<td>0.8016</td>
<td>0.8309</td>
<td>0.7858</td>
<td>0.7952</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>0.8274</td>
<td>0.8501</td>
<td>0.7982</td>
<td>0.8045</td>
<td>0.8456</td>
<td>0.8190</td>
<td>0.8828</td>
<td>0.8447</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.8664</td>
<td>0.8730</td>
<td>0.9018</td>
<td>0.8779</td>
<td>0.8903</td>
<td>0.8373</td>
<td>0.9760</td>
<td>0.9013</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>0.8864</td>
<td>0.8912</td>
<td>0.9059</td>
<td>0.8926</td>
<td>0.9057</td>
<td>0.9392</td>
<td>0.8727</td>
<td>0.9047</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>0.8895</td>
<td>0.8956</td>
<td>0.9049</td>
<td>0.8948</td>
<td>0.9137</td>
<td>0.9378</td>
<td>0.8909</td>
<td>0.9138</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td><bold>0.9367</bold></td>
<td><bold>0.9470</bold></td>
<td><bold>0.9235</bold></td>
<td><bold>0.9339</bold></td>
<td><bold>0.9587</bold></td>
<td><bold>0.9527</bold></td>
<td><bold>0.9675</bold></td>
<td><bold>0.9601</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: The significance of bold value shows the maximum in current column.</p></fn></table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Computational Complexity Analysis</title>
<p>To further demonstrate the superiority of our proposed scheme, we performed a detailed analysis of the complexity and compared the running time of different FDIA detection schemes. To estimate the complexity of each model, we compared the training times of different detection schemes. The running time of five existing state-of-the-art FDIA detection schemes, CNN [<xref ref-type="bibr" rid="ref-20">20</xref>], LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>], GCN [<xref ref-type="bibr" rid="ref-24">24</xref>], CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>] and DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>], and our proposed scheme. The experiments were performed under identical environmental conditions to ensure a fair comparison. The training time per epoch was the average of ten epochs, while the total training time was the average of three complete model training. In addition, on IEEE 14 bus system, the total maximum execution time of the model is 158.12 s; On the larger IEEE 39 bus system, this time is extended to 491.49 s; On the most complex IEEE 118 bus system, the total maximum execution time reached 1472.16 s. In the self-attention mechanism we introduce, we use linear multiplication, the time complexity is mainly with the input node, its time complexity is maintained in <inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and the space complexity is maintained in the range of constant change, that is <inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. In addition, the computational overhead of the bi-directional LSTM network is mainly related to the input sequence length T. Its consumption mainly increases linearly with T, so the time complexity of the bi-directional LSTM network can be simplified to <inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and the space complexity to <inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. In order to verify the feasibility of the method, we carry out experiments to detect the delay and throughput. The average detection delay of IEEE 14 bus system is 0.0034 s, and the throughput is 8245.95 samples per second. On the IEEE 39 bus system, the average detection delay is 0.0106 s and the throughput is 2621.37 samples per second. The IEEE 118 bus system has an average detection delay of 0.0312 s and a throughput of 897.83 samples per second (taking the average of the results of 20 iterations each). Different datasets were tested over the computer platform with Intel Core i9-12900hx CPU, 16 GB RAM and NVIDIA GeForce RTX 4060 GPU.</p>
<p>The corresponding experimental results were presented in <xref ref-type="table" rid="table-7">Table 7</xref>. As can be seen from this table the average single training and total training time of our scheme were significantly lower than that of DAMGAT scheme, whichever IEEE bus system was used. Specifically, our scheme can reduce by 0.74 and 58.35 s for IEEE 14 bus system, 1.83 and 156.04 s for IEEE 39 bus system, and 8.49 and 623.72 s for IEEE 118 bus system. This is because the two-layer multi-head attention mechanism adopted by DAMGAT increases the computational complexity and leads to a long training time. Our proposed spatio-temporal self-attention mechanism simplifies the attention calculation and is more suitable for detection tasks, which significantly shortens the training time. In addition, compared to some traditional detection schemes, e.g., CNN, LSTM, our scheme got an obviously longer average single training and total training times. This is because the internal structure of traditional detection models is simpler, and their corresponding training process is thus faster. However, their feature extraction and representation capability is also correspondingly limited, which inevitably lowers their detection performance.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Average training time for six different FDIA detection schemes. The experiments were performed over IEEE 14, IEEE 39 and IEEE 118 bus systems, and average single training time (s) (S-training in short) and average total training time (s) (T-training in short) were discussed in each test</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left" />
<col align="left" />
<col align="left" />
<col align="left" />
<col align="left" />
<col align="left" />
<col align="left" />
</colgroup>
<thead>
<tr>
<th rowspan="2">Methods</th>
<th align="center" colspan="2">IEEE 14 bus system</th>
<th align="center" colspan="2">IEEE 39 bus system</th>
<th align="center" colspan="2">IEEE 118 bus system</th>
</tr>
<tr>
<th>S-training (s)</th>
<th>T-training (s)</th>
<th>S-training (s)</th>
<th>T-training (s)</th>
<th>S-training (s)</th>
<th>T-training</th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>0.96</td>
<td>78.15</td>
<td>2.72</td>
<td>255.85</td>
<td>9.17</td>
<td>767.43</td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>1.08</td>
<td>87.44</td>
<td>3.14</td>
<td>259.69</td>
<td>10.38</td>
<td>842.04</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>1.31</td>
<td>107.81</td>
<td>4.03</td>
<td>328.92</td>
<td>12.78</td>
<td>1051.14</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>1.77</td>
<td>144.65</td>
<td>4.97</td>
<td>420.93</td>
<td>16.54</td>
<td>1362.60</td>
</tr>
<tr>
<td>DAMGAT [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>2.56</td>
<td>206.89</td>
<td>7.71</td>
<td>620.67</td>
<td>25.90</td>
<td>2062.69</td>
</tr>
<tr>
<td>CGCN-BiLSTM</td>
<td>1.82</td>
<td>148.54</td>
<td>5.88</td>
<td>464.63</td>
<td>17.41</td>
<td>1438.97</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>This paper proposed a deep collaborative self-attention network to achieve effective and robust FDIA detection. The proposed network designed a high-order Chebyshev polynomials-based graph convolution module to aggregate the node information in the power grid and introduced spatial self-attention mechanism to adjust the degree of attention given to different nodes. Furthermore, a bidirectional LSTM network with a self-attention mechanism was introduced to conduct time series modeling and long-term dependence analysis and assign different weights to different time steps. The proposed network model can effectively capture subtle perturbations from spatio-temporal feature information, efficiently achieving robust FDIA detection, and adapting to diverse attack intensities. Extensive experiments demonstrated that the proposed method outperformed existing state-of-the-art FDIA detection schemes in terms of detection accuracy and robustness.</p>
<p>While our scheme can improve the efficiency and robustness of FDIA detection, it should be noted that the training of the proposed network model may be more complex and time-consuming, especially for power grid topology with dynamic changes. In terms of future work, we aim to refine our new scheme in two ways. First, we intend to investigate FDIA detection method applied to dynamic topology changes in the power grid, which may be more practical for new power systems. Second, we intend to explore the lightweight of the deep collaborative network model by optimizing the self-attention structure. These two issues are left for our future work.</p>
</sec>
</body>
<back>
<ack><p>The authors would like to thank anonymous reviewers for their valuable suggestions which helped to improve this article.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported in part by the Research Fund of Guangxi Key Lab of Multi-Source Information Mining &#x0026; Security (MIMS21-M-02).</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: data collection: Tong Zu; analysis and interpretation of results: Tong Zu; draft manuscript preparation: Tong Zu; study conception and design: Fengyong Li; supervision and revision: Fengyong Li. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>All data used or analyzed during this study are included in this article and its references.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ghiasi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Niknam</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mehrandezh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dehghani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ghadimi</surname> <given-names>N</given-names></string-name></person-group>. <article-title>A comprehensive review of cyber-attacks and defense mechanisms for improving security in smart grid energy systems: past, present and future</article-title>. <source>Elect Power Syst Res</source>. <year>2023</year>;<volume>215</volume>:<fpage>108975</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.epsr.2022.108975</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Islam</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Bin Ameedeen</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ajra</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ismail</surname> <given-names>ZB</given-names></string-name>, <string-name><surname>Zain</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Blockchain-enabled cybersecurity provision for scalable heterogeneous network: a comprehensive survey</article-title>. <source>Comput Modeling Eng Sci</source>. <year>2024</year>;<volume>138</volume>(<issue>1</issue>):<fpage>43</fpage>&#x2013;<lpage>123</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2023.028687</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Reda</surname> <given-names>HT</given-names></string-name>, <string-name><surname>Anwar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Tari</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A taxonomy of cyber defence strategies against false data attacks in smart grids</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>14s</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3592797</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Ghiasi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Niknam</surname> <given-names>T</given-names></string-name>, <string-name><surname>Dehghani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ansari</surname> <given-names>HR</given-names></string-name></person-group>. <chapter-title>Cyber-physical security in smart power systems from a resilience perspective: concepts and possible solutions</chapter-title>. In: <source>Power systems cybersecurity: methods, concepts, and best practices</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>, <year>2023</year>. p. <fpage>67</fpage>&#x2013;<lpage>89</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>An</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Data integrity attack in dynamic state estimation of smart grid: attack model and countermeasures</article-title>. <source>IEEE Trans Autom Sci Eng</source>. <year>2022</year>;<volume>19</volume>(<issue>3</issue>):<fpage>1631</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TASE.2022.3149764</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elsisi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Altius</surname> <given-names>M</given-names></string-name>, <string-name><surname>Su</surname> <given-names>S-F</given-names></string-name>, <string-name><surname>Su</surname> <given-names>C-L</given-names></string-name></person-group>. <article-title>Robust kalman filter for position estimation of automated guided vehicles under cyberattacks</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TIM.2023.3250285</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>ZL</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CS-LeCT: chained secure and low-energy consumption data transmission based on compressive sensing</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TIM.2023.3280495</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Coordinated topology attacks in smart grid using deep reinforcement learning</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2020</year>;<volume>17</volume>(<issue>2</issue>):<fpage>1407</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2020.2994977</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Selvam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Tyagi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Residue number system (RNS) and power distribution network topology-based mitigation of power side-channel attacks</article-title>. <source>Cryptography</source>. <year>2023</year>;<volume>8</volume>(<issue>1</issue>):<fpage>1</fpage>. doi:<pub-id pub-id-type="doi">10.3390/cryptography8010001</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ning</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Improved stealthy false data injection attacks in networked control systems</article-title>. <source>IEEE Syst J</source>. <year>2024</year>;<volume>18</volume>(<issue>1</issue>):<fpage>505</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSYST.2024.3350179</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>Z-H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>G-P</given-names></string-name></person-group>. <article-title>Event-based optimal stealthy false data-injection attacks against remote state estimation systems</article-title>. <source>IEEE Trans Cybern</source>. <year>2023</year>;<volume>53</volume>(<issue>10</issue>):<fpage>6714</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCYB.2023.3255583</pub-id>; <pub-id pub-id-type="pmid">37030790</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhattar</surname> <given-names>PL</given-names></string-name>, <string-name><surname>Pindoriya</surname> <given-names>NM</given-names></string-name></person-group>. <article-title>False data injection attack with max-min optimization in smart grid</article-title>. <source>Comput Secur</source>. <year>2024</year>;<volume>140</volume>:<fpage>103761</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2024.103761</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Habib</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Hasan</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Alkhayyat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alkwai</surname> <given-names>LM</given-names></string-name></person-group>. <article-title>False data injection attack in smart grid cyber physical system: issues, challenges, and future direction</article-title>. <source>Comput Electr Eng</source>. <year>2023</year>;<volume>107</volume>:<fpage>108638</fpage>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Detection of false data injection attack in power grid based on spatial-temporal transformer network</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>238</volume>:<fpage>121706</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.121706</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A false data injection attack detection strategy for unbalanced distribution networks state estimation</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2023</year>;<volume>14</volume>(<issue>5</issue>):<fpage>3992</fpage>&#x2013;<lpage>4006</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Georgievitch</surname> <given-names>PM</given-names></string-name></person-group>. <article-title>Detection of false data injection attack in power system based on hellinger distance</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2023</year>;<volume>20</volume>(<issue>2</issue>):<fpage>2119</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2023.3286895</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Detection, differentiation and localization of replay attack and false data injection attack based on random matrix</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>2758</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-52954-z</pub-id>; <pub-id pub-id-type="pmid">38307898</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>James</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>VO</given-names></string-name></person-group>. <article-title>Online false data injection attack detection with wavelet transform and deep neural networks</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2018</year>;<volume>14</volume>(<issue>7</issue>):<fpage>3271</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2018.2825243</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Habibi</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Baghaee</surname> <given-names>HR</given-names></string-name>, <string-name><surname>Dragi&#x010D;evi&#x0107;</surname> <given-names>T</given-names></string-name>, <string-name><surname>Blaabjerg</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Detection of false data injection cyber-attacks in DC microgrids based on recurrent neural networks</article-title>. <source>IEEE J Emerg Sel Top Power Electron</source>. <year>2020</year>;<volume>9</volume>(<issue>5</issue>):<fpage>5294</fpage>&#x2013;<lpage>310</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JESTPE.2020.2968243</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>K-D</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z-G</given-names></string-name></person-group>. <article-title>Representation-learning-based CNN for intelligent attack localization and recovery of cyber-physical power systems</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2023</year>;<volume>35</volume>(<issue>5</issue>):<fpage>6145</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2023.3257225</pub-id>; <pub-id pub-id-type="pmid">37030822</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>KFRNN: an effective false data injection attack detection in smart grid based on kalman filter and recurrent neural network</article-title>. <source>IEEE Internet Things J</source>. <year>2021</year>;<volume>9</volume>(<issue>9</issue>):<fpage>6893</fpage>&#x2013;<lpage>904</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2021.3113900</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ayad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khalaf</surname> <given-names>M</given-names></string-name>, <string-name><surname>Salama</surname> <given-names>M</given-names></string-name>, <string-name><surname>El-Saadany</surname> <given-names>EF</given-names></string-name></person-group>. <article-title>Mitigation of false data injection attacks on automatic generation control considering nonlinearities</article-title>. <source>Elect Power Syst Res</source>. <year>2022</year>;<volume>209</volume>(<issue>6</issue>):<fpage>107958</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.epsr.2022.107958</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boyaci</surname> <given-names>O</given-names></string-name>, <string-name><surname>Narimani</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Davis</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Ismail</surname> <given-names>M</given-names></string-name>, <string-name><surname>Overbye</surname> <given-names>TJ</given-names></string-name>, <string-name><surname>Serpedin</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Joint detection and localization of stealth false data injection attacks in smart grids using graph neural networks</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2021</year>;<volume>13</volume>(<issue>1</issue>):<fpage>807</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2021.3117977</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Graph-based detection for false data injection attacks in power grid</article-title>. <source>Energy</source>. <year>2023</year>;<volume>263</volume>(<issue>5</issue>):<fpage>125865</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.energy.2022.125865</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>X</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Damgat based interpretable detection of false data injection attacks in smart grids</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2024</year>;<volume>15</volume>(<issue>4</issue>):<fpage>4182</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2024.3364665</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bamisile</surname> <given-names>O</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Spatio-temporal correlation-based false data injection attack detection using deep convolutional neural network</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2021</year>;<volume>13</volume>(<issue>1</issue>):<fpage>750</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2021.3109628</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>False data injection attacks detection with modified temporal multi-graph convolutional network in smart grids</article-title>. <source>Comput Secur</source>. <year>2023</year>;<volume>124</volume>:<fpage>103016</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2022.103016</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Si</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Detection of false data injection attacks in cyber-physical power systems: an adaptive adversarial dual autoencoder with graph representation learning approach</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2024</year>;<volume>73</volume>:<fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TIM.2023.3331398</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Bi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Su</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Sparse adversarial learning for FDIA attack sample generation in distributed smart srids</article-title>. <source>Comput Model Eng Sci</source>. <year>2024</year>;<volume>139</volume>(<issue>2</issue>):<fpage>2095</fpage>&#x2013;<lpage>115</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2023.044431</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Heterogeneous network embedding: a survey</article-title>. <source>Comput Model Eng Sci</source>. <year>2023</year>;<volume>137</volume>(<issue>1</issue>):<fpage>83</fpage>&#x2013;<lpage>130</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2023.024781</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An improved dynamic Chebyshev graph convolution network for traffic flow prediction with spatial-temporal attention</article-title>. <source>Appl Intell</source>. <year>2022</year>;<volume>52</volume>(<issue>14</issue>):<fpage>16104</fpage>&#x2013;<lpage>116</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-021-03022-w</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Quantitative characterization of shale gas reservoir properties based on BiLSTM with attention mechanism</article-title>. <source>Geosci Front</source>. <year>2023</year>;<volume>14</volume>(<issue>4</issue>):<fpage>101567</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.gsf.2023.101567</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rahim</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Farid</surname> <given-names>FA</given-names></string-name>, <string-name><surname>SalehMusaMiah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Puza</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Alam</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An enhanced hybrid model based on CNN and BiLSTM for identifying individuals via handwriting analysis</article-title>. <source>Comput Model Eng Sci</source>. <year>2024</year>;<volume>140</volume>(<issue>2</issue>):<fpage>1689</fpage>&#x2013;<lpage>710</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2024.048714</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Fault diagnosis of hydro-turbine via the incorporation of bayesian algorithm optimized CNN-LSTM neural network</article-title>. <source>Energy</source>. <year>2024</year>;<volume>290</volume>:<fpage>130326</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.energy.2024.130326</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Laplace-domain hybrid distribution model based FDIA attack sample generation in smart grids</article-title>. <source>Symmetry</source>. <year>2023</year>;<volume>15</volume>(<issue>9</issue>):<fpage>1669</fpage>. doi:<pub-id pub-id-type="doi">10.3390/sym15091669</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>