<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">EE</journal-id>
<journal-id journal-id-type="nlm-ta">EE</journal-id>
<journal-id journal-id-type="publisher-id">EE</journal-id>
<journal-title-group>
<journal-title>Energy Engineering</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-0118</issn>
<issn pub-type="ppub">0199-8595</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">62667</article-id>
<article-id pub-id-type="doi">10.32604/ee.2025.062667</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Transient Stability Assessment Model and Its Updating Based on Dual-Tower Transformer</article-title>
<alt-title alt-title-type="left-running-head">Transient Stability Assessment Model and Its Updating Based on Dual-Tower Transformer</alt-title>
<alt-title alt-title-type="right-running-head">Transient Stability Assessment Model and Its Updating Based on Dual-Tower Transformer</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Nan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>20102057@neepu.edu.cn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Dong</surname><given-names>Jingxiong</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Tao</surname><given-names>Liang</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Liang</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Key Laboratory of Modern Power System Simulation and Control &#x0026; Renewable Energy Technology, Ministry of Education (Northeast Electric Power University)</institution>, <addr-line>Jilin, 132012</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Electrical Engineering, Northeast Electric Power University</institution>, <addr-line>Jilin, 132012</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>State Grid Jilin Electric Power Co., Ltd., Siping Power Supply Company</institution>, <addr-line>Siping, 136000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Nan Li. Email: <email>20102057@neepu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>06</month><year>2025</year>
</pub-date>
<volume>122</volume>
<issue>7</issue>
<fpage>2957</fpage>
<lpage>2975</lpage>
<history>
<date date-type="received">
<day>24</day>
<month>12</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>3</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_EE_62667.pdf"></self-uri>
<abstract>
<p>With the continuous expansion and increasing complexity of power system scales, the binary classification for transient stability assessment in power systems can no longer meet the safety requirements of power system control and regulation. Therefore, this paper proposes a multi-class transient stability assessment model based on an improved Transformer. The model is designed with a dual-tower encoder structure: one encoder focuses on the time dependency of data, while the other focuses on the dynamic correlations between variables. Feature extraction is conducted from both time and variable perspectives to ensure the completeness of the feature extraction process, thereby enhancing the accuracy of multi-class evaluation in power systems. Additionally, this paper introduces a hybrid sampling strategy based on sample boundaries, which addresses the issue of sample imbalance by increasing the number of boundary samples in the minority class and reducing the number of non-boundary samples in the majority class. Considering the frequent changes in power grid topology or operation modes, this paper proposes a two-stage updating scheme based on self-supervised learning: In the first stage, self-supervised learning is employed to mine the structural information from unlabeled data in the target domain, enhancing the model&#x2019;s generalization capability in new scenarios. In the second stage, a sample screening mechanism is used to select key samples, which are labeled through long-term simulation techniques for fine-tuning the model parameters. This allows for rapid model updates without relying on many labeled samples. This paper&#x2019;s proposed model and update scheme have been simulated and verified on two node systems, the IEEE New England 10-machine 39-bus system and the IEEE 47-machine 140-bus system, demonstrating their effectiveness and reliability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Transient stability assessment</kwd>
<kwd>sample imbalance</kwd>
<kwd>dual-tower transformer network</kwd>
<kwd>self-supervised learning</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, the integration of new energy sources has introduced numerous power electronics into the power system, complicating grid structure and posing significant challenges to its safe, stable operation.</p>
<p>Data-driven transient stability assessment (TSA) methods, leveraging their powerful feature extraction and data processing capabilities, have been widely applied in the TSA of power systems, including methods such as support vector machine (SVM) [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>], convolutional neural network (CNN) [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>], and deep belief network [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. However, these studies mainly focused on binary classification, neglecting the diverse instability modes, which include both aperiodic instability due to insufficient synchronous torque and oscillatory instability caused by factors like inadequate damping torque. In the case of aperiodic instability, the primary solutions involve tripping, reducing generator mechanical power, and maintaining system stability through excitation regulation [<xref ref-type="bibr" rid="ref-7">7</xref>]. In contrast, the main solutions for oscillatory instability involve activating power system stabilizers [<xref ref-type="bibr" rid="ref-8">8</xref>] or adding broadband damping controllers to flexible AC transmission systems devices to suppress oscillations [<xref ref-type="bibr" rid="ref-9">9</xref>]. The causes and countermeasures for these two instability modes differ. Therefore, traditional binary classification assessment methods fail to provide comprehensive and accurate decision support for power grid dispatchers, thereby increasing the safety risks of power grid operation. Against this background, conducting multi-class transient discrimination holds significant research importance.</p>
<p>Currently, compared to simple binary classification for discrimination, research on multi-class TSA is relatively scarce. Wang et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] utilized a noise-resistant Gaussian process model to achieve discrimination among stable, critical, and unstable conditions. The Gaussian process model, which is established through probabilistic theory, demonstrates excellent performance in processing power data containing significant uncertainties and noise. Shi et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] optimized a CNN model using stochastic gradient descent with warm restarts, enabling continuous refinement of extracted features to differentiate between stable, aperiodic instability, and oscillatory instability. Li et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] utilized a fusion model combining XGBoost and decision trees to perform feature extraction, to distinguish between stable conditions and two types of unstable modes. Li et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] refined the extraction of data features using three different deep-learning models for a more comprehensive assessment of various stability margins and instability degrees. Despite these studies having made preliminary progress in multi-class TSA research, they still face challenges when dealing with high-dimensional time-series data in power systems: The time dependence of multivariate variables embedded in high-dimensional power data reflects the evolution patterns of these variables over time, providing crucial information about the macroscopic system state. Meanwhile, the dynamic correlation among multivariate variables uncovers intrinsic links between different variables in the sequence, crucial for understanding interactions and influences within the power system. However, current research lacks full integration of power data&#x2019;s global and local features and overlooks complete data extraction, limiting further enhancements in multi-class discrimination accuracy. Moreover, the aforementioned studies on multi-class TSA have not considered the biased tendency of the model caused by the issue of sample imbalance [<xref ref-type="bibr" rid="ref-14">14</xref>], which reduces the model&#x2019;s discrimination accuracy for minority-class samples.</p>
<p>When the topology or operation mode of the power system changes, the performance of data-driven models tends to decline or even fail [<xref ref-type="bibr" rid="ref-15">15</xref>]. Retraining the model for a new system structure demands time and risks losing accumulated knowledge. Existing research findings can be broadly summarized into two categories: (1) Efforts are focused on achieving rapid model updates under the condition of limited samples in the target domain: Wang et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] utilized the data inheritance method to retain original data information, allowing rapid model updates with minimal new samples. Zhou et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] utilized active learning to screen high-value samples, allowing the model to achieve good evaluation performance based on only a restricted quantity of key samples. (2) Efforts are focused on expanding the number of training samples available for model updates, thereby enhancing the model&#x2019;s evaluation performance: Zhan et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] utilized an improved generative adversarial network to generate many samples with the same distribution, which are then merged with transfer samples from the source domain to fine-tune the model, effectively reducing the number of target domain samples required. Tang et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] employed a method integrating sample and feature transfer to augment training samples for model updates at both levels. However, both of these solutions have limitations: Implementing model updates in the target domain with a small sample size may reduce update time, but it may lead to overfitting of the model, thereby limiting its generalization ability. Generating more available samples for model updates can improve model performance, but it will significantly increase the update time.</p>
<p>Given all the above concerns, this paper proposes a multi-class TSA method and a two-stage update scheme based on an improved Transformer model. The main contributions of the paper are highlighted as follows:
<list list-type="order">
<list-item>
<p>Addressing the issue of information completeness in model feature extraction, this paper proposes a dual-tower Transformer encoder network. This network comprehensively extracts features from both temporal and variable dimensions to enhance the accuracy of multi-class transient state classification. Additionally, to tackle the model bias caused by imbalanced samples, this paper introduces a hybrid sampling strategy based on sample boundaries, which addresses the issue of sample imbalance by increasing the number of boundary samples in the minority class and reducing the number of non-boundary samples in the majority class.</p></list-item>
<list-item>
<p>In response to the need for online model updates, this paper designs a two-stage update scheme based on self-supervised learning: In the first stage, self-supervised learning is employed to learn rich internal representations of unlabeled data, enhancing the model&#x2019;s generalization performance during the transition period. In the second stage, a key sample selection mechanism is used to screen and annotate key samples, which are then utilized to optimize the model.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Feature Extraction Network Based on a Dual-Tower Transformer Encoder Model</title>
<p>When analyzing power system status using a data-driven model, the completeness and adequacy of feature extraction are crucial for determining model performance. Due to the topological interconnectivity among nodes in the power system, various physical quantities also have a certain correlation with each other. As time progresses, the interactions of these physical quantities evolve dynamically in intensity and pattern. Therefore, accurately extracting temporal dependence and dynamic correlations among multivariate variables is crucial in power data analysis. The Standard Transformer model, featuring its unique multi-head attention, focuses adaptively on high-value information [<xref ref-type="bibr" rid="ref-20">20</xref>]. So, it has been widely applied in multiple fields [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. However, when processing time-series data, the standard Transformer only extracts the temporal dependence of multivariate variables, neglecting the dynamic correlation among them. Considering the limitations of the standard Transformer and the unique features of power data, this paper proposes a dual-tower Transformer encoder model, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Dual-tower transformer encoder network</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-1.tif"/>
</fig>
<p>The feature extraction network is composed of two independent encoder structures connected in parallel through an attention-gate control mechanism. Encoder I is designed to focus on the temporal dependence of multivariate variables and deeply extract their evolutionary rules over time. Encoder II focuses on dynamic correlations between multivariate variables, revealing intrinsic relationships within the sequence. The attention gate control mechanism can adaptively allocate weight to the feature information of time and variable dimensions, and enhance the model&#x2019;s sensitivity to important features, making the model have good interpretability. In addition, this paper uses the gated linear unit (GLU) to replace the feedforward neural network (FNN) for the nonlinear transformation of features, reducing the model&#x2019;s computational overhead.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Parallel Embedding Layer Design</title>
<p>The input data <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a multivariate time series consisting of <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>m</mml:mi></mml:math></inline-formula> variables, with the length of each variable being <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>w</mml:mi></mml:math></inline-formula>. The standard Transformer network maps all the variables in <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>X</mml:mi></mml:math></inline-formula> at the same time to a d-dimensional vector through an embedding layer and then performs feature extraction on it. The d-dimensional vectors at different time points form a matrix <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the d-dimensional vectors of all variables at the time <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> after embedding, as is shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represent the sample values of the <italic>i</italic>-th variable at the time <italic>i</italic>; <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the weight coefficient matrix; and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the bias.</p>
<p>The output feature matrix <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> partitions the multi-dimensional time series of the input into blocks along the time dimension, forming a feature representation centered on time, thus enabling the model to better capture the dynamic evolution patterns of the multi-variable time series. However, this approach causes the independent features of different variables to be intertwined into multi-dimensional features, making it difficult for the model to distinguish and capture the dynamic correlations between different variables. Therefore, the paper constructs a parallel embedding layer structure to solve this problem, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Parallel embedding layer</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-2.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-2">Fig. 2a</xref> shows the temporal embedding layer, which retains the feature embedding approach used in the standard Transformer model, and this operation ensures that the model can fully capture the evolution patterns of the sequence along the time axis. <xref ref-type="fig" rid="fig-2">Fig. 2b</xref> shows the variable embedding layer, which embeds all the sample values of a single variable at all time points into a d-dimensional vector, and after embedding, the m variables form the matrix <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, which <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the d-dimensional vector obtained by embedding the sample values of the <italic>j</italic>-th variable at all time points, as shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the sample value of the <italic>j</italic>-th variable at time <italic>j</italic>; <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the weight coefficient matrix; and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the bias.</p>
<p>The output matrix <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> partitions the input time series in the variable dimension, forming a feature representation centered on the variable, allowing the model to capture the dynamic correlation between different variables. Therefore, by designing the parallel embedding layer, the model can simultaneously pay attention to the temporal dependency and dynamic correlation between variables in the time-series data.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Attention-Gate Control Mechanism</title>
<p>To effectively integrate the outputs of the dual-tower encoder, this paper designs an attention-gate control mechanism. The attention gate control mechanism is composed of an attention module and a GLU module, which realizes the nonlinear integration of the output matrices <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of the dual-tower encoder. This operation can adaptively adjust the attention distribution of the output feature information, and focus on the key feature information of the dual-tower encoder outputs, thereby achieving precise gating and flow control. Its operation flow is as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>Feature Linear Transformation: Linear transformation is applied to the feature matrices <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, mapping them to the query, key, and value matrices <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>Q</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>K</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>V</mml:mi></mml:math></inline-formula>, the mapping formula is as shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represent the weight matrices for the query, key, and value, respectively.</p></list-item>
<list-item><label>(2)</label><p>Attention score calculation: Multiply matrix <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>Q</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>K</mml:mi></mml:math></inline-formula> element-wise, and scale them along dimension <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msqrt><mml:mi>d</mml:mi></mml:msqrt></mml:math></inline-formula>; scale the result through the <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>S</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:math></inline-formula> function to convert it into a probability score, which is then multiplied by the matrix <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>V</mml:mi></mml:math></inline-formula> to obtain the attention score <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>F</mml:mi></mml:math></inline-formula>, as shown in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>F</mml:mi><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>Q</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:msqrt><mml:mi>d</mml:mi></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>V</mml:mi></mml:math></disp-formula></p></list-item>
<list-item><label>(3)</label><p>Feature Integration: The obtained attention values are integrated through a gate mechanism to obtain the output feature <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, as shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>F</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2297;</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mi>F</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the weight coefficient matrices, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the biases, and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is the sigmoid activation function.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>A Multi-Stage TSA Based on the Dual-Tower Transformer</title>
<sec id="s3_1">
<label>3.1</label>
<title>Feature Set Construction</title>
<p>This paper employs busbar voltage magnitude data of individual nodes in the power system as the original dataset. The voltage curves under different system states are shown in <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>,<xref ref-type="fig" rid="fig-3">b</xref>. As evident from these figures, when the system is subjected to a certain disturbance, the busbar voltage in a stable state can recover to its original stable state or attain a new stable state over time. However, in an unstable state, the busbar voltage fluctuates continuously with varying vibration trends. Combining the changes in rotor angle curves can demonstrate the differences between the two instability modes, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3c</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Curve diagrams for stable state, aperiodic instability state, and oscillatory instability state. (a) Under stability conditions; (b) Under instability condition; (c) Rotor angle curve diagram</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-3.tif"/>
</fig>
<p>Considering the potential data loss during PMU data collection, this paper constructs a 30-dimensional trajectory cluster dataset based on the original dataset to prevent the data-driven model from becoming inoperable due to the loss of busbar data [<xref ref-type="bibr" rid="ref-24">24</xref>]. The operating state of the power system is divided into three cases: the stable state is labeled 0, the aperiodic unstable state is labeled 1, and the oscillatory unstable state is labeled 2. The stability of the power system is judged by the transient stability index (TSI) of the generator rotor angle after the fault, and the aperiodic instability and oscillatory instability are further distinguished based on the divergence of the power angle curve. TSI calculation formula is as shown in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mn>360</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mn>360</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> represents the maximum relative power angle difference between any two generators. When <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the sample is classified as a stable state. When <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the sample is classified as an unstable state.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>TSA Model Design</title>
<p>This paper designs a TSA model composed of the dual-tower Transformer encoder, fully connected layer, and Softmax classifier, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The model assessment process is as follows:</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Model framework diagram</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-4.tif"/>
</fig>
<p><list list-type="simple">
<list-item><label>(1)</label><p>A 30-dimensional trajectory cluster dataset is obtained from the power system, as shown in the data processing section of <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, where <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the <italic>m</italic>-th feature of the <italic>i</italic>-th sample at the <italic>w</italic>-th time point. Then, each <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>X</mml:mi></mml:math></inline-formula> is normalized to obtain the normalized data <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item><label>(2)</label><p>Input <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> into a dual-tower encoder network for feature extraction. The obtained vector <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> with time-dependent features and vector <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> with inter-variable correlation features are fused through the attention-gate control mechanism to output a feature fusion matrix <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p></list-item>
<list-item><label>(3)</label><p>Convert the vectors in the matrix <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> into column vectors <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, after passing through a fully connected layer, and perform a linear operation to obtain the vector <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>Z</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>Z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <italic>n</italic> is the number of categories for TSA.</p></list-item>
<list-item><label>(4)</label><p>The output vector <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>Z</mml:mi></mml:math></inline-formula> is processed through the Softmax function to obtain the corresponding probabilities <italic>P</italic> for each category. The category corresponding to the maximum probability value is selected as the predicted category, and the probability calculation formula is shown in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the probability value of the input sample belonging to <italic>j</italic>-th class; <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>e</mml:mi></mml:math></inline-formula> is the exponential function; <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the values of the <italic>i-</italic>th and <italic>j-</italic>th elements in the vector <italic>Z</italic>, respectively.</p></list-item>
</list>
</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Assessment Process</title>
<p>In order to effectively reduce the probability of misclassification of unstable samples, this paper has developed a multi-stage TSA method, which includes pre-training, re-training, and online assessment stages.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Pre-Training Stage</title>
<p>Initially trained on a large-scale dataset, the model captures data features comprehensively. To measure and optimize performance, it employs cross-entropy loss with the RAdam algorithm for dynamic, adaptive parameter adjustment. Meanwhile, to effectively curb overfitting caused by excessive model complexity, L2 regularization is introduced to impose constraints on model parameters, generating a pre-trained model.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Re-Training Stage</title>
<p>To address the model bias caused by sample imbalance, this paper designs a hybrid sampling strategy based on sample boundaries:
<list list-type="simple">
<list-item><label>(1)</label><p>Use the pre-trained model to evaluate each training sample individually, set the evaluation probability threshold &#x03B3; &#x003D; 0.9, and when the accuracy rate of a sample is greater than or equal to &#x03B3;, it is classified as a non-boundary sample, otherwise, it is classified as a boundary sample.</p></list-item>
<list-item><label>(2)</label><p>For the minority class samples in the boundary samples, use the smote algorithm for over-sampling, and for the stable samples in the non-boundary samples, perform random under-sampling. Finally, form a hybrid sampling training set.</p></list-item>
<list-item><label>(3)</label><p>Use the hybrid sampling training set to re-train the pre-trained model to improve the model&#x2019;s discriminative accuracy rate for minority classes. When the model&#x2019;s assessment accuracy rate meets the assessment requirements, save the model parameters and generate an assessment model.</p></list-item>
</list></p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Online Assessment Stage</title>
<p>Real-time sampling of grid data is conducted, and the TSA model is used to assess the sampled data. The probabilities <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> of the sampled data being judged as stable, initial swing instability, and multiple swing instability are calculated. When <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, the sample is judged to be stable. When <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, the sample is judged to be aperiodic instability. When <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, the sample is judged to be oscillatory instability.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>A Two-Stage Updating Scheme Based on the Dual-Tower Transformer</title>
<p>When changes occur in the topology or operational mode of the power grid, there often emerges a significant distribution discrepancy between the newly formed target domain data and the source domain data. This discrepancy greatly diminishes the evaluation accuracy of the source domain model on the target domain. To quantify the degree of difference between the target domain and the source domain, this paper adopts maximum mean discrepancy (MMD) [<xref ref-type="bibr" rid="ref-25">25</xref>] technology as an evaluation tool. Based on different MMD values I<sub>MMD</sub>, differentiated transfer paths are designed, with the update process illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Flowchart of the updating scheme</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-5.tif"/>
</fig>
<p>Path 1: When I<sub>MMD</sub> &#x2264; 0.5, the difference between target domain data and source domain data is relatively small. The source domain model is used to evaluate unlabeled samples in the target domain, and the <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:math></inline-formula> samples with the lowest evaluation probabilities are selected as key samples. This operation is referred to as key sample selection mechanism I. Subsequently, long-term simulation techniques are utilized to annotate these key samples. Finally, fine-tuning just the output layer with limited labeled samples yields satisfactory update results.</p>
<p>Path 2: When I<sub>MMD</sub> &#x003E; 0.5, significant data discrepancy exists between target and source domains, rendering output layer fine-tuning insufficient for model evaluation accuracy. Considering the issues present in current research approaches: (1) Relying solely on a small number of samples for updates can improve the update speed but cannot meet the accuracy requirements for evaluation; (2) Generating more training samples can enhance model performance but increases time costs. Therefore, this paper proposes a two-stage updating scheme based on self-supervised learning.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Stage 1: Regression Prediction Based on Self-Supervised Learning</title>
<p>Self-supervised learning [<xref ref-type="bibr" rid="ref-26">26</xref>], as an important branch of unsupervised learning [<xref ref-type="bibr" rid="ref-27">27</xref>], possesses the ability to autonomously mine and extract the intrinsic representation characteristics of unlabeled data. In this paper, we introduce the masked prediction task into self-supervised learning as an auxiliary strategy, utilizing unlabeled samples to conduct pseudo-supervised training of the model. The specific operation of self-supervised learning based on masked prediction is as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>Data masking: For each unlabeled input sample <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>X</mml:mi></mml:math></inline-formula> of the model, create an independent binary noise mask <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The elements of <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> are 0 and 1, with 0 segments being the masked segments, which account for a proportion of r, and the distribution follows a geometric distribution with a mathematical expectation of <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>E</mml:mi></mml:math></inline-formula>. Multiply the noise mask with the input data element-wise, and use the masked segments to obscure the input data, generating the masked data <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the <italic>i</italic>-th vectors.</p></list-item>
<list-item><label>(2)</label><p>Regression prediction: Input the masked input data into the source domain model, predict and reshape the masked segments of the masked data <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, and output the predicted estimated value <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>Y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of the masked data <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>. By optimizing the prediction value continuously, the source domain model&#x2019;s parameters are updated, and finally, the first-stage pre-trained model is generated.</p></list-item>
</list></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Stage 2: Model Optimization</title>
<p>The initial update in Stage 1 alone cannot enhance the model&#x2019;s assessment performance to meet the predetermined standard. However, after the initial update in Stage 1, the model has learned rich data representations of the new system from unlabeled samples. So, when optimizing the pre-trained model in Stage 1, no more labeled samples are needed, greatly reducing the time required for sample generation. The update steps are as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>Key Sample Selection Mechanism II: Using the pre-trained model from the first stage, perform masked predictions on unlabeled samples, and define the top <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:math></inline-formula> samples with the highest predicted loss value as key unlabeled samples.</p></list-item>
<list-item><label>(2)</label><p>By conducting long-term simulation to label key unlabeled samples, and then using parameter fine-tuning techniques to update model parameters, the evaluation performance of the model is improved, and a second stage transient stability assessment model is generated.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Case Analysis</title>
<sec id="s5_1">
<label>5.1</label>
<title>Data Set and Assessment Indicators</title>
<p>In this simulation, the software uses PSD-BPA software, the assembly language is Python, and the deep learning framework is PyTorch. The case study uses the IEEE New England 10-machine 39-bus system and IEEE 47-machine 140-bus system for simulation experiments. The database setting is shown in <xref ref-type="table" rid="table-1">Table 1</xref>, where S1 and S2 are the source domain databases, and A-F are the target domain databases. The case study considers 9 load levels at 80%, 85%, 90%, &#x2026; , 120%, and the fault location is selected at 0%, 10%, 20%, &#x2026; , 90% of the line, with the fault type set as a three-phase short-circuit fault. The fault duration is randomly set within 0.1 to 0.2 s, and the simulation time is 20 s. The S1 database has 12,240 samples, including 7986 stable samples, 3042 aperiodic unstable samples, and 1212 oscillatory unstable samples. The S2 database has 22,140 samples, comprising 15,201 stable samples, 3853 aperiodic unstable samples, and 3086 oscillatory unstable samples. The datasets are divided into training, validation, and test sets in an 8:1:1 ratio.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Settings for different datasets</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="35mm"/>
<col align="center" width="15mm"/>
<col align="center" width="25mm"/>
<col align="center" width="25mm"/>
</colgroup>
<thead>
<tr>
<th align="center">System</th>
<th align="center">Datasets</th>
<th align="center">Topological transformation</th>
<th align="center">Load levels</th>
<th align="center">The location of the fault</th>
<th align="center">The duration of the fault/s</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">IEEE New England 10-machine 39-bus system</td>
<td>S1</td>
<td>All Busbar are in operation.</td>
<td>80&#x007E;120</td>
<td>0%&#x007E;90%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>A</td>
<td>Busbar 4&#x2013;14 outage</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>B</td>
<td>Busbar 6&#x2013;11 outage</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>C</td>
<td>Busbars 2&#x2013;3 and 15&#x2013;16 outage</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>D</td>
<td>Busbars 4&#x2013;14 and 10&#x2013;11 are out of service</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td rowspan="3">IEEE 47-machine 140-bus system</td>
<td>S1</td>
<td>All lines are in operation.</td>
<td>80&#x007E;120</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>E</td>
<td>Busbar 53&#x2013;55 and 78&#x2013;79 outage</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
<tr>
<td>F</td>
<td>Busbar 54&#x2013;56 and 78&#x2013;79 outage</td>
<td>75&#x007E;125</td>
<td>0%&#x007E;80%</td>
<td>0.1&#x007E;0.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To evaluate the model performance comprehensively, we introduced a confusion matrix, as shown in <xref ref-type="table" rid="table-2">Table 2</xref>, and three classification assessment indicators, as shown in <xref ref-type="disp-formula" rid="eqn-8">Eqs. (8)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-10">(10)</xref>:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>F</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mroot><mml:mrow><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mfrac></mml:mrow><mml:mn>3</mml:mn></mml:mroot><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the overall assessment accuracy rate, <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>F</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the harmonic mean of the precision and recall for samples labeled as <italic>i</italic>, and <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the comprehensive performance index, which is an important indicator reflecting the model&#x2019;s discriminative performance on minority sample.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Confusion matrix</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Actual</th>
<th colspan="4">Predicted</th>
</tr>
<tr>
<th/>
<th>Stable</th>
<th>Aperiodic unstable</th>
<th>Oscillatory unstable</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td>Stable</td>
<td><italic>T</italic><sub>00</sub></td>
<td><italic>T</italic><sub>01</sub></td>
<td><italic>T</italic><sub>02</sub></td>
<td><italic>T</italic><sub>0</sub></td>
</tr>
<tr>
<td>Aperiodic unstable</td>
<td><italic>T</italic><sub>10</sub></td>
<td><italic>T</italic><sub>11</sub></td>
<td><italic>T</italic><sub>12</sub></td>
<td><italic>T</italic><sub>1</sub></td>
</tr>
<tr>
<td>Oscillatory unstable</td>
<td><italic>T</italic><sub>20</sub></td>
<td><italic>T</italic><sub>21</sub></td>
<td><italic>T</italic><sub>22</sub></td>
<td><italic>T</italic><sub>2</sub></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Comparison of Assessment Performance among Different Models</title>
<p>To assess the dual-tower Transformer&#x2019;s performance, we compared it with six other models: dual-tower FNN-Transformer, standard Transformer, CNN&#x002B;LSTM, LSTM, CNN, and SVM. The dual-tower FNN-Transformer integrates an FNN module for nonlinear transformation, while the CNN&#x002B;LSTM model separately extracts and analyzes spatial and temporal features. The model parameters are as follows: the dual-tower Transformer and standard Transformer models have 8-head attention mechanisms and 4 sub-modules. The LSTM algorithm uses 4 hidden layers with a hidden layer dimension of 150. The CNN model comprises 4 convolutional layers with a kernel size of 3 &#x00D7; 3, and each layer is followed by a 2 &#x00D7; 2 max pooling layer. The SVM kernel function uses the radial basis function, with hyperparameters C &#x003D; 10 and kernel function parameters. All models are set with an Epoch size of 200, a Batch size of 128, an initial learning rate of Lr &#x003D; 0.001, and an L2 regularization term of 0.1. The average of 20 repeated experiments is taken as the evaluation result, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. Simultaneously, to test the working efficiency of the dual-tower Transformer model, the training durations of both the dual-tower Transformer model and the dual-tower FNN-Transformer model were recorded, with the results shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Assessment performance of different models</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Systems</th>
<th>Model</th>
<th colspan="5">Assessment indicators</th>
</tr>
<tr>
<th/>
<th/>
<th><italic>P</italic><sub><italic><bold>ACC</bold></italic></sub></th>
<th><italic>F1</italic><sup><italic><bold>1</bold></italic></sup></th>
<th><italic>F1</italic><sup><italic><bold>2</bold></italic></sup></th>
<th><italic>F1</italic><sup><italic><bold>3</bold></italic></sup></th>
<th><italic>G</italic><sub><italic><bold>mean</bold></italic></sub></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="8">IEEE New England 10-machine 39-bus system</td>
<td>Mixed Sampling &#x002B; dual-tower Transformer</td>
<td>98.45</td>
<td>99.00</td>
<td>98.07</td>
<td>94.83</td>
<td>97.50</td>
</tr>
<tr>
<td>Dual-tower Transformer</td>
<td>98.10</td>
<td>98.82</td>
<td>97.31</td>
<td>95.22</td>
<td>96.85</td>
</tr>
<tr>
<td>Dual-tower FNN-Transformer</td>
<td>98.12</td>
<td>98.99</td>
<td>97.59</td>
<td>93.67</td>
<td>95.47</td>
</tr>
<tr>
<td>CNN&#x002B;LSTM</td>
<td>96.59</td>
<td>97.95</td>
<td>98.08</td>
<td>87.92</td>
<td>93.60</td>
</tr>
<tr>
<td>Transformer</td>
<td>97.38</td>
<td>98.35</td>
<td>98.83</td>
<td>91.02</td>
<td>94.90</td>
</tr>
<tr>
<td>LSTM</td>
<td>95.88</td>
<td>97.46</td>
<td>97.79</td>
<td>86.08</td>
<td>92.70</td>
</tr>
<tr>
<td>CNN</td>
<td>95.13</td>
<td>98.83</td>
<td>93.06</td>
<td>78.62</td>
<td>89.75</td>
</tr>
<tr>
<td>SVM</td>
<td>95.08</td>
<td>97.35</td>
<td>96.05</td>
<td>81.91</td>
<td>91.95</td>
</tr>
<tr>
<td rowspan="8">IEEE 47-machine 140-bus system</td>
<td>Mixed Sampling &#x002B; dual-tower Transformer</td>
<td>98.78</td>
<td>99.41</td>
<td>97.54</td>
<td>97.90</td>
<td>98.35</td>
</tr>
<tr>
<td>Dual-tower Transformer</td>
<td>98.42</td>
<td>99.21</td>
<td>97.56</td>
<td>95.42</td>
<td>97.91</td>
</tr>
<tr>
<td>Dual-tower FNN-Transformer</td>
<td>98.37</td>
<td>98.87</td>
<td>97.62</td>
<td>97.00</td>
<td>97.08</td>
</tr>
<tr>
<td>CNN&#x002B;LSTM</td>
<td>97.06</td>
<td>98.19</td>
<td>98.58</td>
<td>89.54</td>
<td>95.05</td>
</tr>
<tr>
<td>Transformer</td>
<td>97.70</td>
<td>98.55</td>
<td>98.84</td>
<td>92.01</td>
<td>96.01</td>
</tr>
<tr>
<td>LSTM</td>
<td>96.57</td>
<td>98.00</td>
<td>97.94</td>
<td>87.52</td>
<td>94.36</td>
</tr>
<tr>
<td>CNN</td>
<td>95.83</td>
<td>99.00</td>
<td>92.16</td>
<td>81.67</td>
<td>90.66</td>
</tr>
<tr>
<td>SVM</td>
<td>96.07</td>
<td>97.84</td>
<td>97.08</td>
<td>86.22</td>
<td>92.36</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of transfer time</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Systems</th>
<th>Model</th>
<th>Time/s</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">IEEE New England 10-machine 39-bus system</td>
<td>Dual-tower transformer</td>
<td>947.38</td>
</tr>
<tr>
<td>Dual-tower FNN-transformer</td>
<td>1623.72</td>
</tr>
<tr>
<td rowspan="2">IEEE 47-machine 140-bus system</td>
<td>Dual-tower transformer</td>
<td>1689.64</td>
</tr>
<tr>
<td>Dual-tower FNN-transformer</td>
<td>2926.39</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-3">Table 3</xref> demonstrates that the Transformer model outperforms LSTM, CNN, and SVM in assessment tasks, due to its multi-head attention mechanism that targets key features. The dual-tower Transformer, which extracts data from both time and variable dimensions, exhibits superior performance. Although the CNN&#x002B;LSTM model, which captures features from both dimensions, performs better than each model, it still lags behind the dual-tower Transformer. Retraining the dual-tower Transformer using a mixed sampling set significantly enhanced indicators <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, validating the strategy&#x2019;s effectiveness. According to <xref ref-type="table" rid="table-4">Table 4</xref>, during the training process of the two systems, the dual-tower Transformer model significantly outperformed the dual-tower FNN-Transformer model in terms of training speed. This validates that the Transformer model, which utilizes GLU instead of FNN for nonlinear transformation, possesses higher training efficiency while maintaining superior evaluation performance.</p>

</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Visualization of Attention Weight Distribution</title>
<p>To verify the effectiveness and explainability of the designed dual-tower Transformer in allocating attention to feature information, this experiment randomly selected 50 samples for testing. Heatmaps in <xref ref-type="fig" rid="fig-6">Figs. 6</xref> and <xref ref-type="fig" rid="fig-7">7</xref> visualize the evolution process of attention weight allocation in Encoder I and Encoder II, respectively, with color depth indicating weight (darker means higher).</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Time attention distribution. (a) The 5th round of time attention distribution; (b) The 100th round of time attention distribution</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Variable attention distribution. (a) The 5th round of variable attention distribution; (b) The 100th round of variable attention distribution</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-7.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-6">Figs. 6</xref> and <xref ref-type="fig" rid="fig-7">7</xref> illustrate the changes in attention weights of the tower encoder in time and variable dimensions over training epochs. In <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, as the number of training rounds increases, the attention mechanism of Encoder I increasingly focuses on features from later periods, and this conforms to the system&#x2019;s state gradually becoming clearer as time passes after the fault occurs. In <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the attention weights of Encoder II are evenly distributed among features initially but develop a clear bias by the 100th round, showing the model learns to emphasize key features. The figures demonstrate that the feature extraction network based on the model can dynamically adjust the feature weights and verify its interpretability.</p>

</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Performance Testing of Transferability in Different Target Domains</title>
<p>The source domain model is tested on different target domain datasets, respectively, and the average of 20 repeated experiments is taken as the evaluation result. The results are shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Pre-update model performance</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="2">Target domain</th>
<th colspan="5">Assessment indicators</th>
</tr>
<tr>
<th colspan="2"></th>
<th><italic>P</italic><sub><italic><bold>ACC</bold></italic></sub></th>
<th><italic>F1</italic> <sup><italic><bold>1</bold></italic></sup></th>
<th><italic>F1</italic> <sup><italic><bold>2</bold></italic></sup></th>
<th><italic>F1</italic> <sup><italic><bold>3</bold></italic></sup></th>
<th><italic>G</italic><sub><italic><bold>mean</bold></italic></sub></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4">IEEE New England 10-machine 39-bus system</td>
<td>A</td>
<td>79.03</td>
<td>89.42</td>
<td>92.11</td>
<td>31.40</td>
<td>63.71</td>
</tr>
<tr>
<td>B</td>
<td>74.29</td>
<td>81.73</td>
<td>82.31</td>
<td>38.29</td>
<td>63.63</td>
</tr>
<tr>
<td>C</td>
<td>69.46</td>
<td>74.61</td>
<td>75.73</td>
<td>38.19</td>
<td>59.98</td>
</tr>
<tr>
<td>D</td>
<td>68.66</td>
<td>85.08</td>
<td>70.16</td>
<td>37.36</td>
<td>60.64</td>
</tr>
<tr>
<td rowspan="2">IEEE 47-machine 140-bus system</td>
<td>E</td>
<td>73.66</td>
<td>82.95</td>
<td>65.17</td>
<td>23.18</td>
<td>30.42</td>
</tr>
<tr>
<td>F</td>
<td>76.38</td>
<td>84.00</td>
<td>68.80</td>
<td>37.04</td>
<td>59.82</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows that the evaluation performance of the source domain model in each target domain has dropped sharply, with accuracy rates below 80%, which fails to meet the accuracy requirements for transient stability assessment. Therefore, the source domain model needs to be updated to adapt to the new system structure. Since the I<sub>MMD</sub> values between the six target domains and the source domain dataset are all greater than 0.5, this paper utilizes the two-stage updating scheme to update the source domain model.</p>

<p>In the first stage update, the mask ratio is set to 0.15, with a geometric distribution whose mathematical expectation is 3. In the second stage, 1500 critical samples are used, the initial learning rate is set to 0.001, and 100 epochs are run, with other settings remaining the same as in <xref ref-type="sec" rid="s5_2">Section 5.2</xref>. The assessment result is the average of 20 repeated experiments, as shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Post-update model performance</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="2">Target domain</th>
<th colspan="5">Assessment indicators</th>
</tr>
<tr>
<th colspan="2"></th>
<th><italic>P</italic><sub><italic><bold>ACC</bold></italic></sub></th>
<th><italic>F1</italic> <sup><italic><bold>1</bold></italic></sup></th>
<th><italic>F1</italic> <sup><italic><bold>2</bold></italic></sup></th>
<th><italic>F1</italic> <sup><italic><bold>3</bold></italic></sup></th>
<th><italic>G</italic><sub><italic><bold>mean</bold></italic></sub></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4">IEEE New England 10-machine 39-bus system</td>
<td>A</td>
<td>97.47</td>
<td>98.38</td>
<td>96.19</td>
<td>94.29</td>
<td>95.57</td>
</tr>
<tr>
<td>B</td>
<td>97.52</td>
<td>98.34</td>
<td>99.22</td>
<td>91.51</td>
<td>98.64</td>
</tr>
<tr>
<td>C</td>
<td>97.00</td>
<td>97.13</td>
<td>97.15</td>
<td>96.67</td>
<td>97.03</td>
</tr>
<tr>
<td>D</td>
<td>97.30</td>
<td>97.51</td>
<td>97.49</td>
<td>96.83</td>
<td>97.33</td>
</tr>
<tr>
<td rowspan="2">IEEE 47-machine 140-bus system</td>
<td>E</td>
<td>97.38</td>
<td>98.35</td>
<td>98.83</td>
<td>91.02</td>
<td>94.91</td>
</tr>
<tr>
<td>F</td>
<td>97.24</td>
<td>98.31</td>
<td>98.58</td>
<td>90.91</td>
<td>94.56</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-6">Table 6</xref> shows that the model&#x2019;s assessment performance significantly improves after two-stage updating, achieving over 97% accuracy even in cases of two bus outages, verifying the scheme&#x2019;s effectiveness.</p>

<p>To verify the performance of the key sample selection mechanism based on the first-stage transfer model, 1500 key samples selected by the selection mechanism were compared with 1500 random samples through t-random nearest neighbor embedding dimensionality reduction visualization, the distribution of samples as shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Sample distribution chart. (a) Random sample distribution; (b) Key sample distribution</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-8.tif"/>
</fig>
<p>From <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, it can be seen that the overlap among the key samples is more significant than that among the random samples, and the proportion of minority class samples increases. This is because the samples in the overlapping area and the minority class samples are more difficult to distinguish. Therefore, these key samples have higher value information, verifying the effectiveness of the key sample selection mechanism.</p>

</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Comparison of Different Updating Schemes</title>
<p>To further verify the effectiveness of the two-stage model updating schemes, this paper designs four additional updating schemes for comparison testing:</p>
<p>Option 1: Retrain the model using all labeled samples.</p>
<p>Option 2: Based on the source domain model, randomly select 1500 samples for labeling, and fine-tune the output layer parameters of the model.</p>
<p>Option 3: Use self-supervised learning to initially improve the model&#x2019;s generalization performance. Then, randomly select 1500 unlabeled samples and label them using long-term simulation techniques. Finally, fine-tune the model using the labeled samples.</p>
<p>Option 4: Utilize active learning to screen high-value samples, and subsequently subject the screened 1500 unlabeled samples to long-term simulation to generate labels. These labeled samples are then used to further optimize the model.</p>
<p>Option 5: The two-stage update scheme proposed in this paper.</p>
<p>To test the update effectiveness of each scheme, experiments were conducted on different target domains for the five update schemes, with two performance indicators, <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, selected for comparative analysis, as shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. To further test the update efficiency of the two-stage update scheme, experiments were carried out on target domain F of the IEEE 47-machine 140-bus system using Option 1, Option 4, and Option 5. The comparison benchmark was set at an accuracy rate of 95%, and the update time of the models in the three schemes was compared, as shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Performance comparison of migration solutions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-9.tif"/>
</fig><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of updating time</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Migration plan</th>
<th align="center">Short-term simulation/s</th>
<th align="center">Long-term simulation/s</th>
<th>Model training/s</th>
<th>Total duration/s</th>
</tr>
</thead>
<tbody>
<tr>
<td>Option 1</td>
<td>0</td>
<td>3501.80</td>
<td>1845.58</td>
<td>5347.38</td>
</tr>
<tr>
<td>Option 4</td>
<td>298.49</td>
<td>550.18</td>
<td>326.63</td>
<td>1175.30</td>
</tr>
<tr>
<td>Option 5</td>
<td>298.49</td>
<td>294.74</td>
<td>353.72</td>
<td>946.95</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, although Option 1 can improve the evaluation performance of the model through a large number of training samples and demonstrates similar performance to the proposed update scheme in this paper on the target domain F, according to <xref ref-type="table" rid="table-7">Table 7</xref>, the retraining time required for Option 1 is 5347.38 s, far exceeding that of Option 5, and cannot meet the timely demand for model updates. The performance of Option 2 is inferior to that of Option 5, because when I<sub>MMD</sub> &#x003E; 0.5, the distribution difference between the target domain data and the source domain data is significant, and merely fine-tuning the output layer cannot fully capture the sample characteristics under the new system. Although Option 3 utilizes self-supervised learning in the first stage to enhance the model&#x2019;s generalization performance, the randomly selected labeled samples are not representative, leading to lower update performance across various target domains compared to Option 5. Option 4 employs active learning to update the model, but the performance after the update is significantly lower than that achieved using Option 5.</p>

<p>To verify the impact of sample size on the transfer results, we selected 500 to 3100 key samples from the target domain, with 200 samples as the interval, according to the five updating schemes, and tested them. The results are shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Comparison of migration scheme performance under different training samples</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_62667-fig-10.tif"/>
</fig>
<p>From <xref ref-type="fig" rid="fig-10">Fig. 10</xref>, with just 900 labeled samples, the updated model&#x2019;s accuracy in option 5 is notably higher, outperforming other schemes. This is due to the proposed scheme enhancing model generalization through self-supervised learning in stage 1, and a sample selection mechanism identifying valuable samples. Thus, the model significantly improves with limited labeled samples, validating the two-stage updating scheme&#x2019;s effectiveness.</p>

</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions</title>
<p>This paper presents a dual-tower Transformer-based evaluation model and a two-stage update strategy, aiming to improve the model&#x2019;s multiclass transient assessment performance and update efficiency. Experimental verification was conducted on the two power systems, and the conclusions are as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>The dual-tower Transformer model in this paper effectively extracts features from temporal and variable dimensions, ensuring complete extraction and significantly enhancing accuracy in multiclass transient stability assessment for power systems. Additionally, the hybrid sampling strategy proposed in this paper effectively improves the classification accuracy for minority class samples by balancing the sample distribution.</p></list-item>
<list-item><label>(2)</label><p>This paper employs deep transfer learning, combining self-supervised learning and supervised learning, to achieve efficient model updates. This approach significantly reduces the dependence on external labeled samples during the model update process, greatly saves the time required for generating labeled samples, and improves the efficiency of model updates.</p></list-item>
</list></p>
<p>In this paper, the model update strategy focuses on effectively updating the model using unlabeled samples and a limited number of labeled samples in the target domain, significantly reducing the time cost of model updating. Meanwhile, the source domain boasts abundant labeled sample resources, and some of these samples have similar data distributions to those in the target domain. However, these similarly distributed samples have not been effectively utilized in the model update process. To further enhance the accuracy of model updates, our future research direction will be to leverage these source domain samples with similar distributions to implement sample-level transfer learning. We aim to further improve the accuracy of model updates by incorporating valuable sample information from the source domain.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This project is funded by the National Natural Science Foundation of China (5227-7084).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Nan Li, Jingxiong Dong; data collection: Nan Li, Jingxiong Dong; analysis and interpretation of results: Nan Li, Jingxiong Dong, Liang Huang, Liang Tao; draft manuscript preparation: Nan Li, Jingxiong Dong. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hashim</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Transient stability assessment of a power system using multi-layer SVM method</article-title>. <source>IEEE Trans Power Syst</source>. <year>2021</year>;<volume>35</volume>(<issue>1</issue>):<fpage>821</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPEC51183.2021.9384918</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Real-time transient stability assessment in power system based on improved SVM</article-title>. <source>J Mod Power Syst Clean Energy</source>. <year>2019</year>;<volume>7</volume>(<issue>1</issue>):<fpage>26</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s40565-018-0453-x</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Representational learning approach for power system transient stability assessment based on convolutional neural network</article-title>. <source>J Eng</source>. <year>2017</year>;<volume>2017</volume>(<issue>13</issue>):<fpage>1847</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CIASG.2013.6611493</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Geng</surname> <given-names>G</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Fast transient stability batch assessment using cascaded convolutional neural networks</article-title>. <source>IEEE Trans Power Syst</source>. <year>2019</year>;<volume>34</volume>(<issue>4</issue>):<fpage>2802</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRS.2019.2895592</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Improved deep belief network and model interpretation method for power system transient stability assessment</article-title>. <source>J Mod Power Syst Clean Energy</source>. <year>2019</year>;<volume>8</volume>(<issue>1</issue>):<fpage>27</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.35833/MPCE.2019.000058</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Commented content classification with deep neural network based on attention mechanism</article-title>. In: <conf-name>2017 IEEE 2nd Advanced Information Technology, Electronic and Automation Control Conference (IAEAC)</conf-name>; <year>2017 Mar 25&#x2013;26</year>; <publisher-loc>Chongqing, China</publisher-loc>. p. <fpage>2016</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/IAEAC.2017.8054369</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Kundur</surname> <given-names>P</given-names></string-name></person-group>. <source>Power system stability and control</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>CRC Press</publisher-name>; <year>2007</year>. <fpage>360</fpage> p.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>QH</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Wide-area damping controller of FACTS devices for inter-area oscillations considering communication time delays</article-title>. <source>IEEE Trans Power Syst</source>. <year>2013</year>;<volume>29</volume>(<issue>1</issue>):<fpage>318</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRS.2013.2280216</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Resilient wide-area damping control using GrHDP to tolerate communication failures</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2018</year>;<volume>10</volume>(<issue>3</issue>):<fpage>2547</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2018.2803822</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Transient voltage stability assessment of wind power system based on noisy input multi-class gaussian process</article-title>. <source>South Power Syst Technol</source>. <year>2024</year>;<volume>18</volume>(<issue>9</issue>):<fpage>126</fpage>&#x2013;<lpage>37</lpage>. (In Chinese). doi:<pub-id pub-id-type="doi">10.13648/j.cnki.issn1674-0629.2024.09.014</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ai</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Convolutional neural network-based power system transient stability assessment and instability mode prediction</article-title>. <source>Appl Energy</source>. <year>2020</year>;<volume>263</volume>(<issue>6419</issue>):<fpage>114586</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apenergy.2020.114586</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A transient stability assessment method of power system based on XGboost-DF</article-title>. <source>Electr Meas Instrum</source>. <year>2024</year>;<volume>61</volume>(<issue>10</issue>):<fpage>119</fpage>&#x2013;<lpage>27</lpage>. (In Chinese). doi:<pub-id pub-id-type="doi">10.19753/j.issn1001-390.2024.10.016</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Qiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Adaptive assessment of transient stability for power system based on transfer multi-type of deep learning model</article-title>. <source>Electr Power Autom Equip</source>. <year>2023</year>;<volume>43</volume>(<issue>1</issue>):<fpage>184</fpage>&#x2013;<lpage>92</lpage>. (In Chinese). doi:<pub-id pub-id-type="doi">10.16081/j.epae.202206002</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Dual cost-sensitivity factors based power system transient stability assessment</article-title>. <source>IET Gener Transm Distrib</source>. <year>2020</year>;<volume>14</volume>(<issue>24</issue>):<fpage>5858</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.1049/iet-gtd.2020.0365</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>An</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A data-driven method for transient stability margin prediction based on security region</article-title>. <source>J Mod Power Syst Clean Energy</source>. <year>2020</year>;<volume>8</volume>(<issue>6</issue>):<fpage>1060</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.35833/MPCE.2020.000457</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Data inheritance-based updating method and its application in transient frequency prediction for a power system</article-title>. <source>Int Trans Electr Energy Syst</source>. <year>2019</year>;<volume>29</volume>(<issue>6</issue>):<fpage>e12022</fpage>. doi:<pub-id pub-id-type="doi">10.1002/2050-7038.12022</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A novel data-driven approach for transient stability prediction of power systems considering the operational variability</article-title>. <source>Int J Electr Power Energy Syst</source>. <year>2019</year>;<volume>107</volume>(<issue>3</issue>):<fpage>379</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijepes.2018.11.031</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>N</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A hybrid transfer learning method for transient stability prediction considering sample imbalance</article-title>. <source>Appl Energy</source>. <year>2023</year>;<volume>333</volume>(<issue>2</issue>):<fpage>120573</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apenergy.2022.120573</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Transient stability prediction of time-varying power systems based on inheritance</article-title>. <source>Proc CSEE</source>. <year>2021</year>;<volume>41</volume>(<issue>15</issue>):<fpage>5107</fpage>&#x2013;<lpage>19</lpage>. (In Chinese). doi:<pub-id pub-id-type="doi">10.13334/j.0258-8013.pcsee.200829</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Attention is all you need</article-title>. In: <conf-name>Advances in Neural Information Processing Systems</conf-name>; <year>2017 Dec 4&#x2013;9</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bei</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Transient stability assessment of power systems based on the transformer and neighborhood
rough set</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>2</issue>):<fpage>270</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13020270</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Su</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A data-driven method for online transient stability monitoring with vision-transformer networks</article-title>. <source>Int J Electr Power Energy Syst</source>. <year>2023</year>;<volume>149</volume>(<issue>2</issue>):<fpage>109020</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijepes.2023.109020</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Deep learning based on Transformer architecture for power system short-term voltage stability assessment with class imbalance</article-title>. <source>Renew Sustain Energy Rev</source>. <year>2024</year>;<volume>189</volume>(<issue>4</issue>):<fpage>113913</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.rser.2023.113913</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Using trajectory clusters to define the most relevant features for transient stability prediction based on machine learning method</article-title>. <source>Energies</source>. <year>2016</year>;<volume>9</volume>(<issue>11</issue>):<fpage>898</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en9110898</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Asymptotically optimal one- and two-sample testing with kernels</article-title>. <source>IEEE Trans Inf Theory</source>. <year>2021</year>;<volume>67</volume>(<issue>4</issue>):<fpage>2074</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TIT.2021.3059267</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Self-supervised learning for time series analysis: taxonomy, progress, and prospects</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2024</year>;<volume>46</volume>(<issue>10</issue>):<fpage>6775</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2024.3387317</pub-id>; <pub-id pub-id-type="pmid">38598381</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name>, <string-name><surname>Osindero</surname> <given-names>S</given-names></string-name>, <string-name><surname>Teh</surname> <given-names>YW</given-names></string-name></person-group>. <article-title>A fast learning algorithm for deep belief nets</article-title>. <source>Neural Comput</source>. <year>2006</year>;<volume>18</volume>(<issue>7</issue>):<fpage>1527</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1162/neco.2006.18.7.1527</pub-id>; <pub-id pub-id-type="pmid">16764513</pub-id></mixed-citation></ref>
</ref-list>
</back></article>