<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">EE</journal-id>
<journal-id journal-id-type="nlm-ta">EE</journal-id>
<journal-id journal-id-type="publisher-id">EE</journal-id>
<journal-title-group>
<journal-title>Energy Engineering</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-0118</issn>
<issn pub-type="ppub">0199-8595</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">73912</article-id>
<article-id pub-id-type="doi">10.32604/ee.2025.073912</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Curriculum-Learning-Guided Multi-Agent Deep Reinforcement Learning for N-1 Static Security Prevention and Control</article-title>
<alt-title alt-title-type="left-running-head">Curriculum-Learning-Guided Multi-Agent Deep Reinforcement Learning for N-1 Static Security Prevention and Control</alt-title>
<alt-title alt-title-type="right-running-head">Curriculum-Learning-Guided Multi-Agent Deep Reinforcement Learning for N-1 Static Security Prevention and Control</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zhang</surname><given-names>Ximing</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>zhangxm@csg.cn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Zhuohuan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Quan</surname><given-names>Xuexia</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Cheng</surname><given-names>Kai</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Yu</surname><given-names>Yang</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>China Southern Power Grid Co., Ltd.</institution>, <addr-line>Guangzhou, 510700</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Digital Grid Research Institute Co., Ltd., China Southern Power Grid</institution>, <addr-line>Guangzhou, 510663</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Ximing Zhang. Email: <email>zhangxm@csg.cn</email> or <email>ximing1980@yeah.net</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>06</day><month>08</month><year>2026</year>
</pub-date>
<volume>123</volume>
<issue>9</issue>
<elocation-id>19</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>09</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_EE_73912.pdf"></self-uri>
<abstract>
<p>The &#x201C;N-1&#x201D; criterion represents a fundamental principle for assessing the reliability of power systems in static security analysis. Existing studies mainly rely on centralized single-agent reinforcement learning frameworks, where centralized control is difficult to cope with regional autonomy and communication delays. In high-dimensional state&#x2013;action spaces, these approaches often suffer from low efficiency and unstable policies, limiting their applicability to large-scale grids. To address these issues, this paper proposes a Multi-Agent Deep Reinforcement Learning (MADRL) method enhanced with Curriculum Learning (CL) and Prioritized Experience Replay (PER). The proposed framework adopts a Centralized Training with Decentralized Execution (CTDE) paradigm, where independent agents are assigned to different system regions to enable autonomous decision-making and interregional coordination. In addition, the Actor&#x2013;Critic (AC) architecture is refined with optimized value update rules to mitigate Q-value overestimation. A curriculum learning mechanism based on source&#x2013;load fluctuation intensity further guides agents from simple to complex operating conditions, enhancing convergence and policy robustness. Simulation results on the IEEE 39-bus system demonstrate that the proposed method efficiently generates coordinated multi-region control strategies, eliminates voltage and current violations under N-1 contingencies, and consistently outperforms the baseline MADRL approach in terms of decision performance and robustness under fluctuating source&#x2013;load scenarios.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multi-agent deep reinforcement learning</kwd>
<kwd>static security analysis</kwd>
<kwd>preventive control</kwd>
<kwd>curriculum learning</kwd>
<kwd>N-1 guidelines</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Key technologies for stability analysis and coordinated control of new power systems based on data-mechanism fusion</funding-source>
<award-id>ZBKJXM20232027</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Static security analysis (SSA) is a fundamental approach to ensuring the secure and stable operation of power systems. Over time, SSA has evolved into a mature theoretical framework with established engineering standards. As the core criterion for static security assessment, the N-1 criterion can provide decision-making support in system accident prevention and operational stability improvement [<xref ref-type="bibr" rid="ref-1">1</xref>]. When the system operating point violates the N-1 security boundary, preventive or corrective measures must be implemented to maintain system security [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>At present, static security prevention and control primarily rely on model-driven approaches, where mathematical models and optimization algorithms are employed to derive the optimal control strategy under safety constraints [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. However, this type of method is highly dependent on the precise modeling of the alternating current power flow equations and the complete acquisition of line parameters. With the expansion of the scale of the power system, this method encounters significant challenges in terms of computational efficiency and real-time applicability. To address these issues, data-driven approaches have emerged, offering nonlinear modeling capabilities, reduced dependence on physical models for SSA. However, Deep Reinforcement Learning (DRL) has the advantages of eliminating the need for precise physical models and supports end-to-end decision-making, and has shown good application prospects in power system optimization scheduling [<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>], emergency control [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>] and other fields. Studies have shown that DRL can drive agents to learn control strategies through reward function design, enabling generator output adjustment, load shedding, and other corrective operations to eliminate security risks. For example, [<xref ref-type="bibr" rid="ref-15">15</xref>] achieves the accurate calculation of the power adjustment amount based on Deep Q Network(DQN) and sensitivity analysis; the literature [<xref ref-type="bibr" rid="ref-16">16</xref>] constructs a trend adjustment Markov decision-making process that satisfies the static stability constraints, and achieves the collaborative control of generator action and multi-objective under the framework of parallel DRL; the literature [<xref ref-type="bibr" rid="ref-17">17</xref>] quickly generates an approximate optimal adjustment strategy considering transient stability constraints through distributed Deep Deterministic Policy Gradient (DDPG); the literature [<xref ref-type="bibr" rid="ref-18">18</xref>] uses the PQ decoupling characteristics to propose a dual-agent collaborative architecture to realize the coordinated control of active power and voltage.</p>
<p>However, most of these methods are based on a centralized single-agent framework, which face two major challenges in large-scale systems: (1) centralized control makes it difficult to guarantee communication latency or ensure regional autonomy; and (2) the high-dimensional state&#x2013;action space increases training complexity, often preventing effective convergence. In this context, Multi-Agent Deep Reinforcement Learning (MADRL) has become a promising direction due to its ability of distributed decision-making and global coordination. By decomposing centralized control into a cooperative game among agents, each agent can make independent decisions under a decentralized structure and achieve global coordination through local interaction. This approach not only relieves computational pressure caused by high dimensionality but also better fits the autonomy and coordination needs of large-scale grids. Studies have verified the advantages of MADRL in power system control. For example, Reference [<xref ref-type="bibr" rid="ref-19">19</xref>] proposes a two&#x2014;stage control method combining optimized scheduling and Multi-Agent Deep Deterministic Policy Gradient (MADDPG) for vol/var optimization; Reference [<xref ref-type="bibr" rid="ref-20">20</xref>] formulates multi-regional volt/var coordination as a partially observable Markov game and proposes a robust regional collaborative VVC strategy; Reference [<xref ref-type="bibr" rid="ref-21">21</xref>] develops a centralized training&#x2013;decentralized execution multi-agent voltage control framework under N-1 conditions; Reference [<xref ref-type="bibr" rid="ref-22">22</xref>] employs LSTM-enhanced MADDPG for microgrid frequency regulation under renewable uncertainty.</p>
<p>Although MADRL has demonstrated strong potential in improving the scalability and adaptability of power system control, the baseline MADDPG framework still faces challenges such as Q-value overestimation, unstable convergence, and low sampling efficiency in high-dimensional environments [<xref ref-type="bibr" rid="ref-23">23</xref>]. Compared with other multi-agent reinforcement learning algorithms that rely on value factorization or shared advantage estimation, MADDPG offers a deterministic actor&#x2013;critic structure capable of directly handling continuous control actions. This property is essential for preventive control in large-scale power systems, where generator voltage and power adjustments must satisfy nonlinear physical constraints and maintain high control precision. Discretizing these continuous actions, as required by many value-based algorithms, could reduce control granularity and impair the enforcement of N-1 static security limits.</p>
<p>Building on the advantages of MADDPG, this study focuses on enhancing its training stability and efficiency in complex multi-agent environments. A Curriculum Learning (CL) mechanism [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>] is introduced to guide agents from simple to more challenging N-1 operating conditions, improving adaptability to varying source&#x2013;load fluctuations. Additionally, a Prioritized Experience Replay (PER) mechanism [<xref ref-type="bibr" rid="ref-26">26</xref>] is integrated to emphasize high-value samples during centralized training, thereby accelerating convergence and improving sample utilization. The resulting CL-PER-MADRL framework maintains the physical interpretability of control actions while achieving more stable and efficient policy learning for preventive control tasks in large-scale power systems.</p>
<p>Based on this, this research proposes a multi-agent actor-critic learning framework integrating CL and PER for static security prevention and control of power systems. This approach refines the value update rules to alleviate Q-value overestimation and incorporates progressive learning to guide agents from simple to complex scenarios, thereby improving convergence and policy robustness.</p>
<p>The main contributions of this article are as follows:
<list list-type="simple">
<list-item><label>1.</label><p>A curriculum-learning-guided prioritized replay mechanism is designed to improve convergence efficiency and robustness to source&#x2013;load fluctuations without requiring precise physical modeling.</p></list-item>
<list-item><label>2.</label><p>A distributed control framework based on CTDE in static security prevention and control is developed, effectively solve the problems of dimensional disasters, communication delays and single points of failure, and realize regional autonomy and cross-regional collaboration.</p></list-item>
<list-item><label>3.</label><p>On the basis of the traditional AC architecture, in view of the problem of overestimation of Q value in multi-agent training, an optimized value update rule is proposed, which enhances the stability and control effect of the strategy.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Static Security Prevention and Control Modeling of the Power System</title>
<p>In power system studies, the primary criterion for static security analysis is that, following any N-1 contingency (i.e., the disconnection of a single component), system equipment should not become overloaded and bus voltages must remain within permissible limits. Preventive and corrective control are generally defined as adjustments to the system&#x2019;s operating point to ensure adequate safety and stability margins during normal operation. The basic process of static security prevention and control typically involves modifying the operating point of the system, sequentially disconnecting transmission lines, calculating bus voltages and branch active power flows, and checking whether any violations occur under N-1 contingency conditions. Based on these principles, the static security prevention and control optimization model can be formulated as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>The control variables are the generator&#x2019;s active power and terminal voltage.</p></list-item>
<list-item><label>(2)</label><p>Constraints include both equality and inequality constraints. The equality constraints correspond to the power flow equations after the control variables are applied, as shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref></p></list-item>
</list>
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the active and reactive power injected at bus <italic>i</italic>, respectively; <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the voltage magnitude at bus <italic>i</italic>; <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the voltage phase angle difference, conductance and admittance between bus <italic>i</italic> and <italic>j</italic>, respectively.</p>
<p>The inequality constraints specify the upper and lower bounds of the control variables during the adjustment process, expressed as:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> denote the active power of the synchronous generator <italic>i</italic> and its upper and lower limits; <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> denote the active power variation of synchronous generator <italic>i</italic> and its maximum upward and downward ramping capability; <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> are the reactive power of synchronous generator <italic>i</italic> and its upper and lower limits.
<list list-type="simple">
<list-item><label>(3)</label><p>The control objective is to minimize both the magnitude and number of voltage, current, and power violations in the static security analysis, as represented in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.</p></list-item>
</list>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Specifically:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mo>[</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the sets of branch, bus and equipment apparent power, respectively; <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the bus voltage magnitude, line current magnitude, and equipment apparent power, respectively, under the <italic>k</italic>-th line outage; <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> are their corresponding limits.</p>
<p>Each term in the objective function represents the cumulative violation magnitude of a specific operational variable under all N-1 contingencies. Since all violations are expressed as squared exceedance values normalized by their respective operational limits (e.g., <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>), the overall objective function becomes dimensionless. Equal weighting is applied among the three components <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, ensuring balanced consideration of voltage, current, and apparent power violations without introducing subjective prioritization. This formulation highlights comprehensive static security improvement rather than emphasizing a single type of violation.</p>
<p>In summary, static security preventive control in power systems primarily relies on regulating generator output. A nonlinear optimization model is constructed with power flow equations as equality constraints and equipment operation limits as inequality constraints. However, as the system scale expands and the number of N-1 contingency increases sharply, the model becomes highly nonlinear and computationally intensive, making traditional optimization-based methods difficult to apply in real time.</p>
<p>To address these challenges, this study introduces a MADRL framework to enhance the adaptability and real-time performance of static security preventive control.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Static Security Prevention and Control Method Based on MADRL</title>
<p>In order to enhance the real-time performance of static security preventive control in power systems, this study employs the MADRL approach to accelerate decision-making. Specifically, the power grid is partitioned into regions, with each region assigned an agent responsible for regulating generator outputs; based on the multi-agent Actor&#x2013;Critic framework, an improved training mechanism is developed to enable the agents to jointly learn regulation schemes that satisfy the N-1 security constraints of the entire system.</p>
<p>It should be noted that the partitioning strategy of the system directly affects the distribution characteristics of adjustable units in each region, which in turn influences the convergence and stability of multi-agent training. The algorithm examined in this study assumes that the partitioning is fixed according to the network topology and permits the sharing of state and action information among agents during the centralized training phase to ensure effective learning and coordination of the scheduling scheme.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Distributed Partially Observable Markov Decision Processes</title>
<p>Before applying the MADRL algorithm, the problem is formulated as a distributed partially observable Markov decision process (Dec-POMDP). In this framework, the static security preventive control of the power system is represented as a six-tuple: <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mrow><mml:mo>(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where S is the system state space; <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the local observation set of agent <italic>i</italic>, containing only partial information from the global state; <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the action set of agent <italic>i</italic>; <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the reward function of agent <italic>i</italic>, which satisfies <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>S</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>R</mml:mi></mml:math></inline-formula>; <italic>P</italic> is the state transfer function; and <italic>&#x03B3;</italic> is the discount factor. At each time step <italic>t</italic>, agent <italic>i</italic> observes the operational state <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of its region and selects an action according to its policy <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The system then transitions to the next state <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x223C;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on the current state <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and the joint actions <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of all agents, providing immediate rewards <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msubsup><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to each agent. Subsequently, each agent receives a new local observation and stores the current state, action, reward, and next state in the replay buffer for network updates. This process repeats until a termination condition is met and a training episode is completed.
<list list-type="simple">
<list-item><label>(1)</label><p>The observable state of the agent</p></list-item>
</list></p>
<p>Due to limitations such as communication constraints and privacy protection, each agent&#x2019;s observation is restricted to its local measurement data. Each agent makes decisions solely based on its local state, without requiring complex communication devices to obtain information from other agents. The observable state is defined in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>.
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>
<list list-type="simple">
<list-item><label>(2)</label><p>The action of the agent</p></list-item>
</list></p>
<p>The range of real-time control decision variables is mapped to the interval [&#x2212;1, 1] according to the generator ramping constraint in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, and the mapped variables constitute the agent&#x2019;s action space.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> And <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the generator active power and the change in terminal voltage, respectively.
<list list-type="simple">
<list-item><label>(3)</label><p>Rewards for agents</p></list-item>
</list></p>
<p>The control objective of each agent is to minimize violations of static security constraints. The reward generation process is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. After the joint actions of all agents are executed in the environment, system security constraints under normal conditions are first evaluated, followed by N-1 contingency verification. The resulting immediate reward is then fed back to each agent. When the trend does not converge, a greater negative reward <italic>K</italic><sub>1</sub> is given.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Reward generation process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-1.tif"/>
</fig>
<p>The training goal of MADRL is to find the optimal strategy to maximize the cumulative return. To this end, the static security overstepping problem is transformed into a penalty item and a reward function is introduced. The reward function design is shown in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mspace width="0pt" /><mml:mspace width="0pt" /></mml:mtd><mml:mtd><mml:mrow><mml:mtext>Post-decision power flow non-convergence</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mrow><mml:mtext>Power flow fails to converge for an&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;contingency</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mrow><mml:mtext>Limit violations occur during the&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>contingency analysis</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1.5</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mtext>The system passes all&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>contingency checks without any violations&#xA0;</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the penalty for current non-convergence after agent <italic>i&#x2019;</italic>s action; <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the penalty for non-convergence during the N-1 process; <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the penalty for violations of balancing machines; <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the penalty for violations of regulating units; <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the positive reward assigned when the total number of voltage over-limit <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and current over-limit <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in the system is small after the static security verification, as shown in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>; <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are the penalties for the agent&#x2019;s violation of static security constraints.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the reward coefficient.</p>
<p>The reward function is designed to address the requirements of static security preventive control in power systems, combining a severity measurement method based on utility theory with a discrete reward mechanism to guide agents in accurately perceiving and responding to static security risks. The penalty items <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> for each agent&#x2019;s violation of static security constraints are defined as follows:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>I</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mi>L</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mn>6</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mn>6</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mi>L</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mn>6</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mn>6</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Among them:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext>&#xA0;other</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext>&#xA0;other</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In summary, the reward function consists of three distinct components that guide the learning process. The penalty terms, governed by coefficients <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, serve to penalize power flow non-convergence and operational constraint violations, ensuring all actions yield feasible system states. The guidance reward <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> provides dense, positive feedback for control actions that reduce the number of security violations. Furthermore, a significant sparse reward is granted only when a fully secure state (zero violations) is achieved, collectively steering the agent toward robust N-1 static security compliance.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>MADDPG Algorithm</title>
<p>In order to solve the problem of instability in multi-agent environments, Multi-Agent Deep Deterministic Policy Gradient(MADDPG) adopts a centralized training-decentralized execution framework; in the centralized training phase, each regional agent first observes the regional state information, feeds it into its policy network to generate actions, and then all actions are combined to act on the environment. The environment feeds back the rewards of each agent and transfers them to the next operating state, and at the same time stores the experience in the experience replay buffer. In the update phase, small batches of data are randomly sampled from the experience replay buffer and aggregated into joint experiences <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. When the Critic network is updated, it will consider the other agents&#x2019; policies and actions, and use the joint experience to update the network parameters, so as to achieve coordinated decision-making between agents. The Actor network updates its policy based on the <italic>Q</italic>-value given by the Critic network. After the network update is completed, each agent in the online execution phase makes decisions solely based on local observations, and there is no need to communicate with other agents. This method takes into account the communication restrictions in the actual scenario, making it suitable for the coordinated prevention and control of complex power systems in the sub-regions.</p>
<p>If the environment contains N agents, the system maintains 2N Actor&#x2013;Critic networks. Among them, the policy of the <italic>i</italic>-th agent is denoted as <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, with parameters <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The cumulative expected return of the <italic>i</italic>-th agent is given by:
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Then the deterministic strategy gradient of the <italic>i</italic>-th agent is expressed as:
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the local observation of the <italic>i</italic>-th agent; <italic>D</italic> is the experience replay buffer, which stores the experience trajectories of all agents, and each sample is a quadruple, where: <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>; <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the joint action value function, that is, a centralized state&#x2013;action value function, each sample not only contains the local observation state and the execution action, but also incorporates the state-action information of other agents. This data structure effectively addresses the technical bottlenecks of traditional reinforcement learning in the field of multi-agents, so that agents can not only perceive changes in their own state, but also obtain global action strategies. Through the state-action joint coding in the process of continuous strategy update, the Markov nature of environmental interaction is preserved, and the problem of environmental instability caused by partial observability of traditional methods is effectively mitigated. Even if the strategy is continuously updated, the environment maintains stability. The update of <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is updated according to <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref>:
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mi>Q</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x03BC;</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>y</italic><sub><italic>i</italic></sub> is obtained by <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>:
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msubsup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Improved Actor-Critic Strategy</title>
<p>In the MADDPG algorithm, the target actor network and the target critic network adopt a &#x201C;soft update&#x201D; strategy. This strategy can make the current network too similar to the target network, which hampers the effective separation of action selection from policy evaluation. It can also lead to an excessive pursuit of maximizing long-term discount returns, resulting in overestimation of the true <italic>Q</italic>-value. With multiple rounds of iterative updates, this overestimation tends to accumulate during policy exploration, increasing both deviation and variance. Consequently, the agent&#x2019;s ability to explore globally optimal solutions and make optimal decisions in the current state may be impaired, which may eventually undermine the effectiveness of the learning strategy.</p>
<p>In order to solve overestimation in DRL, this paper incorporates the AC framework from the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm into MADDPG. Two independent critic networks <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are introduced, along with their corresponding target networks <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula>. The minimum of the two <italic>Q</italic>-value estimates is used as the target value. This approach effectively reduces overestimation deviation, thereby improving online learning stability. The <italic>Q</italic>-value calculation in <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref> is therefore replaced by the formulation in <xref ref-type="disp-formula" rid="eqn-21">Eq. (21)</xref>.
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:munder><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msubsup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Importance Sampling</title>
<p>Prioritized Experience Replay (PER) speeds up convergence by prioritizing experiences with large temporal-difference (TD) errors, which are more informative for learning. However, such prioritization introduces a sampling bias, which is corrected by applying importance sampling weights <xref ref-type="disp-formula" rid="eqn-25">Eqs. (25)</xref> and <xref ref-type="disp-formula" rid="eqn-26">(26)</xref> to ensure unbiased gradient estimation.</p>
<p>In the MADDPG framework, the update of the action-value function <italic>Q</italic>(<italic>s,a</italic>) depends on the gradient transfer of the TD error. The TD error directly reflects the discrepancy between the agent&#x2019;s estimated and actual experience values. Based on this characteristic, this study uses |<italic>&#x03B4;</italic>| as a quantitative indicator of the importance of experience, as shown in <xref ref-type="disp-formula" rid="eqn-22">Eq. (22)</xref>:
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mrow><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The larger the <italic>&#x03B4;</italic>, the greater the discrepancy between the target network&#x2019;s prediction and the actual experiential return. The sampling frequency of such experiences should be increased to accelerate convergence between the target network and the current network, thereby improving training efficiency. Therefore, the sampling probability of each experience is defined as <xref ref-type="disp-formula" rid="eqn-23">Eq. (23)</xref>, with its priority given by <xref ref-type="disp-formula" rid="eqn-24">Eq. (24)</xref>.
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="eqn-23">Eq. (23)</xref>: When <italic>&#x03B1;</italic> &#x003D; 0, it corresponds to uniform sampling. In <xref ref-type="disp-formula" rid="eqn-24">Eq. (24)</xref>, <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a small constant that avoids probability collapse. Meanwhile, the TD error of newly added experiences is initialized as <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>Although PER improves learning efficiency by replaying experiences with high TD errors more frequently, it can also lead to data distribution deviations and training instability. Oversampling of high-priority samples changes the empirical data distribution, resulting in a biased gradient expectation during policy evaluation. To address this issue, importance sampling is introduced to compensate for the non-uniform sampling probabilities while maintaining the convergence properties of stochastic gradient descent. The corresponding importance sampling weight formula is shown in <xref ref-type="disp-formula" rid="eqn-25">Eq. (25)</xref>.
<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <italic>N</italic> is the current capacity of the replay buffer, and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the sampling probability, and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is the weight coefficient, which controls the influence of the importance sampling weight in the learning process, and linearly anneals to 1 over the course of training. In order to improve the stability during convergence, the weights are normalized to stabilize the gradient update scale and suppress oscillations during training. Therefore, the <xref ref-type="disp-formula" rid="eqn-25">Eq. (25)</xref> is normalized to obtain the formula:
<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>In the MADDPG algorithm, the parameters of the strategy network depend on the selection of the value network, and the parameters of the value network are updated by the loss function of the value network. Combined with the above-mentioned importance experience playback and improved Actor-Critic strategy, the loss function in <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref> is modified to <xref ref-type="disp-formula" rid="eqn-27">Eq. (27)</xref>:
<disp-formula id="eqn-27"><label>(27)</label><mml:math id="mml-eqn-27" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This section thus incorporates prioritized experience replay with importance sampling into the proposed CTDE-based multi-agent framework. The implementation process and network structure of the algorithm proposed by this research institute are shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The implementation process and network structure configuration of the algorithm proposed in this study</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Curriculum Learning</title>
<p>In the reinforcement learning framework, Curriculum Learning (CL) has emerged as an effective paradigm to improve the efficiency of policy learning. Its core principle lies in organizing training tasks in a &#x201C;from simple to complex&#x201D; manner, thereby guiding agents to progressively acquire stable and efficient decision-making capabilities. In each subsequent stage, load and generator disturbances are gradually increased, enabling the agents to learn from basic stability control to complex system coordination. By structuring the learning environment with progressive levels of difficulty, agents can rapidly acquire fundamental control strategies under low-complexity conditions and subsequently adapt to increasingly challenging scenarios. Through this gradual exposure, the learned policies not only enhance adaptability to multi-source load fluctuations while maintaining robustness and stability under severe disturbances.</p>
<p>In this work, CL is integrated into the preventive control training process of power systems. At the early stage, the focus is placed on ensuring basic stability control, allowing agents to understand and internalize the fundamental operational dynamics of the grid. During the intermediate stages, progressively broader load variations and disturbance conditions are introduced, thereby enhancing policy transferability and generalization. Ultimately, the agents are trained to respond rapidly and reliably to large-scale disturbances within highly complex operating environments, thereby achieving both robustness and practical applicability. This hierarchical design ensures stability and efficiency during training and establishes a solid foundation for subsequent optimization of cooperative control strategies.</p>
<p>To provide an intuitive overview of the proposed training framework, Algorithm 1 summarizes the pseudocode of the CL-PER-MADRL algorithm. This algorithm integrates CL with PER within an MADRL framework. The pseudocode illustrates the key training process, detailing the curriculum-based progression with increasing difficulty levels and the specific performance criteria required to advance to the next stage.</p>
<fig id="fig-10">
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-10.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Example Analysis</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Scene Setup</title>
<p>In this paper, the proposed method is simulated and verified by the IEEE 39-bus transmission system which is partitioned into three regions, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The observation space of each agent includes the active power of the generator, the line load rate, the bus voltage magnitude, and active and reactive power at load buses. The action space is defined as a continuous high-dimensional space for the adjustment of the generator&#x2019;s output, and its adjustment range is constrained to 60%&#x2013;100% of the rated output. During training, seven load levels (85%, 90%, 95%, 100%, 105%, 110%, and 115% of the reference load) are considered, with random source&#x2013;load disturbances applied. Static N-1 verification requires traversing all transmission lines.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Schematic diagram of IEEE 39-bus partition</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-3.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Agent Network Architecture and Parameter Settings</title>
<p>Based on the results of power grid partitioning, a distributed decision-making architecture is developed, where an independent multi-layer fully connected neural network is assigned to each region as the agent model, enabling autonomous regulation and decision-making. The number of neurons in the input layer of the agent actor network corresponding to regions 1, 2, and 3 corresponds to the dimensionality of the regional operating state. The output is the active power of the generator in each region and the adjustment of terminal voltages with the tanh function used as the activation function. Based on experimental evaluation, the number of hidden layers in both the actor and critic networks is set to three; the fully connected layers use the ReLU activation function, and the number of neurons in each layer is (512, 512, 256) for the actor and (1024, 512, 256) for the critic. Each neural network is updated once after every load adjustment. Both networks adopt the Adam optimizer for training, and Gaussian noise is employed for exploration. The specific parameter settings are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Parameter settings</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameters</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Batch size</td>
<td>128</td>
</tr>
<tr>
<td>Learning rates of the actor network</td>
<td>0.0001</td>
</tr>
<tr>
<td>Learning rates of the critic network</td>
<td>0.0003</td>
</tr>
<tr>
<td><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula></td>
<td>0.01</td>
</tr>
<tr>
<td><inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula></td>
<td>0.97</td>
</tr>
<tr>
<td><italic>K</italic><sub>1</sub></td>
<td>2</td>
</tr>
<tr>
<td><italic>K</italic><sub>2</sub></td>
<td>0.04</td>
</tr>
<tr>
<td><italic>K</italic><sub>3</sub></td>
<td>0.004</td>
</tr>
<tr>
<td><italic>K</italic><sub>4</sub></td>
<td>0.01</td>
</tr>
<tr>
<td><italic>K</italic><sub>5</sub></td>
<td>0.000004</td>
</tr>
<tr>
<td><italic>K</italic><sub>6</sub></td>
<td>0.0006</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-2">Table 2</xref> presents the components that exceed operational limits in the static N-1 verification conducted prior to adjustment. Static N-1 verification requires traversing all transmission lines individually; for each contingency, the system is checked for limit violations. The columns report the total number of violations accumulated over all N-1 contingencies: Voltage violations, Current violations, and Slack generator violations respectively denote the number of buses, transmission lines, and slack generator power outputs that exceed permissible limits in any contingency. Numerous voltage, current, and slack generator elements in the system exceed permissible thresholds, thereby violating static security requirements. Therefore, it is necessary to adjust the system operating point by applying preventive control and related measures to ensure static security compliance.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Component overruns after static security verification of the system before adjustment</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Load level</th>
<th>Voltage violations</th>
<th>Current violations</th>
<th>Slack generator violations</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td>0.85</td>
<td>94</td>
<td>253</td>
<td>36</td>
<td>383</td>
</tr>
<tr>
<td>0.9</td>
<td>32</td>
<td>217</td>
<td>36</td>
<td>285</td>
</tr>
<tr>
<td>0.95</td>
<td>5</td>
<td>188</td>
<td>36</td>
<td>229</td>
</tr>
<tr>
<td>1.0</td>
<td>1</td>
<td>152</td>
<td>36</td>
<td>189</td>
</tr>
<tr>
<td>1.05</td>
<td>1</td>
<td>95</td>
<td>34</td>
<td>130</td>
</tr>
<tr>
<td>1.1</td>
<td>1</td>
<td>40</td>
<td>34</td>
<td>75</td>
</tr>
<tr>
<td>1.15</td>
<td>0</td>
<td>12</td>
<td>0</td>
<td>12</td>
</tr>
<tr>
<td>Total</td>
<td>134</td>
<td>957</td>
<td>212</td>
<td>1303</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>The Results of Model Training</title>
<p>The power flow calculation program of the simulation example system, as illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, is developed on the Pandapower platform, which provides the simulation environment for multi-agent training. The proposed algorithm is implemented using PaddlePaddle and is employed to train the multi-agent model. The hardware platform consists of an NVIDIA RTX A4000 GPU, an Intel(R) Core(TM) i7-13700 CPU, and 16 GB of RAM.</p>
<p>During the training process of the MADRL agent, each load is explored 12 times per round, and the power system is statically checked for safety after each current adjustment. The training process adopts intensive CL to reduce the difficulty of training, and the course is divided into four stages for learning. The training scenario settings for each stage are shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Scenario setting for each stage of training</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Training stage</th>
<th>Load variation interval (%)</th>
<th>Load variation interval (%)</th>
<th>Generator disturbance (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Phase I</td>
<td>[0.85, 1.15]</td>
<td>1</td>
<td>5</td>
</tr>
<tr>
<td>Phase II</td>
<td>[0.85, 1.15]</td>
<td>2.5</td>
<td>5</td>
</tr>
<tr>
<td>Phase III</td>
<td>[0.85, 1.15]</td>
<td>2.5</td>
<td>10</td>
</tr>
<tr>
<td>Phase IV</td>
<td>[0.85, 1.15]</td>
<td>2.5</td>
<td>15</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The regional agents in this study are fully cooperative, and the total cumulative rewards obtained by the three agents in each round are shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. Both the reward curve and the actor&#x2013;critic loss converge smoothly, indicating that the policy has reached a stable equilibrium. The total training duration for the proposed MADRL model was approximately 1 h and 55 min. To ensure reproducibility, random seeds were fixed for all experiments, including neural network initialization, noise generation, and environment randomness.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Cumulative rewards of three regional agents and their total reward</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-4.tif"/>
</fig>
<p>In the first stage, the weight parameters of each neural network are randomly initialized, which fluctuate greatly during early training but gradually converge as training progresses. During the first 45 rounds, only the experience buffer was filled without updating the network, resulting in low reward values. From rounds 45 to 100, network updates commenced, causing the reward curve to rise rapidly as the agent gradually learned to satisfy the static security and stability requirements of preventive control measures. At round 140, the training entered the second stage. During the second stage, the load fluctuation range was increased, yet the control success rate remained high, indicating that the proposed MADRL method exhibits strong generalization ability. At round 190, the third stage commenced. The substantial increase in initial generator output disturbance caused the agent to partially exceed limits at the beginning of this stage; however, it stabilized after training. Upon entering the fourth stage, the agent continued to maintain high control performance even under greater disturbance intensities.</p>
<p>To further verify the robustness of the learning framework to hyperparameter variations, an additional experiment was performed in which the learning rate was increased by 20% compared to the value listed in <xref ref-type="table" rid="table-1">Table 1</xref>. The resulting reward curve, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, maintained a smooth and stable convergence pattern similar to the original setting, indicating that the proposed MADRL algorithm exhibits strong robustness and stability against moderate parameter fluctuations.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Reward convergence curves of the proposed MADRL under different learning rates</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-5.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Model Validation and Performance Evaluation</title>
<p>Before validating the proposed MADRL-based preventive control strategy, a conventional Optimal Power Flow (OPF) method was tested under the same system configuration to examine whether traditional optimization can ensure static security compliance. Although the OPF successfully converged to an economically optimal operating point, the resulting solution failed to satisfy N-1 static security criteria. Across multiple load levels, the OPF results exhibited dozens of cumulative voltage and current limit violations after N-1 contingency verification, indicating that OPF-based optimization, which does not explicitly consider contingency constraints, cannot guarantee static security in large-scale grid operations. This highlights the necessity of developing intelligent preventive control approaches that explicitly account for N-1 contingencies.</p>
<p>The effectiveness of the proposed algorithm in enhancing the static security of the power system was evaluated by deploying the trained agents in a simulation environment. During the testing phase, each agent performs real-time control actions based on the grid&#x2019;s operating state under the specified load conditions.</p>
<p>As an illustrative case, at a load level of 85%&#x2014;which represents the most severe limit violations&#x2014;<xref ref-type="fig" rid="fig-6">Figs. 6</xref> and <xref ref-type="fig" rid="fig-7">7</xref> illustrate the bus voltages and line loading rates before and after control through N-1 contingency verification. The results indicate that initially, multiple bus voltages and line currents exceeded their operational limits, met static security requirements. Specifically, the lower voltage limit was violated 94 times, and current overloads occurred 253 times. When line 14&#x2013;15 were disconnected, the voltage at bus 14 dropped to 0.874 p.u., while the load rate of line 2&#x2013;3 peaked at 308%. clearly indicates static insecurity.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Bus voltage and line load ratio before control</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Bus voltage and line load ratio after control</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-7.tif"/>
</fig>
<p>Specifically, at this load level, the overlimit conditions summarized in <xref ref-type="table" rid="table-4">Table 4</xref> required five consecutive control adjustments to restore static security, with N-1 verification conducted after each adjustment of the system operating point. Through progressive tuning of operating points, the system gradually approached compliance with static security standards. Specifically, after the first adjustment, the removal of line 14&#x2013;15 reduced the peak load rate of lines 2&#x2013;3 from 308% to 258% (a 16.2% decrease); the minimum voltage at bus 14 increased from 0.874 p.u. to 0.916 p.u. (a 4.81% increase), with no overvoltage observed. After the second adjustment, the load rate of line 2&#x2013;3 further decreased to 214%, and the voltage at bus 14 increased to 0.974 p.u. marking the first time all bus voltages satisfied safety limits. After the third adjustment, lines 3&#x2013;4 were removed, shifting the overload-dominant line to lines 5&#x2013;10 (load rate of 147%), while all bus voltages remained within safe limits. Following the fourth adjustment, lines 14&#x2013;15 and 2&#x2013;3 were identified as key overloaded lines (132%), with voltage levels remaining within normal limits. After the fifth adjustment, N-1 verification confirmed the elimination of equipment overloads and voltage violations, indicating that static security had been fully restored.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Overlimit situation during light load level regulation</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Load level 0.85</th>
<th>Voltage violations</th>
<th>Current violations</th>
<th>Slack generator violations</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td>Before adjustment</td>
<td>94</td>
<td>253</td>
<td>36</td>
<td>0</td>
</tr>
<tr>
<td>1st adjustment</td>
<td>10</td>
<td>196</td>
<td>36</td>
<td>0</td>
</tr>
<tr>
<td>2nd adjustment</td>
<td>0</td>
<td>94</td>
<td>36</td>
<td>0</td>
</tr>
<tr>
<td>3rd adjustment</td>
<td>0</td>
<td>21</td>
<td>36</td>
<td>0</td>
</tr>
<tr>
<td>4th adjustment</td>
<td>0</td>
<td>3</td>
<td>34</td>
<td>0</td>
</tr>
<tr>
<td>5th adjustment</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
</tr>
</tbody>
</table>
</table-wrap>
 
<p>After completing training, the agent was deployed to evaluate each typical load condition. The results indicate that under various random operating scenarios, line outages&#x2014;whether due to failures or maintenance&#x2014;did not result in system voltages or currents exceeding operational limits.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Comparative Analysis of Algorithm Effectiveness</title>
<p>To verify the effectiveness of the proposed MADRL-based static security preventive control method, MADDPG was employed for comparison, using the fourth stage as the target training benchmark. <xref ref-type="fig" rid="fig-8">Fig. 8</xref> presents the total reward curves of agents for each algorithm. As shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the MADDPG reward curve exhibits substantial fluctuations, indicating poor performance across varying load and source fluctuation scenarios.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Comparison of the proposed MADRL against the baseline MADDPG on total cumulative rewards</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-8.tif"/>
</fig>
<p>The proposed method also began neural network updates at round 45 but incorporated a PER mechanism to prioritize under-learned experiences via importance sampling, thereby accelerating strategy convergence and enabling the agent to better extract local observation features and retain historical state information. Furthermore, the improved Actor-Critic network structure, combined with a well-designed curriculum, allows the agent to initially accumulate successful experiences in stable control under low disturbance conditions, facilitating a smooth transition as task difficulty gradually increases, and enabling the early-learned control strategies and features to transfer to more complex scenarios. This shallow-to-deep training approach effectively enhances strategy adaptability and robustness across varying load and source fluctuation conditions, maintaining high control performance under severe disturbances. Compared to direct training on challenging tasks, the proposed method substantially improves training stability and convergence speed, mitigates excessive reward fluctuations, and enhances strategy adaptability and robustness under varying disturbance intensities.</p>
<p><xref ref-type="fig" rid="fig-9">Fig. 9</xref> depicts the cumulative number of N-1 limit violations for each algorithm following sequential load adjustments in each round. As illustrated, the proposed method significantly reduces limit violations after network updates, while maintaining a consistent downward trend as task complexity increases. These results demonstrate that the CL strategy effectively reduces static security constraint violations in complex disturbance environments and fosters a more robust multi-round control strategy.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Algorithm comparison of cumulative N-1 violations during sequential load adjustments</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="EE_73912-fig-9.tif"/>
</fig>
<p>To evaluate effectiveness, 1000 random scenarios from the fourth training stage were selected for testing and compared with the baseline MADDPG algorithm. Evaluation metrics comprised the success rate of static security preventive control, the average decision time per step, and the total number of agent decisions. Simulation results on the IEEE 39-bus system are summarized in <xref ref-type="table" rid="table-5">Table 5</xref>. Results indicate that the proposed method achieves a preventive control success rate of 99.2%, significantly surpassing the 96.2% achieved by MADDPG. Moreover, the proposed reinforcement learning framework exhibits superior performance in both control efficiency and cross-scenario generalization.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of different algorithms for control</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Control strategy</th>
<th>Success rate</th>
<th>Decision time/ms</th>
<th>Action total number of times</th>
</tr>
</thead>
<tbody>
<tr>
<td>MADDPG</td>
<td>96.20%</td>
<td>2.20</td>
<td>3882</td>
</tr>
<tr>
<td>Proposed method</td>
<td>99.20%</td>
<td>2.16</td>
<td>3770</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>The proposed MADRL-based preventive control framework can be deployed within modern power system operation architectures as an intelligent auxiliary control layer. In practice, the trained agents can operate at the scheduling or regional control level, providing rapid preventive control recommendations based on real-time grid measurements. Because agent inference involves only lightweight matrix operations, the response time is within milliseconds, satisfying the timeliness requirement for short-term preventive control.</p>
<p>In large-scale power grids, communication delays and data acquisition latency may affect decision timeliness. These issues can be mitigated through a distributed control architecture in which regional agents operate independently but coordinate through shared state variables. This hierarchical structure aligns with the existing multi-level dispatching system in power networks. Future work will focus on integrating the MADRL framework with real-time measurement data (e.g., PMU/EMS) and improving robustness against communication uncertainty and data delays.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This paper proposes a curriculum-guided multi-agent deep reinforcement learning framework for static security preventive control of power systems. The method leverages centralized training with distributed execution to enable model-free, online safety regulation. By introducing CL, the instability issues of conventional MADDPG in complex grid scenarios are mitigated, significantly enhancing training robustness. Simulation results on the IEEE 39-bus system demonstrate that the proposed approach can effectively improve the static security compliance rate under N-1 contingencies and maintain stable performance under source&#x2013;load fluctuations. These findings highlight its potential for practical application in distributed, real-time decision-making. Future research will consider extending the framework to larger power system models and incorporating multiple types of uncertainties.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by the China Southern Power Grid Co., Ltd. &#x201C;Key technologies for stability analysis and coordinated control of new power systems based on data-mechanism fusion&#x201D; (ZBKJXM20232027).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Ximing Zhang conceived the study, designed the methodology, conducted simulations, and managed the project. Zhuohuan Li was responsible for data analysis, curation, visualization, and writing the original draft. Xuexia Quan contributed to validation, reviewing, and editing. Kai Cheng provided critical resources and assisted with formal analysis. Yang Yu contributed to investigation and data collection. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are available within the article. And the additional data that support the findings of this study are available on request from the corresponding author, upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alizadeh</surname> <given-names>MI</given-names></string-name>, <string-name><surname>Usman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Capitanescu</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Envisioning security control in renewable dominated power systems through stochastic multi-period AC security constrained optimal power flow</article-title>. <source>Int J Electr Power Energy Syst</source>. <year>2022</year>;<volume>139</volume>(<issue>2</issue>):<fpage>107992</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijepes.2022.107992</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Capitanescu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Martinez Ramos</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Panciatici</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kirschen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Marano Marcolini</surname> <given-names>A</given-names></string-name>, <string-name><surname>Platbrood</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>State-of-the-art, challenges, and future trends in security constrained optimal power flow</article-title>. <source>Electr Power Syst Res</source>. <year>2011</year>;<volume>81</volume>(<issue>8</issue>):<fpage>1731</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.epsr.2011.04.003</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Security-constrained hybrid optimal energy flow model of multi-energy system considering N-1 component failure</article-title>. <source>J Energy Storage</source>. <year>2023</year>;<volume>64</volume>(<issue>1</issue>):<fpage>107060</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.est.2023.107060</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Improved correction strategy for power flow control based on multi-machine sensitivity analysis</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>82391</fpage>&#x2013;<lpage>403</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2020.2989927</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sennewald</surname> <given-names>T</given-names></string-name>, <string-name><surname>Linke</surname> <given-names>F</given-names></string-name>, <string-name><surname>Westermann</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Preventive and curative actions by meshed bipolar HVDC-overlay-systems</article-title>. <source>IEEE Trans Power Deliv</source>. <year>2020</year>;<volume>35</volume>(<issue>6</issue>):<fpage>2928</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRD.2020.3011733</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Poyrazoglu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Optimal topology control with physical power flow constraints and N-1 contingency criterion</article-title>. <source>IEEE Trans Power Syst</source>. <year>2015</year>;<volume>30</volume>(<issue>6</issue>):<fpage>3063</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRS.2014.2379112</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Day-ahead optimal dispatch method of integrated electric-heat-cool-gas energy system based on N-1 safety criterion</article-title>. <source>Energy Build</source>. <year>2024</year>;<volume>323</volume>(<issue>3201&#x2013;5</issue>):<fpage>114800</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.enbuild.2024.114800</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Heidarifar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Andrianesis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ruiz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Caramanis</surname> <given-names>MC</given-names></string-name>, <string-name><surname>Paschalidis</surname> <given-names>IC</given-names></string-name></person-group>. <article-title>An optimal transmission line switching and bus splitting heuristic incorporating AC and N-1 contingency constraints</article-title>. <source>Int J Electr Power Energy Syst</source>. <year>2021</year>;<volume>133</volume>(<issue>3</issue>):<fpage>107278</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijepes.2021.107278</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Intelligent power grid load transferring based on safe action-correction reinforcement learning</article-title>. <source>Energy Eng</source>. <year>2024</year>;<volume>121</volume>(<issue>6</issue>):<fpage>1697</fpage>&#x2013;<lpage>711</lpage>. doi:<pub-id pub-id-type="doi">10.32604/ee.2024.047680</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Enhanced deep reinforcement learning strategy for energy management in plug-in hybrid electric vehicles with entropy regularization and prioritized experience replay</article-title>. <source>Energy Eng</source>. <year>2024</year>;<volume>121</volume>(<issue>12</issue>):<fpage>3953</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.32604/ee.2024.056705</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Coordinated scheduling of electric-hydrogen-heat trigeneration system for low-carbon building based on improved reinforcement learning</article-title>. <source>Energy Eng</source>. <year>2025</year>;<volume>122</volume>(<issue>11</issue>):<fpage>4561</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.32604/ee.2025.067574</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep reinforcement learning-based active network management and emergency load-shedding control for power systems</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2024</year>;<volume>15</volume>(<issue>2</issue>):<fpage>1423</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2023.3302846</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Distributed hierarchical deep reinforcement learning for large-scale grid emergency control</article-title>. <source>IEEE Trans Power Syst</source>. <year>2024</year>;<volume>39</volume>(<issue>2</issue>):<fpage>4446</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRS.2023.3298486</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Distributional deep reinforcement learning-based emergency frequency control</article-title>. <source>IEEE Trans Power Syst</source>. <year>2022</year>;<volume>37</volume>(<issue>4</issue>):<fpage>2720</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPWRS.2021.3130413</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>LJ</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>XP</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>TQ</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>XD</given-names></string-name></person-group>. <article-title>Active power security correction method for power grids based on deep reinforcement learning algorithm</article-title>. <source>Power Syst Prot Control</source>. <year>2022</year>;<volume>50</volume>:<fpage>114</fpage>&#x2013;<lpage>22</lpage> (In Chinese).</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Automatic adjustment method of power flow calculation convergence for large-scale power grid based on knowledge experience and deep reinforcement learning</article-title>. In: <conf-name>Proceedings of the 2020 IEEE 4th Conference on Energy Internet and Energy System Integration (EI2); 2020 Oct 30&#x2013;Nov 1</conf-name>; <publisher-loc>Wuhan, China</publisher-loc>. p. <fpage>694</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ei250167.2020.9346831</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Distributed deep reinforcement learning-based approach for fast preventive control considering transient stability constraints</article-title>. <source>CSEE J Power Energy Syst</source>. <year>2021</year>;<volume>9</volume>(<issue>1</issue>):<fpage>197</fpage>&#x2013;<lpage>208</lpage>. doi:<pub-id pub-id-type="doi">10.17775/CSEEJPES.2020.04610</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>BY</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Han</surname> <given-names>XQ</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Static security preventive control method of power systems based on dual-agent deep reinforcement learning</article-title>. <source>Proc CSEE</source>. <year>2023</year>;<volume>43</volume>:<fpage>1818</fpage>&#x2013;<lpage>30</lpage>. (In Chinese). doi:<pub-id pub-id-type="doi">10.13334/j.0258-8013.pcsee.220005</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Two-stage volt/var control in active distribution networks with multi-agent deep reinforcement learning method</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2021</year>;<volume>12</volume>(<issue>4</issue>):<fpage>2903</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2021.3052998</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>ZY</given-names></string-name></person-group>. <article-title>Robust regional coordination of inverter-based volt/var control via multi-agent deep reinforcement learning</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2021</year>;<volume>12</volume>(<issue>6</issue>):<fpage>5420</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2021.3104139</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Diao</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A data-driven multi-agent autonomous voltage control framework using deep reinforcement learning</article-title>. <source>IEEE Trans Power Syst</source>. <year>2020</year>;<volume>35</volume>(<issue>6</issue>):<fpage>4644</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpwrs.2020.2990179</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hatziargyriou</surname> <given-names>ND</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Load frequency control of multi-microgrids based on deep deterministic policy gradient integrated with online learning</article-title>. <source>IEEE Trans Smart Grid</source>. <year>2025</year>;<volume>16</volume>(<issue>5</issue>):<fpage>4266</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSG.2025.3587312</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A policy gradient algorithm to alleviate the multi-agent value overestimation problem in complex environments</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>23</issue>):<fpage>9520</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23239520</pub-id>; <pub-id pub-id-type="pmid">38067892</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A survey on curriculum learning</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2022</year>;<volume>44</volume>(<issue>9</issue>):<fpage>4555</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2021.3069908</pub-id>; <pub-id pub-id-type="pmid">33788677</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Diao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Fast-converging deep reinforcement learning for optimal dispatch of large-scale power systems under transient security constraints</article-title>. <source>J Mod Power Syst Clean Energy</source>. <year>2024</year>;<volume>13</volume>(<issue>5</issue>):<fpage>1495</fpage>&#x2013;<lpage>506</lpage>. doi:<pub-id pub-id-type="doi">10.35833/mpce.2024.000624</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>X</given-names></string-name>, <string-name><surname>Song</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Prioritized experience replay based on dynamics priority</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>6014</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-56673-3</pub-id>; <pub-id pub-id-type="pmid">38472457</pub-id></mixed-citation></ref>
</ref-list>
</back></article>