<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">64849</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.064849</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Multi-Agent Reinforcement Learning for Moving Target Defense Temporal Decision-Making Approach Based on Stackelberg-FlipIt Games</article-title>
<alt-title alt-title-type="left-running-head">Multi-Agent Reinforcement Learning for Moving Target Defense Temporal Decision-Making Approach Based on Stackelberg-FlipIt Games</alt-title>
<alt-title alt-title-type="right-running-head">Multi-Agent Reinforcement Learning for Moving Target Defense Temporal Decision-Making Approach Based on Stackelberg-FlipIt Games</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Sun</surname><given-names>Rongbo</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Fei</surname><given-names>Jinlong</given-names></name><email>aston_deta@163.com</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Zhu</surname><given-names>Yuefei</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Guo</surname><given-names>Zhongyu</given-names></name></contrib>
<aff id="aff-1"><institution>Key Laboratory of Cyberspace Security, Ministry of Education</institution>, <addr-line>Zhengzhou, 450001</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jinlong Fei. Email: <email>aston_deta@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>03</day><month>07</month><year>2025</year>
</pub-date>
<volume>84</volume>
<issue>2</issue>
<fpage>3765</fpage>
<lpage>3786</lpage>
<history>
<date date-type="received">
<day>25</day>
<month>2</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>5</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_64849.pdf"></self-uri>
<abstract>
<p>Moving Target Defense (MTD) necessitates scientifically effective decision-making methodologies for defensive technology implementation. While most MTD decision studies focus on accurately identifying optimal strategies, the issue of optimal defense timing remains underexplored. Current default approaches&#x2014;periodic or overly frequent MTD triggers&#x2014;lead to suboptimal trade-offs among system security, performance, and cost. The timing of MTD strategy activation critically impacts both defensive efficacy and operational overhead, yet existing frameworks inadequately address this temporal dimension. To bridge this gap, this paper proposes a Stackelberg-FlipIt game model that formalizes asymmetric cyber conflicts as alternating control over attack surfaces, thereby capturing the dynamic security state evolution of MTD systems. We introduce a belief factor to quantify information asymmetry during adversarial interactions, enhancing the precision of MTD trigger timing. Leveraging this game-theoretic foundation, we employ Multi-Agent Reinforcement Learning (MARL) to derive adaptive temporal strategies, optimized via a novel four-dimensional reward function that holistically balances security, performance, cost, and timing. Experimental validation using IP address mutation against scanning attacks demonstrates stable strategy convergence and accelerated defense response, significantly improving cybersecurity affordability and effectiveness.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Cyber security</kwd>
<kwd>moving target defense</kwd>
<kwd>multi-agent reinforcement learning</kwd>
<kwd>security metrics</kwd>
<kwd>game theory</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62302520</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The pervasive digitization of societal infrastructure over two decades has rendered globally distributed entities profoundly interconnected and interdependent, with cyberspace emerging as a contested domain for adversarial engagements. Contemporary network systems increasingly exhibit automation, intelligence, and integration. Paradoxically, sophisticated attacks often exploit simple tools/vectors, whereas defenses necessitate complex coordinated mechanisms&#x2014;an imbalance exacerbated by attackers&#x2019; asymmetric advantages in reconnaissance and persistence.</p>
<p>Inspired by principles of dynamism, randomization, determinism, and diversity, Moving Target Defense (MTD) dynamically manipulates system configurations (e.g., via <italic>shuffling</italic>, <italic>diversity</italic>, and <italic>redundancy</italic>) to reduce attack surface predictability, forcing adversaries into perpetual reconnaissance cycles [<xref ref-type="bibr" rid="ref-1">1</xref>]. MTD research clusters around three axes: &#x201C;<italic>what to move</italic>&#x201D; (spatial decisions), &#x201C;<italic>when to move</italic>&#x201D; (temporal decisions), and &#x201C;<italic>how to move</italic>&#x201D; (technical implementation). The temporal decision-making of MTD aims to select the optimal temporal strategy for transitioning an MTD system from its current state to a new state, thereby invalidating the information or progress acquired by attackers under the current state while ensuring high defensive efficacy and low implementation costs. Consequently, determining how to identify the optimal MTD temporal strategy through balancing system security and availability to guarantee both operational efficiency and defensive effectiveness constitutes a critical research challenge. The primary objective of MTD is to eliminate attackers&#x2019; asymmetric temporal advantage, underscoring that temporal strategy represents an essential component of MTD. Key temporal parameters&#x2014;including the interval between consecutive MTD activations, the duration required for MTD deployment, the time attackers allocate to executing their strategies, and the time required for successful attacks&#x2014;exert significant influence on the overall effectiveness of MTD implementations. Zhuang et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] provided a preliminary exploration of MTD temporal strategies by outlining fundamental issues and hypotheses. They posited that defining the probing surface and attack surface through temporal and spatial MTD strategies could better model the dynamic characteristics of MTD systems. However, their work did not establish a concrete theoretical framework for temporal decision-making in MTD systems.</p>
<p>However, the temporal effectiveness of MTD strategies critically influences the operational success. Overly aggressive MTD activation risks system instability or service degradation due to synchronization overhead, resource contention, or interrupted legitimate workflows. Conversely, insufficiently reactive MTD intervals grant attackers extended time windows to analyze system patterns, exploit vulnerabilities, or escalate privileges, ultimately undermining defense objectives. To balance these trade-offs, Clark et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed a time-based decoy deception detection technique, where a virtual network composed of decoys records temporal log information such as query and response times for attempted node connections. By analyzing the response times of nodes to probing packets, node types could be identified, and they derived a closed-form solution for the expected detection time. While stochastic approaches obscure predictable attack surfaces and complicate adversarial time-based reconnaissance, existing works struggle to rigorously harmonize defensive efficacy with operational efficiency, which necessitates advanced decision frameworks that holistically integrate temporal dynamics, adversarial behavior models, and system constraints. Therefore, this paper addresses MTD timing optimization via a Stackelberg-FlipIt game framework integrated with multi-agent Win or Learn Fast Policy Hill-Climbing (WoLF-PHC) algorithm. Key contributions include:
<list list-type="order">
<list-item>
<p><bold>Abstracting the Network Attack-Defense Process as a Stackelberg-FlipIt Game and Introducing Belief Factors to Control MTD Temporal Decisions.</bold> Tan et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] modeled the MTD temporal decision-making process as a FlipIt game, which effectively depicted the process of alternating control of the attack surface between the attacker and the defender, so as to integrate the temporal strategy into the game decision-making process. However, their design of the game is not accurate enough, ignoring the problem of action sequence and information asymmetry in the real game. Therefore, this paper introduces a Stackelberg game framework, constructing a Stackelberg-FlipIt game model based on the integration of Stackelberg and FlipIt games. By designing and incorporating belief factors to characterize the estimation differences between attackers and defenders regarding defense thresholds, the model more accurately reflects the asymmetric nature of real-world network attack-defense scenarios. This abstraction precisely captures the dynamics of actual network environments, facilitating the dynamic learning and updating of MTD decision-making methods and optimizing the judgment of triggering timing.</p></list-item>
<list-item>
<p><bold>Solving the Game Equilibrium by Using Multi-Agent WoLF-PHC Algorithm.</bold> Existing works [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>] tended to solve the equilibrium of their modeled temporal game by purely mathematical equations, such as Min-Max solving and dynamic programming. However, with the sudden increase of the dimension of the policy space, it is difficult for this kind of solutions to converge in an acceptable time and may only converge to the suboptimal strategies. Therefore, this paper tends to utilize a multi-agent reinforcement learning framework to simulate the game solving process. The WoLF-PHC algorithm is an adaptive reinforcement learning algorithm [<xref ref-type="bibr" rid="ref-7">7</xref>] with low computational complexity, strong convergence properties, and the ability to adapt to dynamic environments. It can quickly learn and adjust strategies, making it highly suitable for MTD temporal decision-making research based on the Stackelberg-FlipIt game. Under the multi-agent framework, defenders and attackers act as independent agents, optimizing their respective strategies through interactive learning. Defenders can dynamically adjust MTD triggering times based on the attackers&#x2019; strategies, thereby achieving more efficient defense.</p></list-item>
<list-item>
<p><bold>Defining a Comprehensive Reward Function.</bold> The aforementioned methods are too simplistic to evaluate the benefits of the strategy, and most of them are evaluated based on the cost of the strategy, ignoring the benefits brought by the strategy in terms of security and other aspects. Therefore, this paper designs a reward function based on Security, Performance, Affordability, and Belief Error to guide the learning direction of the agents. The function references existing research frameworks [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>], ensuring the accuracy and generalizability of reward quantification. It comprehensively balances the costs, defensive effects, and security of different MTD strategies while ensuring the optimal triggering timing.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews the relevant research progress in MTD decision-making methods; <xref ref-type="sec" rid="s3">Section 3</xref> details the Stackelberg-FlipIt model and the multi-agent WoLF-PHC algorithm; <xref ref-type="sec" rid="s4">Section 4</xref> validates the feasibility and effectiveness of the proposed method through a case study on IP address dynamic hopping against scanning attacks; <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper and discusses future directions for improvement.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Decision-oriented MTD research focuses on two pivotal questions: <italic>&#x201C;What to move&#x201D;</italic> and <italic>&#x201C;When to move&#x201D;</italic>. Current studies predominantly employ game theory&#x2014;a mathematical framework for analyzing strategic interactions among rational agents [<xref ref-type="bibr" rid="ref-11">11</xref>]&#x2014;due to two inherent characteristics of MTD scenarios:
<list list-type="order">
<list-item>
<p><bold>Resource Antagonism:</bold> Attackers exploit system vulnerabilities to expand breach surfaces, while defenders constrain exposures via dynamic configuration shifts (e.g., randomization, diversification) [<xref ref-type="bibr" rid="ref-12">12</xref>].</p></list-item>
<list-item>
<p><bold>Interdependent Decision-Making:</bold> The efficacy of adversarial strategies hinges on mutual behavioral adaptations [<xref ref-type="bibr" rid="ref-13">13</xref>].</p></list-item>
</list></p>
<p>These characteristics of MTD attack-defense interactions align with the features of game theory. Consequently, a significant number of game-theoretic approaches have been employed to develop MTD solutions: using game theory to model specific MTD attack-defense processes, proving equilibrium convergence, and ultimately deriving equilibrium-based game strategies [<xref ref-type="bibr" rid="ref-14">14</xref>]. Relevant research can be primarily categorized as follows:
<list list-type="order">
<list-item>
<p><bold>Game Theory-Based MTD Spatial Decision-Making Methods:</bold> MTD spatial decision-making methods are divided into five main categories based on <italic>&#x201C;what to move&#x201D;</italic>: instruction layer, data layer, network layer, platform layer, and runtime environment layer. For example, Ge et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed a game-theoretic MTD approach based on server migration and user service mapping to enhance system real-time performance and throughput elasticity. Carter et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] utilized game theory to derive optimal migration strategies across platforms. Manadhata et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] introduced a two-player stochastic game model, employing the concept of subgame perfect equilibrium to determine optimal MTD strategies based on attack surface diversification. However, due to the single-trigger mechanism of these methods, which only initiate countermeasures upon detecting an attack signal, they often leave room and time for attackers to deploy C&#x0026;C infrastructure. Further temporal methodology refinements remain imperative.</p></list-item>
<list-item>
<p><bold>Game Theory-Based MTD Temporal Decision-Making Methods:</bold> Based on &#x201C;<italic>when to move</italic>&#x201D;, MTD temporal decision-making can be classified into three main categories: time-driven MTD, event-driven MTD, and hybrid time and event-driven MTD.
<list list-type="bullet">
<list-item>
<p>In time-driven MTD decision-making, Li et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] modeled the defender&#x2019;s joint migration and temporal decisions as a semi-Markov decision process. However, their approach is based on a static Stackelberg game, which significantly deviates from real-world network attack-defense dynamics. Moreover, the derived temporal strategies consider only a single factor, achieving only approximate optimality.</p></list-item>
<list-item>
<p>In event-driven MTD decision-making, most existing MTD spatial decision methods and few temporal decision methods rely on specific attack events to trigger defenses. Zhang et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] established an MTD temporal decision-making framework based on the Stochastic Markov Differential Game, which characterizes the continuous-time randomness triggered by the strategy through the It&#x014D; process, focusing on the multi-stage MTD offensive-defensive process and solved the equilibrium by dynamic programming equation. However, the game is under the condition of perfect information, which can hardly be applied to real-world attack and defense process. This kind of method suffers from low accuracy in attack event identification, potential misjudgments, and the risk of defensive responses lagging behind attack deployments. Consequently, such methods fail to ensure both low overhead and security for the maintained systems.</p></list-item>
<list-item>
<p>In hybrid time and event-driven MTD decision-making, Tan et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed an MTD temporal model based on multidimensional transitions of the attack surface. They then integrated this model with the system security state evolution of the FlipIt game [<xref ref-type="bibr" rid="ref-18">18</xref>] to establish an MTD spatiotemporal decision model. They used the decision model to derive an optimal spatiotemporal defense strategy selection algorithm. However, this algorithm relies on mathematical difference methods for solving and cannot output adaptive decisions based on environmental changes.</p></list-item>
</list></p></list-item>
</list></p>
<p>In summary, while game theory-based MTD methods have made significant progress in both spatial and temporal decision-making, challenges remain in terms of adaptability, accuracy, and real-time responsiveness. Further research is needed to address these limitations and enhance the practical applicability of MTD solutions. In <xref ref-type="table" rid="table-1">Table 1</xref>, we compare our methods with existing MTD temporal decision-making methods in order to distinguish our game model and equilibrium algorithm.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison with existing MTD temporal decision-making methods</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">References</th>
<th align="center">Information</th>
<th align="center">Game category</th>
<th align="center">Equilibrium algorithm</th>
<th align="center">Payoff configuration</th>
<th align="center">Temporal decision pattern</th>
</tr>
</thead>
<tbody>
<tr>
<td>Li et al. [<xref ref-type="bibr" rid="ref-5">5</xref>]</td>
<td>Imperfect</td>
<td>Bayesian-Stackelberg game</td>
<td>Min-Max</td>
<td>Migration cost</td>
<td>Time-driven</td>
</tr>
<tr>
<td>Zhang et al. [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Perfect</td>
<td>Markov differential game</td>
<td>Dynamic programming</td>
<td>Attack surface resource rate and cost</td>
<td>Event-driven</td>
</tr>
<tr>
<td>Tan et al. [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>Perfect</td>
<td>FlipIt game</td>
<td>Saddle point</td>
<td>Time and cost</td>
<td>Hybrid Time and Event-driven</td>
</tr>
<tr>
<td>Our method</td>
<td>Imperfect</td>
<td>Stackelberg-FlipIt game</td>
<td>Multi-agent WoLF-PHC</td>
<td>Security, performance, and affordability</td>
<td>Hybrid Time and Event-driven</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<label>3</label>
<title>Model and Methodology</title>
<p>In the real network attack and defense, there is a time series dependence and mutual stimulation of the actions of the two sides: (i) defense triggers may expose defense logic and induce attackers to adjust their strategies, and (ii) attack heuristics may reveal blind spots of the defense system and drive defense strategy iterations. However, existing MTD temporal decision methods often rely on simple random triggers or fixed intervals, neglecting the above dynamic interactions between attackers and defenders, leading to an approximately static defense logic: MTD actions are not directly coupled with the attack behavior, and cannot be optimized based on the attacker&#x2019;s real-time tentative feedback. Besides, information asymmetry is not exploited. The defender does not strategically hide the temporal features of MTD decisions, which reduces the learning cost of the attacker. To address this, we propose a Stackelberg-FlipIt game model, where the defender, as the &#x201C;leader&#x201D;, plans defenses in advance, and the attacker, as the &#x201C;follower&#x201D;, adjusts strategies dynamically. This model better reflects real-world scenarios, enhancing defense precision and effectiveness.</p>
<p>We also extend a multi-agent reinforcement learning framework. Both attackers and defenders use reinforcement learning to optimize their temporal strategies. Defenders minimize costs and maximize effectiveness, while attackers maximize rewards. This approach balances dynamic behaviors and fine-tunes network attack-defense processes.</p>
<p>The network attack-defense processes discussed in this paper are constructed based on the Cyber Kill Chain model to establish a Stackelberg-FlipIt game in cyberspace. Both attackers and defenders are intelligent agents utilizing reinforcement learning algorithms, capable of making decisions by observing the environment. The multi-agent game relies on reinforcement learning algorithms to solve for the game&#x2019;s equilibrium points [<xref ref-type="bibr" rid="ref-19">19</xref>]. <xref ref-type="sec" rid="s3_1">Section 3.1</xref> models the Stackelberg-FlipIt game, defining and explaining each component of the model. <xref ref-type="sec" rid="s3_2">Section 3.2</xref> introduces the MTD decision-making method based on this game.</p>
<sec id="s3_1">
<label>3.1</label>
<title> Stackelberg-FlipIt Game Model</title>
<p>In MTD temporal decision methods, dynamically adjusting the triggering timing of defense strategies is a core challenge. Existing MTD approaches often assume symmetric knowledge of network states between attackers and defenders, neglecting the process by which attackers gradually learn defense thresholds through trial and error. This assumption makes it difficult to precisely control the timing of defense strategies, thereby compromising defense effectiveness and system performance. To address this issue, we introduce a belief factor <italic>B</italic>, capturing the estimation differences between attackers and defenders regarding defense thresholds. The design is based on:
<list list-type="order">
<list-item>
<p><bold>Asymmetric Cognitive Modeling.</bold> Defenders know the actual thresholds but estimate attackers&#x2019; perceptions, while attackers iteratively approximate thresholds through trial and error. The belief factor dynamically reflects these differences, enabling precise defense timing.</p></list-item>
<list-item>
<p><bold>Dynamic Learning and Updating.</bold> The belief factors are updated as the attack-defense process evolves. Defenders adjust estimates by observing attackers, while attackers refine estimates through feedback. This allows defense strategies to adapt to changing conditions, enhancing system adaptability.</p></list-item>
<list-item>
<p><bold>Optimized Defense Timing.</bold> Using the belief factor, defenders can trigger strategies before attackers fully learn thresholds, minimizing attack surface exposure. Trigger frequency can also be adjusted to balance defense effectiveness and system performance.</p></list-item>
</list></p>
<p>By introducing the belief factor, the proposed MTD temporal decision method effectively addresses the following issues:
<list list-type="order">
<list-item>
<p><bold>Precise Control of Defense Timing.</bold> It quantifies cognitive differences, enabling dynamic adjustments to avoid premature or delayed triggers.</p></list-item>
<list-item>
<p><bold>Dynamic Characterization of Attack-Defense Interactions.</bold> Its updates reflect strategic adjustments, capturing the dynamics of network attack-defense.</p></list-item>
<list-item>
<p><bold>Balancing Defense Effectiveness and Performance Overhead.</bold> Defense strategies based on the belief factor ensure effective protection while minimizing system performance overhead, thereby improving the overall efficiency of the MTD system.</p></list-item>
</list></p>
<p>We propose a Stackelberg-FlipIt game under the leader-follower paradigm, modeled as a 9-tuple <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo>&#x003C;</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula> to comprehensively account for real-world network attack-defense parameters and variables.
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> denote the participants in the attack-defense game, where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the defender and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the attacker.</p></list-item>
<list-item>
<p><italic>AS</italic> denotes the common resource contested by both parties in the attack-defense game, i.e., the attack surface, which represents the exploitable surfaces in the system that attackers can discover and utilize. In the game, attacker aim to take control of it and then compromise it, while defenders aim to shift it and protect it from attackers&#x2019; detection. For convenience, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>A</mml:mi><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the <italic>AS</italic> controlled by participant <italic>P</italic> at time <italic>t</italic>.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> are the possible network states at time <italic>t</italic>. Each network state <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the attack frequency, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the system load, and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the <italic>AS</italic> exposure time. The transition of each state is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.
<list list-type="simple">
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>: The network is in a <italic>normal</italic> state, with no detected attack behaviors or potential threats. In this state, the attack frequency is low, the system load remains within normal ranges, and the <italic>AS</italic> is either not exposed or has minimal exposure time.</p>
</list-item>
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>: The network is in a <italic>fragile</italic> state, indicating the presence of potential threats and a possible imminent attack. In this state, the attack frequency gradually increases, the system load exhibits abnormal fluctuations, and the exposure time of the <italic>AS</italic> is prolonged.</p></list-item>
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>: The network is in an under-<italic>damage</italic> state, where some systems or services have already been compromised. In this state, the attack frequency is high, the system load has significantly increased, and the exposure time of the <italic>AS</italic> is prolonged.</p></list-item>
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>: The network is in a <italic>recovery</italic> state, where the attack has been mitigated, and the system is gradually returning to normal operation. In this state, the attack frequency decreases, the system load is progressively restored to normal levels, and the exposure time of the <italic>AS</italic> is reduced.</p></list-item>
</list></p></list-item>
</list></p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>State transitions in the Stackelberg-FlipIt game</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-1.tif"/>
</fig>
<p>The conditions for classifying the four states are as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>S</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>S</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>S</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>S</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the vector of thresholds for <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. These three thresholds are only transparent to the defender, while the attacker can only gradually approximate them through continuous exploration. This setup aligns more closely with real-world network attack-defense dynamics. Whether these thresholds are exceeded depends on the actions of both the attacker and the defender. For example, a state transition <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> hints that at time <italic>t AS</italic> is FLIPPED and controlled by the attacker. This aspect will be further elaborated in the subsequent description of actions and the algorithm.
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which denotes the total time required for the attack-defense game, which is the sum of the total time the attacker controls the <italic>AS</italic> (<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>) and the total time the defender controls the <italic>AS</italic> (<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>). To simplify the analysis, we assume that the attack-defense game unfolds within a finite and discrete time frame, i.e., <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> denote the actions in the attack-defense game. The attacker&#x2019;s action set is represented as <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> which, for simplicity, is categorized into three levels, as shown in <xref ref-type="table" rid="table-2">Table 2</xref>. Similarly, the defender&#x2019;s action set is represented as <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, and categorized into three levels. At any discrete time <italic>t</italic>, both the attacker and the defender may take actions to gain control of the <italic>AS</italic>.</p>
</list-item>
</list></p>
<p><list list-type="bullet">
<list-item>
<p>Belief factor <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the participants&#x2019; estimated values of the thresholds <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow></mml:math></inline-formula> perceived by the attacker, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. It is a crucial factor influencing the temporal decisions of participants in the Stackelberg-FlipIt game. As the game progresses, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> gradually approaches the actual values. Therefore, for the attacker, the primary goal is to quickly converge <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> to the actual values, as mastering the true thresholds enables prolonged control over the <italic>AS</italic>. For the defender, although the thresholds are transparent, waiting until the attacker learns the actual values before taking action would prolong the loss of control over <italic>AS</italic>, potentially leading to defense failure. Thus, the defender also needs to estimate <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, denoted as <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The defender determines the timing of strategy triggers based on <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula>, which represents the difference between the defender&#x2019;s estimated network state values and the attacker&#x2019;s estimated thresholds. This ensures targeted and timely implementation of defense measures.</p>
</list-item>
</list></p>
<p><list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> denote the reward functions, satisfying:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>F</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mi>B</mml:mi><mml:mi>E</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the weights that satisfy <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. <italic>BE</italic> is the belief error, satisfying
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>B</mml:mi><mml:msup><mml:mi>E</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>B</mml:mi><mml:msup><mml:mi>E</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p></list-item>
</list></p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The relationship among action levels, states, and costs</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Action level</th>
<th align="center">For attacker</th>
<th align="center">For defender</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Low</bold></td>
<td>High attack frequency and low costs</td>
<td>Low system load, high <italic>AS</italic> exposure time, and low costs</td>
</tr>
<tr>
<td><bold>Medium</bold></td>
<td>Medium attack frequency and medium costs</td>
<td>Medium system load, medium <italic>AS</italic> exposure time, and medium costs</td>
</tr>
<tr>
<td><bold>High</bold></td>
<td>Low attack frequency and high costs</td>
<td>High system load, low <italic>AS</italic> exposure time, and high costs</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Initial and updated values of belief factor</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th>Belief factor</th>
<th align="center">Initial value</th>
<th align="center">Update</th>
<th>Formula</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>Random values or estimates based on prior knowledge</td>
<td>by trial and observation on defender&#x2019;s reaction</td>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (2)</td>
</tr>
<tr>
<td><inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>Known threshold</td>
<td>by observation on attacker&#x2019;s reaction and network</td>
<td><inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (3)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-3fn1" fn-type="other">
<p>Note: <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are the learning rate of participants&#x2019; belief factor, and defaulted as (0.1, 0.1, 0.1).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><italic>R</italic> mainly rely on three categories of reward evaluation <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>F</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which evaluate the security, performance, and affordability respectively. The design of <italic>BE</italic> is primarily aimed at guiding strategy updates. In our presumption, the weights are configurable, so that our method can be applicable to various MTD systems, e.g., a lightweight system that aims to offer modest security and system performance with the least cost, or, a large defense system that pursues as much as possible a highly secure defense without focusing much on cost. This design is motivated by the following two reasons:
<list list-type="order">
<list-item>
<p>A multi-dimensional evaluation framework provides a more holistic assessment of strategy effectiveness, capturing various aspects of performance, security, and cost.</p></list-item>
<list-item>
<p>Existing MTD decision-making methods often lack clear and well-defined reward structures, which limits their practical applicability.</p></list-item>
</list></p>
<p>Clarification of reward design in our model addresses it by designing a robust and transparent reward function considering former evaluation metrics of MTD technologies over the past decade. These metrics ensure that the reward function is both comprehensive and grounded in established research, enhancing the practicality and effectiveness of the proposed method.</p>
<p><bold>Assumption 1:</bold> The values of each component in the reward function are normalized to the range [0,1]. This design choice is justified because the evaluations underlying the reward function are all derived from probabilities or ratios, which naturally fall within this range.
<list list-type="simple">
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For defender: Availability Probability</p></list-item>
</list></p></list-item>
</list></p>
<p>To quantify the effectiveness of the defense system in terms of security, we use <italic>AP</italic> (Availability Probability) [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>] to measure the probability that the defender controls the <italic>AS</italic> over a period <italic>T</italic>. The formula is defined as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:mi>T</mml:mi></mml:mfrac></mml:math></disp-formula>
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For attacker: Mean Time to Compromise Ratio</p></list-item>
</list></p>
<p><italic>MTTCR</italic> (Mean Time to Compromise Ratio) [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>] represents the ratio of the time the attacker controls the <italic>AS</italic> to the total time, which is essentially the probability of the attacker controlling <italic>AS</italic>. This metric indirectly reflects the attacker&#x2019;s ability to fully access and exploit <italic>AS</italic>.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>M</mml:mi><mml:mi>T</mml:mi><mml:mi>T</mml:mi><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mi>T</mml:mi></mml:mfrac></mml:math></disp-formula>
<list list-type="simple">
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For defender: Escape Probability</p></list-item>
</list></p></list-item>
</list></p>
<p><italic>EP</italic> (Escape Probability) [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>] represents the probability that the defender transfers the <italic>AS</italic> after detecting that the attacker has gained control of it. This metric emphasizes the system&#x2019;s ability to promptly defend itself upon detecting an attack. The formula is expressed as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>E</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the probability of reaching state <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from state <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in which the defender takes action <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and the attacker takes action <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For attacker: Successful Attack Probability</p></list-item>
</list></p>
<p><italic>SAP</italic> (Successful Attack Probability) [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>] represents the probability that the attacker, after gaining control of the <italic>AS</italic>, proceeds with the next attack while avoiding detection and triggering an <italic>AS</italic> transfer by the defender. This metric primarily reflects the attacker&#x2019;s ability to conduct stealthy attacks. The formula is expressed as:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<list list-type="simple">
<list-item>
<label>&#x025AA;</label><p><inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>F</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, the same as <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For defender: Defense Cost</p></list-item>
</list></p></list-item>
</list></p>
<p><italic>DC</italic> (Defense Cost) represents the cost incurred by the defender when executing defense strategies to counter attacks. This metric measures the expenses associated with defensive actions. Due to the multi-stage nature of the attack-defense design, this indicator indirectly incorporates the time costs of both parties into consideration. Since the costs of strategies vary, no specific formula is provided; instead, the costs are comprehensively calculated after each action.
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>For attacker: Attacker Cost</p></list-item>
</list></p>
<p>Similar to <italic>DC</italic>, <italic>AC</italic> (Attack Cost) measures the overhead incurred by the attacker when executing attack strategies. Likewise, it is only necessary to comprehensively calculate the costs after each action.
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the policy function, representing the probability that participant <italic>P</italic> takes action <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> at a given time t and state <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Its update formula is:</p></list-item>
</list>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the policy learning rate for participant <italic>P</italic>. The learning rate dynamically adjusts based on policy performance to adapt to rapidly changing environments [<xref ref-type="bibr" rid="ref-25">25</xref>]. When the current policy outperforms the average policy, the learning rate decreases as <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to reduce the magnitude of policy adjustments. Conversely, when the current policy underperforms, the learning rate increases as <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to quickly learn better strategies, improving the responsiveness and precision of MTD temporal decisions. This fast-slow adaptation balances the trade-off between the effectiveness and cost of temporal strategies. Its selective formula is:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#xA0;if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:msup><mml:mover><mml:mi>&#x03C0;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#xA0;otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the policy gradient:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title> Multi-Agent WoLF-PHC for Stackelberg-FlipIt-MTD Scenarios</title>
<p>Given the setting of the game model, we model the MTD attack-defense process as a multi-stage Stackelberg-FlipIt game, incorporating the MTD scenario into the attack-defense game evolution through the asymmetric projections of both parties. Finally, the proposed multi-agent WoLF-PHC algorithm is applied to solve the MTD temporal decision-making method.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Stackelberg-FlipIt Game for MTD</title>
<p>For the attacker, the knowledge about the system network evolves as the penetration deepens. The attacker&#x2019;s belief factor about the network system continuously changes, and the attacker can only perceive and attack the network system through continuous probing and observation of potential defense strategies. Moreover, due to the dynamic movement of the <italic>AS</italic> in MTD, this local projection may also deviate from objective reality. In contrast, the system administrator (defender) has absolute physical control over the system, along with sufficient security assessment and penetration testing capabilities.</p>
<p>In MTD scenarios, the existence of unknown vulnerabilities does not diminish the defender&#x2019;s absolute cognitive advantage over the attacker. However, since MTD typically involves high defense costs, the set of movable strategies is generally limited, and defensive measures are not taken in every game. Therefore, MTD temporal decision-making methods are crucial. The defender needs to adjust the estimated belief factor of the attacker by observing the attacker&#x2019;s behavior, enabling precise and effective transfer of the attack surface to resist attacks while minimizing costs and avoiding aimless actions.</p>
<p>In summary, in MTD scenarios, the defender can proactively adjust system security configurations, while the attacker can only passively respond to these dynamic changes. Given this, we model the proactive and reactive relationship between the attacker and defender using the Stackelberg-FlipIt game model, defining the defender as the &#x201C;leader&#x201D; and the attacker as the &#x201C;follower&#x201D;. Each game stage is divided into two sub-stages:
<list list-type="order">
<list-item>
<p>Defender&#x2019;s turn: The defender selects one of the MTD movable states or chooses not to act as a defense strategy based on <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, aiming to hinder or delay the attacker&#x2019;s updates of <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and maintain control over <italic>AS</italic> for as long as possible within the time period <italic>T</italic>. In most cases, the defender does not need to trigger defense strategies aimlessly. Relevant experiments [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>] have shown that even when the attacker perceives potential security configuration changes (even if none actually occur), it can increase cognitive load and delay or prevent attack activities.</p></list-item>
<list-item>
<p>Attacker&#x2019;s turn: The attacker initiates attacks based on the initial <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and continuously adjusts it during the process to avoid detection and triggering <italic>AS</italic> transferring by the defender, also aiming to gain control over <italic>AS</italic> for as long as possible within the time period <italic>T</italic>.</p></list-item>
</list></p>
<p>It is evident that the differences in the estimation of the belief factor cause the attacker and defender to operate within entirely different frameworks during the game. This asymmetric interaction is a key feature of the proposed Stackelberg-FlipIt game model, enabling a more realistic and dynamic representation of MTD attack-defense processes.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Multi-Agent WoLF-PHC for MTD Temporal Decision-Making</title>
<p>In the Stackelberg-FlipIt game, the participants&#x2019; actions occur in a specific sequence. Typically, the &#x201C;leader&#x201D; guides the game toward a globally optimal state for themselves. Leveraging cognitive advantages, the defender can conduct multi-stage risk analysis on <italic>AS</italic> to assess the impact of various MTD strategies on risk, which serves as the basis for the &#x201C;leader&#x2019;s&#x201D; decisions. Meanwhile, the attacker, as a rational &#x201C;follower&#x201D;, selects the optimal attack strategy based on their current cognition. The Stackelberg game mechanism is implemented through a subsequential action signal. When the <italic>signal</italic> is 0, the defender observes and selects an action, and the <italic>signal</italic> changes to 1. Upon receiving the <italic>signal</italic> as 1, the attacker observes and selects an action, and the <italic>signal</italic> changes back to 0. This cycle repeats until the game period <italic>T</italic> ends.</p>
<p>The multi-agent WoLF-PHC algorithm continuously optimizes strategies while minimizing environmental influences to achieve faster learning speeds and better performance. Based on its characteristics, we design Algorithm 1 to solve the MTD temporal decision-making method.</p>
<fig id="fig-6">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-6.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment</title>
<p>In this section we aim to validate the effectiveness of our proposed method through experiments on IP address dynamic hopping (short as IP hopping) against scanning attacks. The experimental design focuses on evaluating algorithm performance, optimizing strategy triggering timing, and assessing practical application value, ensuring that the proposed model and algorithm are not only theoretically innovative but also feasible in real-world scenarios.</p>
<p>We use simulation experiments to demonstrate the superior performance of the proposed algorithm. <xref ref-type="sec" rid="s4_1">Section 4.1</xref> details the design of the simulation environment, including the strategies of both attackers and defenders, as well as other critical elements. <xref ref-type="sec" rid="s4_2">Section 4.2</xref> compares and selects hyperparameters for the algorithm. <xref ref-type="sec" rid="s4_3">Section 4.3</xref> compares our algorithm with other classical algorithms and provides experimental analysis.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<p>The IP address is a critical target for attackers launching scanning attacks, DDoS attacks, and other types of subsequent attacks. Therefore, IP hopping defense [<xref ref-type="bibr" rid="ref-28">28</xref>] is one of the key technologies in MTD. Defenders use MTD techniques to randomly change host IP addresses, achieving attack evasion. Existing research has shown that IP hopping is an effective method against scanning attacks. However, blindly triggering IP hopping without perceiving or accurately perceiving the attacker&#x2019;s behavior can lead to two issues:
<list list-type="order">
<list-item>
<p>It may fail to achieve the desired defensive effect and could provide attackers with more information to counteract.</p></list-item>
<list-item>
<p>It may result in unnecessary system resource overhead, leading to inefficiencies.</p></list-item>
</list></p>
<p>The experimental algorithm is run on OpenAI Gym, which provides a multi-agent game simulator for simulating complex network attack-defense scenarios. The experimental environment is deployed in an SDN network controlled by an OpenFlow controller [<xref ref-type="bibr" rid="ref-29">29</xref>]. We use Mininet [<xref ref-type="bibr" rid="ref-30">30</xref>] to create an OpenFlow switch network as the data plane of the SDN, while the control plane is constructed and deployed using the Ryu platform to implement the IP address dynamic defense system, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. In this network, 400 hosts are randomly distributed across 20 subnets, with the subnets connected through OpenFlow switches.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Virtual network topology on Mininet</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-2.tif"/>
</fig>
<p>The experiment simulates real-world scanning attacks and IP hopping defense techniques. The network system state <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> involved information in the attack-defense process that includes attack frequency, system load, and exposure time of real IP addresses. The system load is influenced by both the attacker and the defender, as the defense system may stop services after receiving a large number of data packets, thereby reducing QoS. The individual threshold parameters of the states are normalized as a middle vector with <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to simplify the multi-agent reinforcement learning process.</p>
<p>The attacker employs a coordinated scanning approach, where a certain number of hosts simultaneously initiate scanning attacks by sending probe packets to the IP addresses of the defense network. In the experiment, we assume 50 to 200 scanning hosts targeting 400 defense network hosts. The attacker is assumed to know the IP address resource pool of the defense network and can divide the entire IP address space of the defense network, evenly distributing it among the scanning hosts. These hosts perform random scans on their assigned IP address segments with a uniform distribution. The attacker&#x2019;s goal is to detect the IP addresses currently in use by online hosts to build an attack list, providing a foundation for subsequent attacks. Cautious attackers typically avoid rescanning already scanned IP addresses to minimize the likelihood of detection failure [<xref ref-type="bibr" rid="ref-31">31</xref>]. In the experiment, the scanning attack is designed to be propagative, meaning that detected online hosts are compromised and become new scanning hosts. The unscanned IP address ranges are redistributed among the new set of scanning hosts, with an infection time step of 1.</p>
<p>Each scanning host has two attribute characteristics:
<list list-type="order">
<list-item>
<p>Frequency of Scanning Packet (<italic>FSP</italic>): The average number of probe packets sent per second.</p></list-item>
<list-item>
<p>Proportion of Scanning Packet (<italic>PSP</italic>): The ratio of probe packets to the total packets sent.</p></list-item>
</list></p>
<p>Based on these two characteristics, the attacker&#x2019;s scanning behavior can be classified into three levels, corresponding to the action classification in the Stackelberg-FlipIt game model, as shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Attacker&#x2019;s action level</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Scanning level</th>
<th><italic>FSP</italic></th>
<th><italic>PSP</italic></th>
<th><italic>Cost</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Low</bold></td>
<td>0&#x007E;0.5/s</td>
<td>0%&#x007E;20%</td>
<td>0.1</td>
</tr>
<tr>
<td><bold>Medium</bold></td>
<td>0.5&#x007E;1/s</td>
<td>20%&#x007E;40%</td>
<td>0.5</td>
</tr>
<tr>
<td><bold>High</bold></td>
<td>1&#x007E;2/s</td>
<td>40%&#x007E;80%</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The defense mechanism assigns each host within the system both a fixed real IP address and a time-varying virtual IP address. IP address alteration operates via two distinct modes: periodic time-triggered IP hopping (low-level defense), and event-driven IP hopping (mid-level and high-level defense). The periodic mode introduces fault tolerance for defensive event adjudication, providing a buffer mechanism to mitigate security risks when event determination errors occur. IP hopping behaviors are hierarchically classified into three tiers based on two metrics as systematically detailed in <xref ref-type="table" rid="table-5">Table 5</xref>:</p>
<p><list list-type="order">
<list-item>
<p>Hopping Frequency (<italic>HF</italic>): The frequency of IP address hops.</p></list-item>
<list-item>
<p>Hopping Range (<italic>HR</italic>): The span of address space covered during hopping.</p></list-item>
</list></p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Defender&#x2019;s action level</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>IP hopping Level</th>
<th><italic>HP</italic></th>
<th><italic>HR</italic></th>
<th><italic>Cost</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Low</bold></td>
<td>0&#x007E;0.05/s</td>
<td>0%&#x007E;25%</td>
<td>0.5</td>
</tr>
<tr>
<td><bold>Medium</bold></td>
<td>0.05&#x007E;0.1/s</td>
<td>25%&#x007E;50%</td>
<td>1</td>
</tr>
<tr>
<td><bold>High</bold></td>
<td>0.1&#x007E;0.2/s</td>
<td>50%&#x007E;75%</td>
<td>2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Hyperparameters Configuration</title>
<p>The experiment employs 1000 episodes with <italic>T</italic> &#x003D; 100 steps per episode and 100 independent trials. Hyperparameter configurations are summarized in <xref ref-type="table" rid="table-6">Table 6</xref>, with the following design rationales:</p>
<p><list list-type="bullet">
<list-item>
<p>Learning rate <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is parameterized with two tiers:
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p><inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> must be sufficiently large to ensure rapid convergence when policy performance is favorable but must avoid excessive magnitudes that induce destabilizing oscillations. Usually, <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>.</p></list-item>
<list-item>
<label>&#x02666;</label><p><inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> should remain small to prioritize cautious adjustments during suboptimal policy execution, preventing abrupt deviations from near-optimal strategies, while avoiding undersized values that hinder convergence speed. Usually, <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.01</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>.</p></list-item>
</list></p></list-item>
<list-item>
<label>&#x02022;</label><p>Smoothing parameter <italic>&#x03B2;</italic> (average policy update) functions as a learning rate for updating the average policy. A smaller <italic>&#x03B2;</italic> emphasizes historical policy performance through slower updates, while a larger <italic>&#x03B2;</italic> prioritizes recent policy behaviors with faster adaptation. Empirically, <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.01</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item>
<label>&#x02022;</label><p>Reward weight coefficient <italic>w</italic> directly governs optimization objectives and algorithmic performance. Weight assignments must jointly consider:
<list list-type="simple">
<list-item>
<label>&#x02666;</label><p>Security (high priority, as the core objective of MTD), <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.4</mml:mn><mml:mo>,</mml:mo><mml:mn>0.6</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item>
<label>&#x02666;</label><p>Operational performance (moderate weight to ensure user experience and service continuity), <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.3</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item>
<label>&#x02666;</label><p>Defense overhead (constrained to reasonable levels to prevent resource overconsumption), <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></list-item>
</list></p></list-item>
</list></p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Values configuration of learning rate, smoothing parameter, and reward weight coefficient</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>No.</th>
<th><inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">l</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><bold><italic>&#x03B2;</italic></bold></th>
<th><inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mn mathvariant="bold">2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mn mathvariant="bold">3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>0.5</td>
<td>0.1</td>
<td>0.1</td>
<td>0.6</td>
<td>0.3</td>
<td>0.1</td>
</tr>
<tr>
<td>2</td>
<td>0.2</td>
<td>0.05</td>
<td>0.05</td>
<td>0.5</td>
<td>0.3</td>
<td>0.2</td>
</tr>
<tr>
<td><bold>3</bold></td>
<td><bold>0.2</bold></td>
<td><bold>0.1</bold></td>
<td><bold>0.1</bold></td>
<td><bold>0.6</bold></td>
<td><bold>0.2</bold></td>
<td><bold>0.2</bold></td>
</tr>
<tr>
<td>4</td>
<td>0.1</td>
<td>0.05</td>
<td>0.05</td>
<td>0.5</td>
<td>0.3</td>
<td>0.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Through a series of preliminary experiments, we rigorously validated the guaranteed convergence of the algorithm within 1000 episodes. These experiments further enabled the determination of optimal hyperparameter configurations to maximize algorithmic performance, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Defender&#x2019;s reward after 1000 episodes under different hyperparameters</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-3.tif"/>
</fig>
<p>As demonstrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref> and <xref ref-type="table" rid="table-7">Table 7</xref>, Config. 3 (0.2, 0.1, 0.1, 0.1, 0.6, 0.2, 0.2) achieves the optimal performance, exhibiting both the fastest convergence and the highest final policy reward.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of different configurations</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>No.</th>
<th>Convergence episode</th>
<th>Policy reward</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>453</td>
<td>1.387</td>
</tr>
<tr>
<td>2</td>
<td>587</td>
<td>1.563</td>
</tr>
<tr>
<td><bold>3</bold></td>
<td><bold>421</bold></td>
<td><bold>1.970</bold></td>
</tr>
<tr>
<td>4</td>
<td>504</td>
<td>1.475</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Configs. 1 and 3 exhibit accelerated convergence due to their larger <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, enabling rapid stabilization of near-optimal policies. Conversely, Configs. 2 and 4 converge more slowly and prolong policy refinement. Despite its rapid convergence, Config. 1 yields the lowest reward because its excessive allocation to <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, which leads to the time-occupancy weight prioritizing adversarial engagement duration at the expense of strategy cost optimization, of which the imbalance diminishes the overall reward efficacy. Hence, Config. 3 achieves superior rewards by balancing weights to harmonize security, performance, and affordability objectives. In actual application, the configuration of reward weight coefficients can be manually changed regarding what the goal of the task is, e.g., a lightweight defense system aiming at modest security and performance with the least cost. Here, we mainly focus on the balance among security, performance, and affordability with the weight coefficient being (0.6, 0.2, 0.2).</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experiment Analysis</title>
<p>In the following experiments, by comparing three IP hopping methods&#x2014;OF-RHM [<xref ref-type="bibr" rid="ref-32">32</xref>], SEHT [<xref ref-type="bibr" rid="ref-33">33</xref>], and DDS [<xref ref-type="bibr" rid="ref-34">34</xref>]&#x2014;we comprehensively analyze the results from three perspectives: security, performance, and affordability. As for security analysis, Host Survival Ratio (<italic>HSR</italic>) and Host Survival Average Time (<italic>HSAT</italic>) are utilized to evaluate the network security under different algorithms. As for performance and affordability analysis, Actual IP Hopping Frequency (<italic>AHF</italic>) and Network Channel Resource Occupancy (<italic>NCRO</italic>) are utilized.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Security Analysis</title>
<p>The experiment assumes that the attacker performs a uniformly distributed and non-repetitive random scan on the IP address space used by the defense network. The metric <italic>HSR</italic> represents the proportion of remaining hosts <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> that are not detected and included in the attacker&#x2019;s attack list, relative to the total number of hosts <italic>H</italic> in the defense network [<xref ref-type="bibr" rid="ref-35">35</xref>]. This metric is a crucial indicator of the security performance of the relevant algorithms.</p>
<p>Under the settings of 50, 100, 150, and 200 scanning hosts performing coordinated scans on 400 hosts in the defense network, the comprehensive results are shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. In a network without MTD deployment, the <italic>HSR</italic> approaches 0 after approximately 100&#x2013;300 s. The DDS method performs similarly to OF-RHM, while our method and SEHT significantly reduce the decline in survival rate. The final <italic>HSR</italic> of our method consistently remains above 0.5. It is evident that our algorithm performs slightly better than DDS, as our method can precisely determine the timing of defense strategy triggers, guiding the defense network to perform IP hopping to evade scanning attacks.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title><italic>HSR</italic>s under different scanner numbers. (<bold>a</bold>) <italic>HSR</italic>s under 50 scanners; (<bold>b</bold>) <italic>HSR</italic>s under 100 scanners; (<bold>c</bold>) <italic>HSR</italic>s under 150 scanners; (<bold>d</bold>) <italic>HSR</italic>s under 200 scanners</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-4a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-4b.tif"/>
</fig>
<p>After a successful scanning attack, the attacker may launch follow-up attacks such as worm propagation, DDoS, or other subsequent attacks based on the Hit List, aiming to gain deeper control over the <italic>AS</italic>. These attacks require a certain amount of time to deploy and achieve their intended effects. Therefore, the longer the <italic>HSAT</italic> of hosts in the attack list, the more opportunities it provides for the attacker. This metric is primarily interpreted as the duration between the scanning host receiving a response packet from a malicious probe and the IP hopping of the defense network host. The smaller this metric, the lower the success rate of the attacker&#x2019;s subsequent attacks [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>Under the settings of 50, 100, 150, and 200 scanning hosts performing coordinated scans on 400 hosts in the defense network, the comprehensive results are shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. Compared to other methods, as the number of scanning hosts increases, our method significantly reduces <italic>HSAT</italic>, effectively suppressing the development of subsequent attacks.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title><italic>HSAT</italic>s (s) under different scanner numbers. (<bold>a</bold>) <italic>HSR</italic>s under 50 scanners; (<bold>b</bold>) <italic>HSR</italic>s under 100 scanners; (<bold>c</bold>) <italic>HSR</italic>s under 150 scanners; (<bold>d</bold>) <italic>HSR</italic>s under 200 scanners</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-5a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64849-fig-5b.tif"/>
</fig>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Performance and Affordability Analysis</title>
<p>When hosts in the defense network undergo IP hopping, new forwarding flow tables must be promptly configured to maintain normal network services. This process can cause transient high latency, reducing the network QoS of the defense system and impacting normal network operations. Therefore, the flow table configuration rate and IP hopping frequency are critical factors affecting network performance. Since the former is beyond the scope of this paper, we focus on analyzing the <italic>AHF</italic> and <italic>NCRO</italic> metrics. As for <italic>AHF</italic>, A low <italic>AHF</italic> alone does not indicate the quality of the method; it must be analyzed in conjunction with the <italic>HSR</italic>. Specifically, under the same <italic>HSR</italic> condition, a lower <italic>AHF</italic> results in less impact on network QoS and better network performance. Additionally, the <italic>AHF</italic> reflects defense costs due to less triggering.</p>
<p>As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, our method achieves the lowest <italic>AHF</italic> across different <italic>HSR</italic> levels, indicating the lowest defense cost. This is because our method dynamically adjusts the IP hopping strategy trigger frequency based on the belief factor, ensuring that each hopping is effective.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title><italic>AHF</italic>s (time/s) of different algorithms under different <italic>HSR</italic>s</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>HSR</th>
<th>OF-RHM</th>
<th>SEHT</th>
<th>DDS</th>
<th>Our method</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>0.2</bold></td>
<td>0.032</td>
<td>0.027</td>
<td>0.031</td>
<td>0.030</td>
</tr>
<tr>
<td><bold>0.4</bold></td>
<td>0.066</td>
<td>0.051</td>
<td>0.061</td>
<td>0.042</td>
</tr>
<tr>
<td><bold>0.6</bold></td>
<td>0.123</td>
<td>0.114</td>
<td>0.125</td>
<td>0.098</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Real-time sensitive strategies may require the reservation of dedicated channel resources to avoid congestion, which can limit the flexibility of network resource reuse and increase the marginal cost per unit of traffic. Therefore, <italic>NCRO</italic> can intuitively reflect the performance and affordability of the output strategies. The lower NCRO, the temporal strategy can perform better under the same <italic>HSR</italic> condition.</p>
<p>As shown in <xref ref-type="table" rid="table-9">Table 9</xref>, our method outperforms <italic>NCRO</italic> under different security conditions, close to SEHT. Although these temporal strategies may enhance system robustness, their frequent triggering may lead to a decrease in effective data throughput, essentially converting overhead into bandwidth procurement costs. Therefore, providing modest security with <italic>NCRO</italic> under 150 kb/s [<xref ref-type="bibr" rid="ref-35">35</xref>] is the ideal solution for related methods. Under this circumstance, both SEHT and our method can be applicable.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title><italic>NCRO</italic>s (kb/s) of different algorithms</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>HSR</th>
<th>OF-RHM</th>
<th>SEHT</th>
<th>DDS</th>
<th>Our method</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>0.2</bold></td>
<td>137</td>
<td>101</td>
<td>125</td>
<td>97</td>
</tr>
<tr>
<td><bold>0.4</bold></td>
<td>199</td>
<td>135</td>
<td>161</td>
<td>122</td>
</tr>
<tr>
<td><bold>0.6</bold></td>
<td>273</td>
<td>196</td>
<td>225</td>
<td>191</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, our method outperforms existing approaches in the comparative experiments. We attribute this to: (i) A game model that more closely aligns with real-world network attack-defense processes; (ii) The positive influence of the belief factor on strategy trigger timing; (iii) The decision-making adaptability of the multi-agent WoLF-PHC algorithm; (iv) A comprehensive reward function that guides holistic strategy optimization. These factors collectively contribute to the superior performance of our proposed method.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>We proposed an innovative MTD temporal decision-making method in the field of MTD. The core components of this method include the novel Stackelberg-FlipIt game and the Multi-agent WoLF-PHC algorithm, which can be mainly concluded as follows:
<list list-type="bullet">
<list-item>
<p>Firstly, by introducing the hierarchical decision-making structure of Stackelberg game and combining it with the dynamic process modeling of FlipIt game, we constructed the Stackelberg-FlipIt game model. The concept of belief factor was incorporated to influence the timing and planning of decisions for both attackers and defenders.</p></list-item>
<list-item>
<p>Secondly, the multi-agent framework was integrated into the WoLF-PHC algorithm, enabling both attackers and defenders to autonomously learn and adapt their strategies in dynamic environments, which allowed them to respond to evolving network states and adversarial behaviors in real time.</p></list-item>
<list-item>
<p>Thirdly, for the first time, four key metrics&#x2014;security, system performance, affordability, and belief error&#x2014;were integrated into the reward function of multi-agent reinforcement learning, which provided a new perspective for the multi-dimensional quantification of network defense strategies, helping our method output strategies with the strongest overall capabilities.</p></list-item>
<list-item>
<p>Lastly, the effectiveness of the proposed model and algorithm was validated through experiments on IP hopping against scanning attacks. The results demonstrated the algorithm&#x2019;s performance, showing that, with appropriate parameters, the proposed model and algorithm significantly enhanced the adaptability and efficiency of MTD temporal decision-making methods.</p></list-item>
</list></p>
<p>In conclusion, the proposed method offers a robust and practical solution for MTD temporal decisions, balancing security, performance, and affordability while adapting to dynamic network environments. Although the model and algorithm proposed in this paper have achieved good results in simulation experiments, there are still some limitations and directions for future work:
<list list-type="bullet">
<list-item>
<p>A more adversarial game model: The premise given is that both the attacker and the defense are rational, so the attacker in this paper will not consume defense resources through deception, that is, the &#x201C;intelligence level&#x201D; is not high enough. Future research can lead to the construction of more sophisticated attacker models to simulate more realistic cyberattack behaviors.</p></list-item>
<list-item>
<p>Dynamically adjust hyperparameters: The selection of hyperparameters is mainly based on experiments and experience, and in the future, hyperparameters can be dynamically adjusted through automated methods to adapt to different network environments and attack and defense scenarios.</p></list-item>
<list-item>
<p>Actual deployment and testing: Our model and algorithm have not yet been deployed and tested in a real-world network environment, and these methods can be applied to real-world network systems in the future to evaluate their effectiveness and robustness in the real world.</p></list-item>
</list></p>
</sec>
</body>
<back>
<ack>
<p>Due to the nature of this research, participants of this study agree for their data to be shared publicly as demand.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work is funded by National Natural Science Foundation of China No. 62302520.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Study conception and design: Rongbo Sun, Yuefei Zhu, Jinlong Fei; data collection: Rongbo Sun, Zhongyu Guo; analysis and interpretation of results: Rongbo Sun; draft manuscript preparation: Rongbo Sun, Jinlong Fei, Yuefei Zhu. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data and materials used to support the findings of this study are available from the corresponding author upon request after acceptance.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cho</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Alavizadeh</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ben-Asher</surname> <given-names>N</given-names></string-name>, <string-name><surname>Moore</surname> <given-names>TJ</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Toward proactive, adaptive defense: a survey on moving target defense</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2020</year>;<volume>22</volume>(<issue>1</issue>):<fpage>709</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2019.2963791</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhuang</surname> <given-names>R</given-names></string-name>, <string-name><surname>DeLoach</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Ou</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Towards a theory of moving target defense</article-title>. In: <conf-name>Proceedings of the First ACM Workshop on Moving Target Defense</conf-name>; <year>2014 Nov 7</year>; <publisher-loc>Scottsdale, AZ, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Clark</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>K</given-names></string-name>, <string-name><surname>Poovendran</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Effectiveness of IP address randomization in decoy-based moving target defense</article-title>. In: <conf-name>52nd IEEE Conference on Decision and Control</conf-name>; <year>2013 Dec 10&#x2013;13</year>; <publisher-loc>Florence, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>C</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Optimal temporospatial strategy selection approach to moving target defense: a FlipIt differential game model</article-title>. <source>Comput Secur</source>. <year>2021</year>;<volume>108</volume>(<issue>3</issue>):<fpage>102342</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2021.102342</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Optimal timing of moving target defense: a Stackelberg game model</article-title>. <comment>arXiv:1905.13293. 2019</comment>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Moving target defense decision-making method</article-title>. In: <conf-name>Proceedings of the 7th ACM Workshop on Moving Target Defense</conf-name>; <year>2020 Nov 9</year>; <publisher-name>Virtual</publisher-name>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cook</surname> <given-names>PR</given-names></string-name></person-group>. <article-title>Limitations and extensions of the WoLF-PHC algorithm [master&#x2019;s thesis]. Provo, UT, USA: Brigham Young University</article-title>; <year>2007</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Connell</surname> <given-names>W</given-names></string-name>, <string-name><surname>Albanese</surname> <given-names>M</given-names></string-name>, <string-name><surname>Venkatesan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A framework for moving target defense quantification</article-title>. <source>IFIP Adv Inf Commun Technol</source>. <year>2017</year>;<volume>502</volume>:<fpage>124</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-58469-0_9</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A system attack surface based MTD effectiveness and cost quantification framework</article-title>. In: <conf-name>Proceedings of the 2nd International Conference on Cryptography, Security and Privacy</conf-name>; <year>2018 Mar 16&#x2013;18</year>; <publisher-loc>Guiyang, China</publisher-loc>. p. <fpage>175</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zaffarano</surname> <given-names>K</given-names></string-name>, <string-name><surname>Taylor</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hamilton</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A quantitative framework for moving target defense effectiveness evaluation</article-title>. In: <conf-name>MTD 2015&#x2014;Proceedings of the 2nd ACM Workshop on Moving Target Defense, Co-Located with: CCS</conf-name>; <year>2015 Oct 12</year>; <publisher-loc>Denver, CO, USA</publisher-loc>. p. <fpage>3</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Barron</surname> <given-names>EN</given-names></string-name></person-group>. <source>Game theory: an introduction</source>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Reinforcement learning algorithms for adaptive cyber defense against heartbleed</article-title>. In: <conf-name>Proceedings of the ACM Conference on Computer and Communications Security</conf-name>; <year>2014 Nov 3&#x2013;7</year>; <publisher-loc>Scottsdale, AZ, USA</publisher-loc>. p. <fpage>51</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wright</surname> <given-names>M</given-names></string-name>, <string-name><surname>Venkatesan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Albanese</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wellman</surname> <given-names>MP</given-names></string-name></person-group>. <article-title>Moving target defense against DDoS attacks: an empirical game-theoretic analysis</article-title>. In: <conf-name>MTD 2016&#x2014;Proceedings of the 2016 ACM Workshop on Moving Target Defense, Co-Located with CCS</conf-name>; <year>2016 Oct 24&#x2013;28</year>; <publisher-loc>Vienna, Austria</publisher-loc>. p. <fpage>93</fpage>&#x2013;<lpage>104</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>A survey: when moving target defense meets game theory</chapter-title>. In: <source>Computer science review</source>. Vol. <volume>48</volume>. <publisher-loc>Amsterdam, The Netherlands</publisher-loc>: <publisher-name>Elsevier Ireland Ltd</publisher-name>. <year>2023</year>. doi:<pub-id pub-id-type="doi">10.1016/j.cosrev.2023.100544</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ge</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Toward effectiveness and agility of network security situational awareness using moving target defense (MTD)</article-title>. <source>Sens Syst Space Appl VII</source>. <year>2014</year>;<volume>9085</volume>:<fpage>185</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Carter</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Riordan</surname> <given-names>JF</given-names></string-name>, <string-name><surname>Okhravi</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A game theoretic approach to strategy determination for dynamic platform defenses</article-title>. In: <conf-name>Proceedings of the First ACM Workshop on Moving Target Defense</conf-name>; <year>2014 Nov 7</year>; <publisher-loc>Scottsdale, AZ, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>30</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Manadhata</surname> <given-names>PK</given-names></string-name></person-group>. <chapter-title>Game theoretic approaches to attack surface shifting</chapter-title>. In: <source>Moving target defense II</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2012</year>. p. <fpage>1</fpage>&#x2013;<lpage>13</lpage> doi:<pub-id pub-id-type="doi">10.1007/978-1-4614-5416-8_1</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>van Dijk</surname> <given-names>M</given-names></string-name>, <string-name><surname>Juels</surname> <given-names>A</given-names></string-name>, <string-name><surname>Oprea</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rivest</surname> <given-names>RL</given-names></string-name></person-group>. <article-title>FlipIt: the game of &#x201C;
stealthy takeover&#x201D;</article-title>. <source>J Cryptol</source>. <year>2013</year>;<volume>26</volume>(<issue>4</issue>):<fpage>655</fpage>&#x2013;<lpage>713</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00145-012-9134-5</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Reinforcement learning based self-adaptive moving target defense against DDoS attacks</article-title>. <source>J Phys Conf Ser</source>. <year>2021</year>;<volume>1812</volume>(<issue>1</issue>):<fpage>012039</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1742-6596/1812/1/012039</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cremer</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sheehan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Fortmann</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kia</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Mullins</surname> <given-names>M</given-names></string-name>, <string-name><surname>Murphy</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Cyber risk and cybersecurity: a systematic review of data availability</article-title>. <source>Geneva Pap Risk Insur-Issues Pract</source>. <year>2022</year>;<volume>47</volume>(<issue>3</issue>):<fpage>698</fpage>&#x2013;<lpage>736</lpage>. doi:<pub-id pub-id-type="doi">10.1057/s41288-022-00266-6</pub-id>; <pub-id pub-id-type="pmid">35194352</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Prakash</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wellman</surname> <given-names>MP</given-names></string-name></person-group>. <article-title>Empirical game-theoretic analysis for moving target defense</article-title>. In: <conf-name>Proceedings of the Second ACM Workshop on Moving Target Defense&#x2014;MTD &#x0027;15</conf-name>; <year>2015 Oct 12</year>; <publisher-loc>Denver, CO, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>J-H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Anwar</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Kamhoua</surname> <given-names>CA</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>MP</given-names></string-name></person-group>. <article-title>Foureye: defensive deception against advanced persistent threats via hypergame theory</article-title>. <source>IEEE Trans Netw Serv Manag</source>. <year>2021</year>;<volume>19</volume>(<issue>1</issue>):<fpage>112</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnsm.2021.3117698</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ping</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Estimating probability of success of escape, evacuation, and rescue (EER) on the offshore platform by integrating Bayesian network and fuzzy AHP</article-title>. <source>J Loss Prev Process Ind</source>. <year>2018</year>;<volume>54</volume>(<issue>4</issue>):<fpage>57</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jlp.2018.02.007</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Halgamuge</surname> <given-names>MN</given-names></string-name></person-group>. <article-title>Estimation of the success probability of a malicious attacker on blockchain-based edge network</article-title>. <source>Comput Netw</source>. <year>2022</year>;<volume>219</volume>(<issue>2</issue>):<fpage>109402</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comnet.2022.109402</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mounjid</surname> <given-names>O</given-names></string-name>, <string-name><surname>Lehalle</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Improving reinforcement learning algorithms: towards optimal learning rate policies</article-title>. <source>Math Financ</source>. <year>2024</year>;<volume>34</volume>(<issue>2</issue>):<fpage>588</fpage>&#x2013;<lpage>621</lpage>. doi:<pub-id pub-id-type="doi">10.1111/mafi.12378</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ferguson-Walter</surname> <given-names>KJ</given-names></string-name>, <string-name><surname>LaFon</surname> <given-names>DS</given-names></string-name>, <string-name><surname>Shade</surname> <given-names>TB</given-names></string-name></person-group>. <article-title>Friend or faux: deception for cyber defense</article-title>. <source>J Inf Warf</source>. <year>2017</year>;<volume>16</volume>(<issue>2</issue>):<fpage>28</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Johnson</surname> <given-names>CK</given-names></string-name></person-group>. <source>Decision-making biases in cybersecurity: measuring the impact of the sunk cost fallacy to delay attacker behavior</source>. <publisher-loc>Tempe, AZ, USA</publisher-loc>: <publisher-name>Arizona State University</publisher-name>; <year>2022</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>S-Y</given-names></string-name>, <string-name><surname>Park</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Babu</surname> <given-names>BA</given-names></string-name></person-group>. <article-title>Fast IP hopping randomization to secure hop-by-hop access in SDN</article-title>. <source>IEEE Trans Netw Serv Manag</source>. <year>2019</year>;<volume>16</volume>(<issue>1</issue>):<fpage>308</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnsm.2018.2889842</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>McKeown</surname> <given-names>N</given-names></string-name>, <string-name><surname>Anderson</surname> <given-names>T</given-names></string-name>, <string-name><surname>Balakrishnan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Parulkar</surname> <given-names>G</given-names></string-name>, <string-name><surname>Peterson</surname> <given-names>L</given-names></string-name>, <string-name><surname>Rexford</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>OpenFlow: enabling innovation in campus networks</article-title>. <source>ACM Spec Interest Group Data Commun</source>. <year>2008</year>;<volume>38</volume>(<issue>2</issue>):<fpage>69</fpage>&#x2013;<lpage>74</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rodr&#x00ED;guez</surname> <given-names>P</given-names></string-name>, <string-name><surname>Andrea</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Using mininet for emulation and prototyping software-defined networks</article-title>. In: <conf-name>2014 IEEE Colombian Conference on Communications and Computing (COLCOM)</conf-name>; <year>2014 Jun 4&#x2013;6</year>; <publisher-loc>Bogot&#x00E1;, DC, Colombia</publisher-loc>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aibekova</surname> <given-names>A</given-names></string-name>, <string-name><surname>Selvarajah</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Offensive security: study on penetration testing attacks, methods, and their types</article-title>. In: <conf-name>2022 IEEE International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE)</conf-name>; <year>2022 Apr 23&#x2013;24</year>; <publisher-loc>Ballari, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jafar</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ehab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>D</given-names></string-name></person-group>. <article-title>OpenFlow random host mutation: transparent moving target defense using software defined networking</article-title>. In: <conf-name>Proceedings of the First Workshop on Hot Topics in Software Defined Networks</conf-name>; <year>2013 Aug 16</year>; <publisher-loc>Hong Kong, China. New York, NY, USA</publisher-loc>: <publisher-name>ACM Digital Library</publisher-name>; <year>2013</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>HQ</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>YJ</given-names></string-name></person-group>. <article-title>Network moving target defense technique based on self-adaptive end-point hopping</article-title>. <source>Arab J Sci Eng</source>. <year>2017</year>;<volume>42</volume>(<issue>8</issue>):<fpage>3249</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s13369-017-2430-5</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Miao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>G</given-names></string-name></person-group>. <article-title>The design and implementation of a dynamic IP defense system accelerated by vector packet processing</article-title>. In: <conf-name>Proceedings of the International Conference on Industrial Control Network and System Engineering Research (ICNSER2019)</conf-name>; <year>2019 Mar 15&#x2013;16</year>; <publisher-loc>Shenyang, China. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>64</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>An adaptive IP hopping approach for moving target defense using a light-weight CNN detector</article-title>. <source>Secur Commun Netw</source>. <year>2021</year>;<volume>2021</volume>(<issue>1</issue>):<fpage>8848473</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/8848473</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>











