<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79895</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.079895</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An HRMCTS-Based Optimization Method for Efficient Multi-Objective Path Planning</article-title>
<alt-title alt-title-type="left-running-head">An HRMCTS-Based Optimization Method for Efficient Multi-Objective Path Planning</alt-title>
<alt-title alt-title-type="right-running-head">An HRMCTS-Based Optimization Method for Efficient Multi-Objective Path Planning</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Qianshu</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-1422-1096</contrib-id>
<name name-style="western"><surname>Liu</surname><given-names>Shuangxi</given-names></name><email>lsxdouble@163.com</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wu</surname><given-names>Xianyu</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Zhao</surname><given-names>Wei</given-names></name></contrib>
<aff id="aff-1"><institution>Advanced Propulsion Technology Laboratory, National University of Defense Technology</institution>, <addr-line>Changsha</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Shuangxi Liu. Email: <email>lsxdouble@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>8</day><month>5</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>1</issue>
<elocation-id>93</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_79895.pdf"></self-uri>
<abstract>
<p>Path planning for unmanned systems in complex environments must simultaneously satisfy safety, kinematic feasibility, and real-time performance requirements. Monte Carlo Tree Search (MCTS) offers advantages such as model-free operation, strong interpretability, and anytime planning capability, but it suffers from large branching factors, excessive search depths, and poor convergence under sparse reward conditions in high-dimensional state spaces. To address these challenges, this paper proposes a Heuristic Rolling Monte Carlo Tree Search (HRMCTS) framework. First, the path planning problem is formulated as a constrained Markov decision process, where the state consists of position and heading, and actions are discretized heading changes. Second, a heuristic selection strategy incorporating goal-directed guidance and obstacle safety margins is introduced to improve search directionality, while a limited-depth forward simulation with branch pruning is employed during the rollout phase. A multi-objective reward function is designed to integrate distance, goal progress, tangent-based obstacle avoidance, smoothness, and efficiency, thereby jointly optimizing path quality and computational performance. Experiments are conducted in three scenarios: static polygonal environments, dynamic circular obstacle environments, and dynamic polygonal obstacle environments. Simulation results demonstrate that the proposed method offers significant advantages in terms of planning efficiency, environmental adaptability, generalization capability, and interpretability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Unmanned system</kwd>
<kwd>path planning</kwd>
<kwd>MCTS</kwd>
<kwd>heuristic algorithm</kwd>
<kwd>reward shaping</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Hunan Provincial Natural Science Foundation of China</funding-source>
<award-id>2025JJ60072</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Open Research Subject of State Key Laboratory of Intelligent Game</funding-source>
<award-id>ZBKF-24-01</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Postdoctoral Fellowship Program of CPSF</funding-source>
<award-id>GZB20240989</award-id>
</award-group>
<award-group id="awg4">
<funding-source>China Postdoctoral Science Foundation</funding-source>
<award-id>2024M754304</award-id>
</award-group>
<award-group id="awg5">
<funding-source>Military High-Level Talent Program</funding-source>
<award-id>202401-RCGC-ZZ-004</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Unmanned systems, such as unmanned aerial vehicles (UAVs), unmanned ground vehicles (UGVs), autonomous underwater vehicles (AUVs), and unmanned surface vehicles (USVs), have found extensive applications across various domains. In recent years, the path planning issues concerning these unmanned systems have garnered significant research attention [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>]. In this work, UAV is selected as the case study to examine path planning, which has emerged as a central and challenging research focus in the field of UAV autonomy.</p>
<p>The objective of path planning is to generate an optimal trajectory from starting point to goal point, ensuring both safety and dynamic feasibility. Recent advancements underscore a significant transition from traditional shortest path optimization to a multidimensional equilibrium that encompasses energy efficiency, risk mitigation, and temporal constraints. This shift is particularly vital in high-dynamic and stringently constrained environments, such as autonomous flight through dense obstacle fields or penetration missions in uncertain no-fly zones, where numerous complex and irregularly distributed obstacles render real-time planning particularly challenging. In such scenarios, a high-quality path planner must not only generate collision-free geometric paths but also rigorously respect physical constraints, such as nonholonomic kinematics, maximum angular velocity, and speed limits, while processing vast state spaces within decision windows. This necessity has driven the development of advanced path planning algorithms that tightly integrate environmental perception, kinematic constraints, and online decision-making capabilities, establishing it as a central research focus in the UAV domain [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Existing path planning methodologies are generally categorized into grid-based, sampling-based, and optimization-based paradigms. Within these groups, heuristic optimization algorithms like particle swarm optimization (PSO) have gained significant traction due to their proficiency in global searching and effective constraint handling in complex, dynamic environments. Such capabilities provide a clear advantage over deterministic or exhaustive search techniques [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Deep reinforcement learning (DRL) has also proven highly effective by utilizing deep neural networks to optimize policies within high-dimensional, continuous state spaces [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. However, the DRL decision-making process is often contained within learned network parameters. This reliance on automatic feature extraction lacks transparent reasoning steps, which makes it difficult to map state observations to final actions and potentially compromises the reliability of agent behavior [<xref ref-type="bibr" rid="ref-13">13</xref>]. Consequently, research into RL explainability seeks to clarify how state information influences decisions to foster more transparent and trustworthy policies. Conversely, Monte Carlo Tree Search (MCTS) provides an inherently interpretable framework for decision-making. In this search tree, nodes represent concrete states while branches correspond to feasible actions. Statistics such as visit counts and simulation results directly quantify the value of each action [<xref ref-type="bibr" rid="ref-14">14</xref>]. This explicit structure facilitates human-understandable reasoning because inspecting high-value paths allows one to reconstruct the rationale for a chosen action and clearly state the intent of every planning step.</p>
<p>Common approaches to path planning in dynamic environments include MCTS, A&#x002A; and its hybrid variants, sampling methods, and DRL. MCTS enhances computational efficiency and reduces collision rates through its capacity for real-time decision-making. These benefits are particularly evident when the algorithm incorporates velocity obstacle techniques or adaptive action sets within dynamic environments [<xref ref-type="bibr" rid="ref-15">15</xref>]. Recent research has also emphasized modeling obstacle persistence to facilitate robust multi-agent replanning [<xref ref-type="bibr" rid="ref-16">16</xref>]. Furthermore, extending MCTS for long-term horizons assists in overcoming the complexities associated with high-dimensional action spaces [<xref ref-type="bibr" rid="ref-17">17</xref>]. In structured settings, A&#x002A; and its hybrids are favored for their low latency and ability to generate smooth trajectories. These methods utilize pre-training to prune unnecessary nodes, which makes them highly effective for rapid response in constrained environments [<xref ref-type="bibr" rid="ref-18">18</xref>]. Sampling methods maintain probabilistic completeness and optimality while improving computational speed. These algorithms often employ temporal reasoning and pruning strategies to achieve further optimization [<xref ref-type="bibr" rid="ref-19">19</xref>]. While DRL displays significant adaptability after training, its inherent sample inefficiency often limits real-time implementation. Techniques such as prioritized experience replay and attention mechanisms are therefore required to accelerate policy convergence [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. Ultimately, MCTS serves as a compelling foundation for efficient dynamic path planning due to its balance of adaptability, speed, and interpretability. It avoids the black-box nature of DRL, the potential lack of optimization guarantees in sampling methods, and the structural constraints inherent to A&#x002A;.</p>
<p>Building upon the advancements in single-agent navigation, the focus of recent scholarship has increasingly pivoted toward the complexities of multi-agent collaborative planning. A primary concern in this domain is the deployment of heterogeneous Amphibious Unmanned Surface Vehicles (AUSV) under stringent energy budgets. To address this, He et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] utilized swarm-based coordination to optimize coverage efficiency, while Li et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] introduced dynamic path-adjustment models grounded in real-time energy consumption telemetry.</p>
<p>The challenge of heterogeneity extends to USV-UAV synergistic operations, where Huang et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] demonstrated that incorporating hull dynamics into the planning phase significantly enhances target search efficacy. Such robustness is equally critical in chaotic underwater environments. For instance, in scenarios characterized by 3D terrain and volatile currents, multi-objective frameworks are now frequently employed to navigate the inherent trade-offs between path safety and mission timeliness [<xref ref-type="bibr" rid="ref-24">24</xref>]. Recent trends, as exemplified by Sun and Lv [<xref ref-type="bibr" rid="ref-25">25</xref>], favor a hybridized approach: integrating PSO for global strategic mapping with the Dynamic Window Algorithm (DWA) for tactical obstacle avoidance. This dual-layer architecture, often complemented by evolutionary algorithms to identify Pareto-optimal solutions, ensures that conflicting metrics, such as communication quality and energy overhead, are systematically balanced. Furthermore, in communication-sparse zones, distributed frameworks based on consensus or leader-follower theory continue to serve as the benchmark for maintaining formation stability [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p>MCTS frames path planning as a sequential decision-making problem within the context of a Markov Decision Process (MDP). Although traditional greedy heuristics often stall in local optima during long-horizon missions, MCTS maintains a robust balance between exploration and exploitation. It achieves this equilibrium by incrementally building a search tree through selective simulations. This mechanism offers a clear advantage in interpretability. Rather than functioning as an opaque black box, MCTS populates its nodes with explicit state statistics. Data such as visitation frequency and cumulative rewards provide a transparent record of the perceived utility of each action. From an operational perspective, the anytime property is particularly valuable in dynamic environments. This feature allows the immediate retrieval of a valid solution once the computational budget is exhausted. However, applying these methods to high-dimensional continuous spaces reveals significant bottlenecks. The combination of exponential branching and sparse reward signals can severely effect convergence rates. To mitigate these structural constraints, researchers have proposed various enhanced MCTS variants to improve sampling efficiency and search depth [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<p>Efforts to augment MCTS generally converge on three critical fronts: narrowing the action space, optimizing search depth, and leveraging structural or hardware acceleration. When confronted with the &#x201C;curse of dimensionality&#x201D; and excessive branching factors, the literature has moved toward more selective growth strategies. These include progressive expansion [<xref ref-type="bibr" rid="ref-29">29</xref>], dominance-based pruning [<xref ref-type="bibr" rid="ref-30">30</xref>&#x2013;<xref ref-type="bibr" rid="ref-32">32</xref>], and the use of macro-actions or &#x201C;options&#x201D; to abstract long-range decision sequences [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. Interestingly, even Large Language Models have recently been tapped to evolve the heuristic functions that steer these searches [<xref ref-type="bibr" rid="ref-35">35</xref>]. To refine simulation efficiency and depth, contemporary research has pivoted toward more sophisticated selection policies&#x2014;utilizing weighted UCB algorithm [<xref ref-type="bibr" rid="ref-36">36</xref>] or value gradients [<xref ref-type="bibr" rid="ref-37">37</xref>], while employing cost-approximations for early simulation termination [<xref ref-type="bibr" rid="ref-38">38</xref>&#x2013;<xref ref-type="bibr" rid="ref-41">41</xref>]. This focus on resource allocation is further evident in the adoption of nested structures [<xref ref-type="bibr" rid="ref-42">42</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>] and neural-network-based value proxies [<xref ref-type="bibr" rid="ref-45">45</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>], both of which concentrate computational effort on the most promising sub-trees. Beyond algorithmic refinements, the integration of transposition tables [<xref ref-type="bibr" rid="ref-47">47</xref>] and information sets [<xref ref-type="bibr" rid="ref-48">48</xref>] has streamlined the underlying search structure. When these are hybridized with deep learning [<xref ref-type="bibr" rid="ref-45">45</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>] or evolutionary meta-heuristics [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>], and further backed by hardware-level parallelization [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-56">56</xref>], MCTS becomes significantly more viable for the intricate demands of complex, real-world environments.</p>
<p>Modern path planning for unmanned systems has undergone a fundamental shift, moving beyond simple shortest-path heuristics toward a complex multi-objective equilibrium that balances energy conservation, risk mitigation, and strict temporal deadlines. This transition is especially critical in highly dynamic or adversarial environments, such as autonomous navigation through dense obstacle fields or penetration missions within contested no-fly zones. In these settings, the primary challenge for any planning algorithm is the need to reconcile high-fidelity kinematic constraints with the necessity of navigating vast state spaces within millisecond-level decision windows.</p>
<p>While MCTS offers a theoretically sound foundation for such problems, standard implementations and common enhancements (e.g., action pruning or parallelization) often fail when tasked with maintaining solution quality under extreme computational and kinematic pressure. There remains a pressing need for MCTS architectures that provide both the scalability and search intensity required for the next generation of intelligent systems. To this end, this work introduces a Heuristic Rolling Monte Carlo Tree Search (HRMCTS) framework. By synthesizing a receding-horizon logic with targeted heuristic guidance, HRMCTS significantly accelerates search convergence while preserving the algorithm&#x2019;s inherent transparency. Although the present study centers on the high-fidelity navigation of a single agent, the core pillars of HRMCTS specifically its localized planning strategy, multi-objective reward design, and heuristic-informed evaluation, like establish a robust computational template for future extensions into heterogeneous multi-agent coordination and energy-constrained missions. The primary contributions of this research are structured as follows:<list list-type="simple">
<list-item>
<label>(1)</label>
<p>We introduce a heuristic upper confidence bound for trees (UCT) selection and rollout strategy that embeds goal-directed guidance and obstacle safety margins directly into the search process. By mitigating the raw stochastic exploration and high branching factors inherent in traditional MCTS, this approach significantly sharpens search directionality and minimizes redundant computational overhead.</p></list-item>
<list-item>
<label>(2)</label>
<p>To overcome the challenges of excessive search depth and dynamic threats, this work develops a receding-horizon planning mechanism characterized by limited-depth forward simulation and strategic branch pruning. The resulting closed-loop &#x201C;plan&#x2013;execute&#x2013;replan&#x201D; architecture effectively reconciles the tension between immediate real-time responsiveness and global path feasibility within complex, dynamic environments.</p></list-item>
<list-item>
<label>(3)</label>
<p>A multifaceted reward function is formulated to unify disparate metrics including goal proximity, progress, and trajectory smoothness with predictive dynamic obstacle avoidance. This integrated design facilitates the simultaneous optimization of path quality and efficiency, notably accelerating MCTS convergence in scenarios typically degenerated by sparse reward signals.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Background</title>
<sec id="s2_1">
<label>2.1</label>
<title>MDP</title>
<p>MDP is a probabilistic modeling framework for sequential decision-making, formally characterizing how an agent selects actions in a stochastic environment to maximize cumulative long-term reward. Rather than being a specific problem-solving algorithm, MDP constitutes a general theoretical framework that bridges dynamic system modeling with globally optimal decision-making. It abstracts complex decision problems into a cyclic structure of &#x201C;<italic>state&#x2013;action&#x2013;transition&#x2013;reward</italic>&#x201D;, quantifying environmental uncertainty through a probabilistic transition function while explicitly defining a long-term optimization objective. This formulation provides a unified mathematical representation and optimization logic for sequential decision problems such as path planning [<xref ref-type="bibr" rid="ref-13">13</xref>]. The core assumption of an MDP is the Markov property (memorylessness): the evolution of future states depends solely on the current state and action, independent of the historical trajectory. In path planning, this implies that the agent need not retain its full motion history; instead, only essential information&#x2014;such as current position, environmental constraints (e.g., obstacle distribution), and task progress&#x2014;is sufficient for effective future path decisions. This abstraction reduces state-space complexity while preserving decision efficacy, forming the foundation for MDP&#x2019;s applicability across diverse dynamic scenarios [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p>An MDP is formally defined by a tuple <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, often extended to a quintuple with a discount factor <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>. In the context of path planning, these components are interpreted as follows:<list list-type="simple">
<list-item>
<label>1.</label>
<p>State space <italic>S</italic>: The set of all possible system configurations, commonly represented as coordinate pairs <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in 2D planning scenarios.</p></list-item>
<list-item>
<label>2.</label>
<p>Action space <italic>A</italic>: The set of admissible actions <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>A</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x2286;</mml:mo><mml:mi>A</mml:mi></mml:math></inline-formula> available at state <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>s</mml:mi></mml:math></inline-formula>. While low-branching cases may utilize atomic moves, higher-dimensional problems often rely on action sampling or macro-actions to manage computational complexity.</p></list-item>
<list-item>
<label>3.</label>
<p>Transition function <italic>P</italic>: The probabilistic mapping <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>P</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>:</mml:mo><mml:mi>S</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> that defines environmental dynamics. Although <italic>P</italic> facilitates modeling stochasticity, many path planning implementations treat these transitions as deterministic.</p></list-item>
<list-item>
<label>4.</label>
<p>Reward function <italic>R</italic>: The mapping <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>R</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>:</mml:mo><mml:mi>S</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula> that formalizes the optimization objective. In this paper, <italic>R</italic> must reconcile mission goals with state and path constraints to dictate trajectory quality as a requirement that aligns with RL paradigms where reward density and scale are paramount for convergence [<xref ref-type="bibr" rid="ref-4">4</xref>].</p></list-item>
</list></p>
<p>As a foundational framework for sequential decision-making, MDP provides a abstract modeling paradigm, serving as both the theoretical basis and problem description language for RL and MCTS.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>MCTS</title>
<p>MCTS optimizes decision-making by merging stochastic simulation with structured tree exploration. The algorithm dynamically constructs a search tree to reconcile the tension between exploring unvisited states and exploiting known high-reward trajectories. While MCTS shares a formal problem structure with RL and their conceptual intersections are well-documented [<xref ref-type="bibr" rid="ref-14">14</xref>], this study characterizes MCTS as a heuristic search algorithm to ensure clarity. By employing stochastic sampling in lieu of exhaustive enumeration, MCTS effectively addresses the complexities of state spaces characterized by high branching factors and significant depths. As depicted in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the algorithm approximates the value of specific states through a cyclic four-phase process consisting of selection, expansion, evaluation, and backpropagation. During the selection phase, the agent traverses the tree from the root to a leaf node <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>s</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> based on accumulated statistics. The expansion phase then appends new child nodes to <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>s</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> to represent potential actions. In the evaluation phase, the utility of these new states is estimated through simulation or statistical aggregation. Finally, backpropagation updates the statistics of all ancestral nodes along the path to the root. This iterative refinement allows MCTS to identify optimal policies in environments where conventional value or policy iteration remains computationally prohibitive [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Description of MCTS process.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-1.tif"/>
</fig>
<p>As discussed above, the core of MCTS lies in the selection and rollout phases. By refining the algorithms in these two phases, MCTS variants tailored to specific problems can be developed. The goal of the selection phase is to traverse from the root node down to an expandable leaf node along the branch with the highest exploration value. Its essence is to avoid blind exploration and prioritize actions that yield high information gain. A commonly used selection policy is the UCT formula [<xref ref-type="bibr" rid="ref-57">57</xref>]:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msup><mml:mi>a</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>A</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:mrow><mml:mo>{</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msqrt><mml:mfrac><mml:mrow><mml:mi>ln</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:msqrt><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the reward for executing action <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>a</mml:mi></mml:math></inline-formula> in state <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>s</mml:mi></mml:math></inline-formula>, while <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> track the selection frequency of specific actions and the total visitation count for the state, respectively. The first component of the expression prioritizes the exploitation of historically favorable actions. Specifically, the ratio <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula> serves as an empirical approximation of the expected reward <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This standard approach encounters two significant bottlenecks in long-horizon path planning characterized by sparse rewards and high-dimensional spaces. A primary issue is that meaningful reward signals typically only emerge upon goal attainment or collision. Because intermediate steps yield nearly uniform rewards, the exploitation term struggles to distinguish node quality during the early stages of the search. Furthermore, the exploration term relies exclusively on visitation statistics and lacks awareness of the goal&#x2019;s direction. In the absence of goal-oriented guidance, MCTS often requires an excessive number of iterations to locate the target. This leads to the inefficient expenditure of computational resources on irrelevant regions and necessitates a modified heuristic to improve search efficiency. The objective of the rollout phase is to approximate the value of newly expanded nodes by simulating trajectories toward terminal states using a default policy. By substituting exact evaluations with stochastic sampling, this process significantly accelerates the iteration cycle [<xref ref-type="bibr" rid="ref-58">58</xref>]. However, this approach faces two fundamental challenges. Increasing simulation fidelity introduces substantial computational overhead, which is a problem magnified by the vast search spaces inherent in path planning. Additionally, the extreme depth required to reach terminal states often renders purely random simulations computationally prohibitive. These constraints require a refined rollout policy that can effectively reconcile evaluation accuracy with limited computational budgets.</p>

<p><bold>Remark 1:</bold> <italic>The interpretability of MCTS is rooted in the structural clarity of its search tree and the transparency of its underlying node-level statistics. Unlike black-box models that rely on opaque computations, MCTS establishes decisions through the incremental expansion of a visible search space. The individual components of <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> clarify exactly why a specific action is prioritized. These terms provide a clear distinction between an action&#x2019;s historical performance and its exploration potential. Furthermore, each node preserves essential metrics such as visitation frequency and cumulative rewards. This organizational structure ensures that the entire decision chain remains traceable. By backtracking from the selected node to the root, the specific logic informing the final output can be fully audited and analyzed [<xref ref-type="bibr" rid="ref-59">59</xref>]</italic>.</p>

</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methods</title>
<p><bold>Assumption 1:</bold> <italic>The unmanned system in this paper is modeled as a point mass, denoted as an agent</italic>.</p>
<p>In this scenario, the agent must autonomously plan an optimal path from a given initial state to a target region, subject to its own kinematic constraints, environmental boundary constraints, and obstacle avoidance requirements, while simultaneously optimizing multiple objectives: path smoothness, efficiency, and safety. The core challenge lies in balancing exploration and exploitation under dynamically distributed obstacles and limited computational resources, ensuring that the planned path remains feasible in complex environments populated by polygonal obstacles and maximizes cumulative reward. This work achieves multi-objective path optimization, encompassing safe obstacle avoidance, path efficiency, and motion smoothness, through the synergistic integration of a quantified reward mechanism and a heuristic UCT strategy.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Path Planning Problem in MDP Modeling</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>State Definition</title>
<p>By default, the MCTS algorithm assumes discrete states and actions with finite cardinality. The instantaneous state vector of the agent is defined as:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> denotes the position vector of the agentcenter at time <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>t</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>y</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represent the coordinates along the horizontal and vertical axes, respectively; <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the heading angle of the agent at time <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>t</mml:mi></mml:math></inline-formula>, measured counterclockwise from the positive <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis with a range of <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This state vector comprises two core dimensions&#x2014;position and heading.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Dynamic Model</title>
<p>In order to reduce unnecessary complexity, the agent adopts a kinematic model of constant speed <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>v</mml:mi></mml:math></inline-formula> with discrete heading adjustments. The action is defined as the heading angle offset <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which corresponds to the angular velocity in continuous control. The discrete-time update rule over a time step <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> is as follows:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>v</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Constraints</title>
<p>Path planning for the agent is governed by four primary constraint categories. These include environmental boundaries, obstacle collisions, action space limitations, and the specific requirements of the MCTS search process. It is illustrated within a predefined two-dimensional rectangular map <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <italic>W</italic> and <italic>H</italic> represent the map width and height. To account for the agent&#x2019;s physical dimensions, a safety radius <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is utilized. This radius incorporates both the agent&#x2019;s body size and sensor uncertainty. Consequently, the agent must remain entirely within the map boundaries throughout the mission. The positional constraint for all <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is defined as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <italic>T</italic> denotes the termination time of the trajectory, and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> represents an additional safety margin. If the position exceeds the boundary, it is constrained to the valid range via a clamp operation.</p>
<p>The environment is modeled with a set of <italic>M</italic> static obstacles, denoted by <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mi>M</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. Each obstacle <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>O</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is defined as a closed polygon comprising <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>N</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> vertices, where the <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>k</mml:mi></mml:math></inline-formula>-th vertex is represented as <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mrow><mml:mtext>obs</mml:mtext><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mtext>obs</mml:mtext><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mtext>obs</mml:mtext><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula>. These vertices are ordered to form a contiguous boundary. To ensure navigation safety, the agent must maintain a minimum separation distance equal to its safety radius <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> from any obstacle edge. This collision avoidance constraint is expressed as <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> for all <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>M</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. The term <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the shortest Euclidean distance from the agent&#x2019;s position <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to the boundary of polygon <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>O</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>. Furthermore, no portion of the planned trajectory may lie within an obstacle. This condition is enforced by performing segment-wise sampling between consecutive path points to verify that all samples remain in free space. The agent begins its mission with an initial heading angle <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>init</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. Subsequent heading deviations <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are restricted by the vehicle&#x2019;s physical turning capabilities. To facilitate efficient computation, the continuous action space is discretized into a finite set <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow></mml:math></inline-formula> by uniformly sampling <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> values within the allowable range <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. Search efficiency and path quality are managed through three primary MCTS constraints. Computational expenditure is capped by a maximum iteration limit <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>mcts,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. Additionally, the search depth <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>d</mml:mi></mml:math></inline-formula> is restricted to <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>tree,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> to prevent the generation of excessively long paths that could compromise real-time performance. Finally, the search terminates once the agent reaches the goal region. This region is defined as a circle of radius <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>r</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula> centered at <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula>. The terminal condition <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula> provides the necessary planning flexibility while ensuring the agent successfully reaches the target neighborhood.</p>
<p>To accommodate environmental complexity, the proposed framework integrates specific constraints for dynamic obstacles based on their instantaneous velocity and projected trajectories. Unlike the treatment of static boundaries, these dynamic constraints are evaluated across a finite prediction horizon <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>h</mml:mi></mml:math></inline-formula>. The algorithm employs constant-velocity extrapolation to forecast the future positions of <italic>M</italic> moving obstacles. The predicted state for any dynamic obstacle <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>j</mml:mi></mml:math></inline-formula> at a subsequent time step is modeled as <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula>. This predictive model also incorporates boundary reflection to ensure realistic environmental dynamics. A proactive safety buffer is maintained by requiring the agent to satisfy a predictive minimum clearance <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>. This metric accounts for both the agent&#x2019;s safety radius <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and the specific radius <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the obstacle.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Optimization Objectives and Reward Mechanisms</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Heuristic UCT Strategy</title>
<p>The selection phase of MCTS requires choosing one child node from all children of the current node for expansion, with the UCT formula serving as the core decision criterion. As previously discussed, a common variant is the UCB1 formula, which relies solely on historical statistics of the nodes, whereas heuristic UCT incorporates domain knowledge to accelerate convergence [<xref ref-type="bibr" rid="ref-36">36</xref>]. To accommodate the domain-specific characteristics of agent path planning, a heuristic bias term <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is introduced, extending the formula to:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left" rowspacing="3pt" columnspacing="0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>UCB1</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>heuristic</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msqrt><mml:mfrac><mml:mrow><mml:mi>ln</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>parent</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is the weight coefficient for heuristic information, and the heuristic function <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is defined as:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left" rowspacing="3pt" columnspacing="0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1..</mml:mn><mml:mi>M</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the first term is the goal-directed bias: the closer to the goal, the higher the value, guiding the search toward the goal; the second term is the safety bias: the greater the safety margin from obstacles, the higher the value. <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> are heuristic coefficients (<inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>&#x03B1;</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> to ensure <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>), and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:msqrt></mml:math></inline-formula> denotes the maximum diagonal distance of the map. <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>3.0</mml:mn></mml:math></inline-formula> represents the safety distance threshold, indicating the minimum safe distance that should be maintained from the center of the agent to the boundary of obstacles.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Reward Definition</title>
<p>The core objective of path planning is to maximize the cumulative discounted reward, where the reward function integrates flight safety, path efficiency, heading smoothness, and task completion through a weighted sum. Let <italic>T</italic> denote the termination time of the trajectory, with the action sequence given by <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and the corresponding state sequence by <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. The cumulative reward <italic>J</italic> obtained during the rollout is then defined as:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>J</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mrow><mml:mtext>disc</mml:mtext></mml:mrow></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>step</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mrow><mml:mtext>disc</mml:mtext></mml:mrow></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>step</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denotes the step reward at time <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>t</mml:mi></mml:math></inline-formula>, reflecting the performance of a single action; <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> (no heading adjustment occurs before the initial action, so the smoothness reward is zero); the discount factor <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mtext>disc</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>; and <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the terminal reward, which is assigned if the goal is reached:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>len</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>path</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is the terminal reward constant, <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>len</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is the path-length penalty weight, and <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>path</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denotes the total path length; otherwise, <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>err</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, with <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>err</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> being the penalty weight for failing to reach the goal. The search process terminates when the agent successfully reaches the goal region, defined by <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>d</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula>, also when the maximum search depth <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>tree,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> or the maximum number of iterations <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> are exhausted. At the termination state, the cumulative reward is formulated as:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>step</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>dist</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>tangent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>prog</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Each item of the above formula jointly accounts for efficiency in approaching the goal, obstacle avoidance, and heading smoothness. Each component is detailed below. Specifically, <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>dist</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the distance-based reward, which increases as the agent gets closer to the goal, thereby encouraging progress toward the target. It is defined as:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>dist</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:math></inline-formula> denotes the distance to the goal at time <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>t</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>w</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is the weight for the distance-based reward. <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>tangent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the tangent reward, designed to encourage the agent to navigate along a tangent trajectory within a safe distance <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>min,clear</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> from obstacle boundaries, thereby reducing detour length when circumnavigating obstacles. This is necessary because path-efficiency rewards alone are insufficient to guide the agent toward such tangent paths. Let <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> be the tangent distance to the goal, and let <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>obs</mml:mtext><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the actual distance to obstacle <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>k</mml:mi></mml:math></inline-formula> that requires circumnavigation. Defining the tolerance as <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>obs</mml:mtext><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula>, the tangent reward is given by:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>tangent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>obs</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is the weight for the tangential reward. <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>prog</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denotes the progress reward, which incentivizes the agent for moving closer to the goal and provides no reward when moving away, defined as:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>prog</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is the weight for the progress reward. The smoothness reward <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is based on the change in heading deviation between consecutive actions. It penalizes large and frequent heading changes that increase energy consumption and reduce stability, thereby encouraging actions with small heading variations to ensure trajectory smoothness.
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>w</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>w</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> denote the weights for smoothness reward and smoothness penalty, respectively, and <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> represents the maximum heading deviation in the action space; for the initial action, <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. The path efficiency reward <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> encourages the agent to plan efficient paths by rewarding progress toward the goal per unit path length:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo></mml:math></inline-formula> denotes the path length traversed from time <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>t</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the efficiency reward weight.</p>
<p>The predictive dynamic avoidance reward <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> evaluates the future collision risk posed by moving obstacles over a finite prediction horizon <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>h</mml:mi></mml:math></inline-formula>. At each planning step, it computes the minimum clearance between the candidate position and all dynamic obstacles across future time steps <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> through <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>h</mml:mi></mml:math></inline-formula>, under constant-velocity motion prediction. This encourages the agent to preemptively deviate from trajectories that, while currently safe, will be occupied by approaching obstacles in the near future:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>1.3</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x003C;</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x03F5;</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2265;</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where the predicted minimum clearance is defined as:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:munder><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>new</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">p</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">p</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">p</mml:mi></mml:mrow><mml:mi>j</mml:mi><mml:mn>0</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> denotes the predicted position of the <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>j</mml:mi></mml:math></inline-formula>-th dynamic obstacle under constant-velocity extrapolation with boundary reflection, <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msub><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is its radius, <italic>M</italic> is the total number of dynamic obstacles, and <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mtext>safe</mml:mtext></mml:mrow><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the predictive safety margin. The exponent <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mn>1.3</mml:mn></mml:math></inline-formula> provides a superlinear penalty gradient that intensifies rapidly as the predicted clearance diminishes, while remaining less aggressive than the instantaneous clearance penalty to avoid excessive conservatism at longer prediction horizons. When the predicted clearance exceeds the safety threshold, a small positive bonus <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mo>&#x03F5;</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula> is awarded to mildly encourage trajectories that remain well-separated from future obstacle positions.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Overview of Proposed Method</title>
<p>Combining the aforementioned constraints and rewards, the agent&#x2019;s path planning problem can be formulated as follows: at each time step, find an action sequence <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and the corresponding state sequence <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> within the intersection of all constraints <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:math></inline-formula> that maximizes the cumulative reward <italic>J</italic>, i.e.:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:mi>J</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:mtext>s.t.&#xA0;</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mspace width="1em" /></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the dynamics update function, and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:math></inline-formula> represents the intersection of all constraints. As outlined in Algorithm 1, the algorithm initiates from the agent&#x2019;s current position and orientation, invoking the MCTS path planner to generate candidate trajectories. Only the first <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mtext>exec</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> steps of the selected candidate trajectory are executed to mitigate error accumulation from long-horizon planning. After execution, the agent&#x2019;s state is updated, and the planning process repeats from this new state, establishing a closed-loop &#x201C;plan&#x2013;execute&#x2013;replan&#x201D; cycle. The procedure terminates when the agent reaches the goal region or the maximum number of outer-loop iterations <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is reached, at which point the final executed trajectory is returned. This mechanism enables high-frequency replanning to adapt to complex environments while leveraging MCTS&#x2019;s global search capability to ensure trajectory feasibility.</p>
<p>The HRMCTS framework iteratively optimizes path planning through a structured four-phase cycle: selection, expansion, rollout, and backpropagation. During the selection phase, the algorithm traverses the search tree from the root using the UCT strategy. This approach maintains a principled balance between the exploitation of known high-value trajectories and the exploration of uncharted state spaces. Following the expansion of a leaf node, the rollout phase executes a rapid forward trajectory simulation up to a maximum depth of <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>rollout,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. To enhance computational efficiency and avoid the sub-optimality of purely random exploration, this phase incorporates a heuristic action selection strategy. At each simulation step <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>rollout,max</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, the action space is constrained by the target&#x2019;s relative bearing, <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>&#x03C8;</mml:mi><mml:mrow><mml:mtext>target</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>acrtan</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. By prioritizing actions that minimize this angular deviation, the algorithm imposes a forward branching constraint <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mrow><mml:mtext>forward</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> that effectively prunes search paths deviating significantly from the goal. Each step within the rollout strictly adheres to kinematic constraints via state transition equations while simultaneously performing boundary and collision checks. If the simulation detects a collision or a breach of a no-fly zone, it terminates prematurely and applies a penalty. The terminal cumulative quality of the trajectory is then quantified by <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. Finally, in the backpropagation phase, the total reward is propagated upward through the traversed path to update the value estimates of all ancestral nodes. Specifically, the visit count for each node in the path is incremented, and the node value is updated using the incremental mean formula:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mrow><mml:mtext>new</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mrow><mml:mtext>old</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mrow><mml:mtext>old</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>This iterative update ensures that the statistical value of each node converges toward the true expected reward, ultimately yielding the optimal, kinematically feasible decision for the current planning step.</p>
<fig id="fig-11">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-11.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Simulation Results and Analysis</title>
<p>In this section, we first conduct comparative experiments to demonstrate the effectiveness of the proposed HRMCTS method. We then evaluate the performance of MCTS under various parameter settings, with a particular comparison against the PSO algorithm. Highlighting the real-time capability, robustness, and rapid convergence of HRMCTS. These results substantiate the ability of MCTS to address sequential decision-based path planning problems. Specifically, we first describe our experimental setup, followed by test results across different tasks and conditions, and conclude with a performance analysis of HRMCTS. The simulation environment runs on a system equipped with an AMD Ryzen 7 5800H CPU, an NVIDIA GeForce RTX 3060 laptop GPU with 6 GB of memory.</p>
<p>To ensure a fair comparison, all algorithms were assigned identical path planning tasks. The parameters and iteration limits of the conventional MCTS and HRMCTS were set identically, and the fitness function of PSO was aligned with the rollout value evaluation function of the improved MCTS, ensuring that both algorithms shared the same optimization objective. All necessary parameters mentioned above are listed in <xref ref-type="table" rid="table-1">Table 1</xref>. Systematic experiments established the core parameters for the HRMCTS algorithm to balance real-time UAV constraints with safe navigation in complex environments. Following classical UCT theory, the exploration constant is defined as <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mn>1.4</mml:mn></mml:math></inline-formula>. To meet the real-time threshold of <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.8</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow></mml:math></inline-formula> per planning step, the maximum number of iterations is capped at <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>180</mml:mn></mml:math></inline-formula>, as additional iterations yield diminishing marginal improvements. Weighting follows a magnitude hierarchy designed around a &#x201C;safety-first, goal-oriented&#x201D; framework. A dominant terminal reward ensures that reaching the goal takes precedence over cumulative step-level gains. With <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>terminal</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2000</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>progress</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>18.0</mml:mn></mml:math></inline-formula>, the algorithm maintains a strong attraction toward the goal. This configuration ensures that a unit displacement of <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mn>1</mml:mn><mml:mtext>m</mml:mtext></mml:math></inline-formula> results in a positive increment of roughly <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mn>18.0</mml:mn></mml:math></inline-formula>, which effectively alleviates &#x201C;dead zone&#x201D; issues in cornering scenarios. Regarding dynamic safety, the predictive penalty weight <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>35.0</mml:mn></mml:math></inline-formula> is intentionally higher than the instantaneous clearance weight. This reflects the principle that anticipating future collisions is more critical than current proximity because early evasion is kinematically more efficient than last-moment maneuvers. The smoothness weight <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>6.0</mml:mn></mml:math></inline-formula> penalizes steering magnitude and inter-step fluctuations. This filters out kinematically infeasible zigzag paths while allowing for essential evasive actions. This design ensures that safety-critical objectives override efficiency when required. Meanwhile, the global path reference provides firm but flexible guidance that permits local deviations when encountering dynamic threats. A discount factor of <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.97</mml:mn></mml:math></inline-formula> ensures that cumulative rewards over a 20-step rollout remain significant at <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mn>20</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2248;</mml:mo><mml:mn>0.54</mml:mn></mml:math></inline-formula>, effectively balancing long-term planning with immediate collision avoidance. These parameters were validated through ablation studies spanning diverse scenarios with varying obstacle densities and velocities.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Algorithm parameters and details.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter Symbol</th>
<th>Meaning</th>
<th>Value</th>
<th>Unit</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="4"><bold>Map and Environment</bold></td>
</tr>
<tr>
<td><italic>W</italic>, <italic>H</italic></td>
<td>2D map width/height</td>
<td><inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>S</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>S</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>S</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>Start coordinates (vehicle)</td>
<td><inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mo stretchy="false">(</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mn>10</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>Target coordinates</td>
<td><inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mo stretchy="false">(</mml:mo><mml:mn>90</mml:mn><mml:mo>,</mml:mo><mml:mn>90</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:msub><mml:mi>r</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula></td>
<td>Target region radius (goal tolerance)</td>
<td><inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mn>5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="4"><bold>Vehicle Kinematics</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Agent collision radius</td>
<td><inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Max steering angle offset</td>
<td><inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>4</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mtext>rad</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Sampled steering angles count</td>
<td><inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mn>9</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Vehicle max linear velocity</td>
<td><inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mn>5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mtext>m/s</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula></td>
<td>State update time step</td>
<td><inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mtext>s</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>init</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Initial heading angle of vehicle</td>
<td><inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>4</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mtext>rad</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="4"><bold>Constraints</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mtext>bound</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Agent-map boundary margin</td>
<td><inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>min,clear</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Obstacle/boundary safe distance</td>
<td><inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mn>3.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="4"><bold>Reward and Penalty</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Distance-to-target reward weight</td>
<td><inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:mn>2.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>pr</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Progress-to-target reward weight</td>
<td><inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mn>18.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Trajectory smoothness reward weight</td>
<td><inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:mn>6.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>pred</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Predictive penalty weight</td>
<td><inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mn>35.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Obstacle tangent movement reward weight</td>
<td><inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mn>9.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Path efficiency reward weight</td>
<td><inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mn>8.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>term</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Terminal reward constant</td>
<td><inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mn>2000</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>len</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Overlong path penalty weight</td>
<td><inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mn>10.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>err</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Unreached target penalty weight</td>
<td><inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mn>30.0</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula></td>
<td>Cumulative reward discount factor</td>
<td><inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mn>0.95</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mrow><mml:mtext>tan</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Tangent reward clearance tolerance</td>
<td><inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:mn>0.55</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:mtext>m</mml:mtext></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="4"><bold>MCTS Algorithm</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mtext>explore</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>UCB1 exploration coefficient</td>
<td><inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:mn>1.4</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>rollout,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Single rollout max steps</td>
<td><inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:mn>35</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-224"><mml:math id="mml-ieqn-224"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mtext>forward</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Forward candidate branch limit</td>
<td><inline-formula id="ieqn-225"><mml:math id="mml-ieqn-225"><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-226"><mml:math id="mml-ieqn-226"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>mcts,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>MCTS max iterations per plan</td>
<td><inline-formula id="ieqn-227"><mml:math id="mml-ieqn-227"><mml:mn>120</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-228"><mml:math id="mml-ieqn-228"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>tree,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>MCTS tree max depth</td>
<td><inline-formula id="ieqn-229"><mml:math id="mml-ieqn-229"><mml:mn>200</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-230"><mml:math id="mml-ieqn-230"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>MCTS outer iterations</td>
<td><inline-formula id="ieqn-231"><mml:math id="mml-ieqn-231"><mml:mn>150</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-232"><mml:math id="mml-ieqn-232"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mtext>exec</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>MCTS execute horizon per step</td>
<td><inline-formula id="ieqn-233"><mml:math id="mml-ieqn-233"><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td align="center" colspan="4"><bold>PSO Algorithm</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-234"><mml:math id="mml-ieqn-234"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>pso,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>PSO max iterations</td>
<td><inline-formula id="ieqn-235"><mml:math id="mml-ieqn-235"><mml:mn>100</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-236"><mml:math id="mml-ieqn-236"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mtext>pso</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Number of particles (PSO)</td>
<td><inline-formula id="ieqn-237"><mml:math id="mml-ieqn-237"><mml:mn>20</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-238"><mml:math id="mml-ieqn-238"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>inert</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>PSO inertia weight</td>
<td><inline-formula id="ieqn-239"><mml:math id="mml-ieqn-239"><mml:mn>0.7</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-240"><mml:math id="mml-ieqn-240"><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula></td>
<td>PSO cognitive coefficient</td>
<td><inline-formula id="ieqn-241"><mml:math id="mml-ieqn-241"><mml:mn>1.5</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-242"><mml:math id="mml-ieqn-242"><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula></td>
<td>PSO social coefficient</td>
<td><inline-formula id="ieqn-243"><mml:math id="mml-ieqn-243"><mml:mn>1.5</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-244"><mml:math id="mml-ieqn-244"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>decay</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>PSO inertia weight decay rate</td>
<td><inline-formula id="ieqn-245"><mml:math id="mml-ieqn-245"><mml:mn>0.9</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-246"><mml:math id="mml-ieqn-246"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mtext>wp</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Number of waypoints</td>
<td><inline-formula id="ieqn-247"><mml:math id="mml-ieqn-247"><mml:mn>18</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-248"><mml:math id="mml-ieqn-248"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>stall,max</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>PSO max fitness stall iterations</td>
<td><inline-formula id="ieqn-249"><mml:math id="mml-ieqn-249"><mml:mn>8</mml:mn></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-250"><mml:math id="mml-ieqn-250"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mtext>fitness</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>PSO fitness stall tolerance</td>
<td><inline-formula id="ieqn-251"><mml:math id="mml-ieqn-251"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This paper focuses on several representative path planning parameters. Notably, all effective path planning results presented below are averaged over multiple simulation runs and thus represent typical algorithmic behavior. The conventional MCTS algorithm serves as a baseline to demonstrate the effectiveness of the proposed approach. As illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, the baseline failed to complete the path planning task in all 10 trials. For this study, a trial is defined as a &#x201C;failure&#x201D; if the algorithm cannot identify a collision-free path to the goal within a limit of 50,000 iterations. This shortcoming stems from the reliance on purely random sampling during the rollout phase, which is particularly problematic when the action space is constrained by kinematic requirements. In such high-dimensional environments, it is statistically improbable for a random policy to select a sequence of actions that consistently nears the goal. These rollouts often terminate in collisions or reach the iteration cap without success. Consequently, the algorithm fails to provide accurate value evaluations for state-action pairs. Without effective simulation-based estimates, the search tree is unable to converge on a feasible solution within a practical computational budget.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Comparison of random MCTS and HRMCTS results.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-2.tif"/>
</fig>
<p>The results indicate a significant improvement over conventional MCTS, featuring an average planning time of approximately 0.001 s per step. This duration is substantially shorter than the execution time step and ensures that planning concludes before execution commences. <xref ref-type="fig" rid="fig-2">Fig. 2b</xref> compares the computation time and path efficiency across 100 repeated simulations. The findings for HRMCTS and PSO represent stable outcomes verified through these 100 runs to remove the impact of randomness. PSO serves as a representative baseline for traditional heuristic algorithms, employing a fitness function identical to the MCTS rollout reward function as detailed in <xref ref-type="table" rid="table-1">Table 1</xref>. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>,<xref ref-type="fig" rid="fig-3">b</xref>, PSO achieves effective path planning with convergence occurring after roughly 30 iterations and a path efficiency of 94.48%. However, its computational cost is significant, averaging 211.8 s per run. Note that all reported computational costs, including this average, are derived from 10 independent repeated experiments to ensure statistical reliability.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>PSO path planning results.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-3.tif"/>
</fig>
<p>This study utilizes the A&#x002A; algorithm to generate the global reference path within a <inline-formula id="ieqn-252"><mml:math id="mml-ieqn-252"><mml:mn>100</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> continuous workspace. The environment is discretized into an occupancy grid with a resolution of <inline-formula id="ieqn-253"><mml:math id="mml-ieqn-253"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> units. A grid cell is classified as occupied if its Euclidean distance to the nearest polygonal obstacle boundary is less than the vehicle collision radius <inline-formula id="ieqn-254"><mml:math id="mml-ieqn-254"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mtext>agent</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>. Starting from the grid cell corresponding to the start position <inline-formula id="ieqn-255"><mml:math id="mml-ieqn-255"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>, the algorithm expands nodes across an 8-connected neighborhood. This process employs Euclidean distance as the heuristic function <inline-formula id="ieqn-256"><mml:math id="mml-ieqn-256"><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, where <inline-formula id="ieqn-257"><mml:math id="mml-ieqn-257"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula> denotes the goal. Finally, a goal tolerance of <inline-formula id="ieqn-258"><mml:math id="mml-ieqn-258"><mml:msub><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mrow><mml:mtext>goal</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is applied to define the completion of the path. Subsequent experiments were conducted in more complex environments, and the results in <xref ref-type="fig" rid="fig-4">Fig. 4a</xref> demonstrate that the proposed HRMCTS method can effectively perform path planning. In contrast, the PSO methods exhibite significant failure, yielding unsuccessful planning outcomes even after prolonged execution, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4b</xref>. This highlights the effectiveness and superiority of the proposed approach.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison of HRMCTS and PSO results in polygon obstacle environment.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-4.tif"/>
</fig>
<p>To further validate HRMCTS in challenging scenarios, simulation tests were conducted in dynamic environments where obstacles moved along predefined trajectories. These tests demanded high-fidelity online replanning and real-time responsiveness. To evaluate obstacle avoidance, a circular dynamic obstacle with a 2.5 m radius was introduced, moving at 1.6 m/s. Its initial position and velocity were set to intercept the agent at a narrow passage along the shortest path. This point is a critical location that the A&#x002A; algorithm inherently traverses. Since A&#x002A; cannot adjust its path online, a collision is inevitable in this scenario. In contrast, HRMCTS uses its receding-horizon replanning and heuristic guidance to perceive motion and adjust the trajectory dynamically. This rolling-horizon mechanism is well-suited for time-varying environments because it allows for continuous updates to maintain safety. For a quantitative assessment, a reinforcement learning baseline using a Double Deep Q-Network (ddQn) was also implemented. The ddQn agent observes a 22-dimensional state vector <inline-formula id="ieqn-259"><mml:math id="mml-ieqn-259"><mml:msub><mml:mrow><mml:mi mathvariant="bold">s</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> involving position, heading, goal-relative metrics, and 8-directional obstacle-distance rays. It outputs actions from a discrete space of <inline-formula id="ieqn-260"><mml:math id="mml-ieqn-260"><mml:msub><mml:mi>N</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>9</mml:mn></mml:math></inline-formula> steering angles within <inline-formula id="ieqn-261"><mml:math id="mml-ieqn-261"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>4</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The Q-function is parameterized by a <inline-formula id="ieqn-262"><mml:math id="mml-ieqn-262"><mml:mn>22</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>256</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>128</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> fully connected network and was trained via the Adam optimizer (<inline-formula id="ieqn-263"><mml:math id="mml-ieqn-263"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) over 2000 episodes. To ensure a fair comparison, both planners share identical kinematics, obstacle representations, and reward functions within the same MPC rolling-horizon framework. Training performance for the RL method is illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The significant drop in success rate during training likely stems from hard target network updates and stale samples in the replay buffer. Hard updates occur every 20 episodes and cause sudden jumps in target Q-values. These jumps increase TD error and disrupt the learned policy. Furthermore, stale samples from early exploration exacerbate estimation bias. The success rate eventually recovers as the network adapts to new target values and higher-quality experiences replace older samples.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The gray vertical line represents the corresponding fluctuation range, and the solid line represents the average value. (<bold>a</bold>) The trend of training reward changes. (<bold>b</bold>) The trend of single round task steps. (<bold>c</bold>) The trend of training loss changes. (<bold>d</bold>) The trend of task success rate changes.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-5.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-6">Fig. 6a</xref> illustrates the optimal dynamic avoidance performance for the circular obstacle scenario, confirming that both HRMCTS and RL algorithms successfully achieve collision avoidance. Analysis of the execution times in <xref ref-type="fig" rid="fig-6">Fig. 6b</xref> shows that the computational cost of HRMCTS is highly targeted. The algorithm performs intensive calculations only when a dynamic obstacle is in close proximity and poses an immediate threat. This selective need for resources demonstrates the efficiency of the hierarchical framework. In comparison, the RL method requires planning times more than an order of magnitude shorter than those of HRMCTS. Data from ten independent trials in <xref ref-type="fig" rid="fig-7">Fig. 7a</xref>,<xref ref-type="fig" rid="fig-7">b</xref> further establish the RL approach as highly efficient and deterministic. This consistency is a primary advantage stemming from its extensive offline training phase.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Dynamic obstacle avoidance results and per-step running time of HRMCTS and RL under circular obstacle conditions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>10-path statistics and operation result statistics: dynamic obstacle avoidance results of HRMCTS and RL under circular obstacle conditions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-7.tif"/>
</fig>
<p>Performance results diverge significantly in complex environments featuring polygonal obstacles. As illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8a</xref>,<xref ref-type="fig" rid="fig-8">b</xref>, HRMCTS consistently finished planning tasks. This stability underscores the algorithm&#x2019;s inherent robustness. Conversely, the RL algorithm fails across all trials in this setting, as shown in <xref ref-type="fig" rid="fig-9">Fig. 9a</xref>. While HRMCTS&#x2019;s performance characteristics, including mean and standard deviation, remain highly consistent with the circular obstacle results shown in <xref ref-type="fig" rid="fig-9">Fig. 9b</xref>. This failure is primarily due to the limited generalization of the RL policy. Since the model was trained only in static circular environments, it cannot adapt to the combined demands of dynamic avoidance and complex polygonal constraints. These findings suggest that while HRMCTS requires a slightly higher computational overhead, it provides greater utility for engineering applications. Its superior generalization and &#x201C;anytime&#x201D; characteristics allow it to handle complex, real-world dynamic scenarios effectively without environment-specific training.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Dynamic obstacle avoidance results and per-step running time of HRMCTS and RL under polygongal obstacle conditions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-8.tif"/>
</fig><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>10-path statistics and operation result statistics: dynamic obstacle avoidance results of HRMCTS and RL under polygonal obstacle conditions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-9.tif"/>
</fig>
<p>To evaluate the computational scalability of the MCTS planner, a series of experiments examined various action-space granularities by adjusting the number of actions <inline-formula id="ieqn-264"><mml:math id="mml-ieqn-264"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. As illustrated in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>, these tests explored the trade-off between the path quality benefits of increased action flexibility and the computational overhead caused by a larger branching factor. Each configuration underwent 100 independent trials within a complex polygonal obstacle environment. Results show that as <inline-formula id="ieqn-265"><mml:math id="mml-ieqn-265"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> scales from 5 to 15, both per-step and total planning times exhibit a clear upward trend. Computational cost and its variance reach a peak at <inline-formula id="ieqn-266"><mml:math id="mml-ieqn-266"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>15</mml:mn></mml:math></inline-formula> due to the drastic expansion of the search space. Reliability also varied significantly. Success rates remained above 98% for <inline-formula id="ieqn-267"><mml:math id="mml-ieqn-267"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mn>7</mml:mn></mml:math></inline-formula> but dropped to approximately 80% at <inline-formula id="ieqn-268"><mml:math id="mml-ieqn-268"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>. This performance gap is primarily caused by &#x201C;control blind spots&#x201D; resulting from sparse action discretization. Limited steering options provide insufficient degrees of freedom for navigating sharp vertices or narrow passages. This constraint often prevents the search tree from finding feasible solutions within a finite depth. Although a higher action count can reduce total path distance, the improvement at <inline-formula id="ieqn-269"><mml:math id="mml-ieqn-269"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>15</mml:mn></mml:math></inline-formula> comes at an excessive temporal cost. A value of <inline-formula id="ieqn-270"><mml:math id="mml-ieqn-270"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>9</mml:mn></mml:math></inline-formula> maintains high success rates and path efficiency while offering much better cost-effectiveness. Because its planning times are significantly lower than those of higher-dimensional spaces, <inline-formula id="ieqn-271"><mml:math id="mml-ieqn-271"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>9</mml:mn></mml:math></inline-formula> is selected as the optimal granularity for satisfying both robustness and real-time requirements.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Scalability analysis of numbers of action space. (<bold>a</bold>) Planning time vs. action count. (<bold>b</bold>) Per-step time vs. action count. (<bold>c</bold>) Path distance vs. action count. (<bold>d</bold>) Success rate vs. action count. (<bold>e</bold>) Path efficiency vs. action count. (<bold>f</bold>) Time&#x2013;quality trade-off.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79895-fig-10.tif"/>
</fig>
<p>Finally, each method was independently executed 100 times to generate the comparative results summarized in <xref ref-type="table" rid="table-2">Table 2</xref>. These findings reveal significant differences in path planning performance among the three approaches. The proposed HRMCTS method demonstrates superior adaptability compared to raw MCTS and ddQn, particularly within complex dynamic environments. It also achieves better computational efficiency than the traditional PSO heuristic, reaching a performance level comparable to the A&#x002A; algorithm. Ultimately, the experimental results indicate that HRMCTS delivers high-quality path planning while balancing efficiency and interpretability through its inherent tree-based structure.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comprehensive comparison of path planning methods across different environments.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Environment</th>
<th>Dynamic Obs.</th>
<th>Success Rate</th>
<th>Path Efficiency</th>
<th>Total Execution Time</th>
<th>Path Length Std.</th>
</tr>
</thead>
<tbody>
<tr>
<td>Raw MCTS</td>
<td>Any</td>
<td>No</td>
<td>0% (Failed)</td>
<td>N/A</td>
<td>Fail</td>
<td>N/A</td>
</tr>
<tr>
<td>PSO</td>
<td>Circular</td>
<td>No</td>
<td>High</td>
<td>94.18%</td>
<td>211.8s</td>
<td>5.15</td>
</tr>
<tr>
<td>PSO</td>
<td>Complex Polygonal</td>
<td>No</td>
<td>0% (Failed)</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
</tr>
<tr>
<td>A&#x002A;</td>
<td>Circular</td>
<td>No</td>
<td>100%</td>
<td>97.0%</td>
<td>0.004s</td>
<td>0</td>
</tr>
<tr>
<td>A&#x002A;</td>
<td>Complex Polygonal</td>
<td>No</td>
<td>100%</td>
<td>88.8%</td>
<td>0.001s</td>
<td>0</td>
</tr>
<tr>
<td>RL (ddQn)</td>
<td>Circular</td>
<td>Yes</td>
<td>100%</td>
<td>90.6%</td>
<td>0.099s</td>
<td>0</td>
</tr>
<tr>
<td>RL (ddQn)</td>
<td>Complex Polygonal</td>
<td>Yes</td>
<td>0% (Failed)</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
</tr>
<tr>
<td>HRMCTS</td>
<td>Any</td>
<td>No</td>
<td>100%</td>
<td>96.7%</td>
<td>0.9s</td>
<td>0.0043</td>
</tr>
<tr>
<td>HRMCTS</td>
<td>Circular</td>
<td>Yes</td>
<td>100%</td>
<td>83.96%</td>
<td>7.60s</td>
<td>0.596</td>
</tr>
<tr>
<td>HRMCTS</td>
<td>Complex Polygonal</td>
<td>Yes</td>
<td>100%</td>
<td>78.0%</td>
<td>7.37s</td>
<td>1.17</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper introduces the HRMCTS framework for unmanned system path planning in complex environments. The approach enhances MCTS performance through several critical innovations. First, a heuristic UCT selection strategy incorporates goal distance and obstacle clearance to improve search directionality and convergence. Second, a multi-objective reward function applied during the rollout phase optimizes path length, goal progress, obstacle avoidance, trajectory smoothness, and time efficiency. This ensures a balanced trade-off between safety and performance. Third, a receding-horizon mechanism integrates local path generation with online replanning to maintain global feasibility while ensuring real-time responsiveness. To systematically evaluate the method, simulation experiments were conducted in four scenarios: static and dynamic environments with both circular and polygonal obstacles. The proposed framework is compared against standard MCTS, PSO, and ddQn baselines. The experimental results demonstrate:<list list-type="bullet">
<list-item>
<p>HRMCTS achieves a 100% success rate and 96.7% path efficiency in static environments, notably outperforming raw MCTS, which fails across all trials. The proposed method also provides a substantial speed advantage over PSO, which requires 211.8 s to complete a single run. While the A&#x002A; algorithm establishes the performance ceiling with 97.0% efficiency, HRMCTS produces nearly identical results and maintains a minimal path length std of 0.0043.</p></list-item>
<list-item>
<p>Both HRMCTS and the ddQn achieve 100% success in dynamic circular environments. Although ddQn yields a higher path efficiency of 90.6% in circular environment due to its offline training, HRMCTS demonstrates robust adaptability by achieving success in both circular and polygonal environments. This outcome is particularly significant because HRMCTS operates effectively without any prior knowledge of the environment or the obstacle field.</p></list-item>
<list-item>
<p>When navigating complex polygonal dynamic environments, the performance of alternative methods diverges significantly as both ddQn and PSO fail to resolve the planning tasks. Conversely, HRMCTS maintains a 100% success rate with a stable path efficiency of 78.0% and a relatively low standard deviation. These findings confirm the superior generalization and reliability of the HRMCTS framework when addressing the coupled constraints of intricate geometries and moving obstacles.</p></list-item>
</list></p>
<p>HRMCTS effectively improves the efficiency constraints found in traditional MCTS when navigating sparse-reward and high-dimensional environments. The framework consistently demonstrates superior planning performance in some scenarios relative to traditional methods such as PSO, A&#x002A; and ddQn. By utilizing a transparent tree structure alongside node-level statistical metrics, the algorithm ensures that the decision-making process is both interpretable and traceable. This approach successfully combines the explainability typically absent in RL with the dynamic adaptability that conventional search methods frequently lack. Moreover, the &#x201C;anytime&#x201D; planning capabilities of the algorithm facilitate efficient real-time responses even under limited computational resources. This functionality maintains a necessary balance between engineering practicality and operational reliability. This makes HRMCTS particularly well-suited for dynamic scenarios characterized by dense obstacles and rapidly changing states, such as urban low-altitude drone navigation and multi-agent coordination in warehouse logistics, where it enables rapid obstacle avoidance and local re-planning based on limited environmental perception. In contrast to black-box models that rely on extensive training, sampling strategies that lag in abruptly changing environments, or graph search methods dependent on structured settings, HRMCTS preserves decision transparency while delivering enhanced robustness and real-time performance, making it a good choice for safety-critical autonomous systems operating in unpredictable environments that demand human-machine collaborative trust.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Hunan Provincial Natural Science Foundation of China (No. 2025JJ60072), the Open Research Subject of State Key Laboratory of Intelligent Game (No. ZBKF-24-01), the Postdoctoral Fellowship Program of CPSF (No. GZB20240989), the China Postdoctoral Science Foundation (No. 2024M754304) and the Military High-Level Talent Program (No. 202401-RCGC-ZZ-004).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Qianshu Yang: conceptualization, methodology design, data collection and manuscript drafting; Shuangxi Liu: supervision, project administration, funding acquisition and manuscript revision; Xianyu Wu: formal analysis, simulation implementation and result verification; Wei Zhao: resource support, technical guidance and review of the final manuscript. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Some or all data, models, or code generated or used during the study are proprietary or confidential in nature and may only be provided with restrictions.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Dynamic integration of Q-learning and A-APF for efficient path planning in complex underground mining environments</article-title>. <source>Comput Mater Contin</source>. <year>2026</year>;<volume>86</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.071319</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A UAV path-planning approach for urban environmental event monitoring</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>83</volume>(<issue>3</issue>):<fpage>5575</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.061954</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xing</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>An algorithm of complete coverage path planning for deep-sea mining vehicle clusters based on reinforcement learning</article-title>. <source>Adv Theory Simul</source>. <year>2024</year>;<volume>7</volume>(<issue>4</issue>):<fpage>2300970</fpage>. doi:<pub-id pub-id-type="doi">10.1002/adts.202300970</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rivi&#x00E8;re</surname> <given-names>B</given-names></string-name>, <string-name><surname>Lathrop</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>SJ</given-names></string-name></person-group>. <article-title>Monte Carlo tree search with spectral expansion for planning with dynamical systems</article-title>. <source>Sci Robot</source>. <year>2024</year>;<volume>9</volume>(<issue>97</issue>):<fpage>eado1010</fpage>. doi:<pub-id pub-id-type="doi">10.1126/scirobotics.ado1010</pub-id>; <pub-id pub-id-type="pmid">39630878</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Technical development and future prospects of cooperative terminal guidance based on knowledge graph analysis: a review</article-title>. <source>J Zhejiang Univ-SCI A</source>. <year>2025</year>;<volume>26</volume>(<issue>7</issue>):<fpage>605</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1631/jzus.a2400317</pub-id>; <pub-id pub-id-type="pmid">39308071</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>M</given-names></string-name></person-group>. <article-title>COM trajectory planning and disturbance-resistant control of a bipedal robot based on CP-ZMP-COM dynamics</article-title>. <source>J Zhejiang Univ-SCI A</source>. <year>2025</year>;<volume>26</volume>(<issue>5</issue>):<fpage>492</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1631/jzus.a2400062</pub-id>; <pub-id pub-id-type="pmid">39308071</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sheng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Path planning for the dynamic UAV-aided wireless systems using monte carlo tree search</article-title>. <source>IEEE Trans Veh Technol</source>. <year>2022</year>;<volume>71</volume>(<issue>6</issue>):<fpage>6716</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tvt.2022.3160746</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Ascent trajectory optimization for hypersonic vehicle based on improved chicken swarm optimization</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>151836</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2947297</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Obstacle avoidance path planning method for intelligent vehicle based on obstacle projection</article-title>. <source>Meas Sci Technol</source>. <year>2025</year>;<volume>36</volume>(<issue>9</issue>):<fpage>096215</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ae0508</pub-id>; <pub-id pub-id-type="pmid">30793291</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Su</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Trajectory optimization of hypersonic periodic cruise using an improved PSO algorithm</article-title>. <source>Int J Aerosp Eng</source>. <year>2021</year>;<volume>2021</volume>(<issue>1</issue>):<fpage>2526916</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/2526916</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Transformer in reinforcement learning for decision-making: a survey</article-title>. <source>Front Inf Technol Electron Eng</source>. <year>2024</year>;<volume>25</volume>:<fpage>763</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1631/FITEE.2300548</pub-id>; <pub-id pub-id-type="pmid">39308071</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Csipp&#x00E1;n</surname> <given-names>G</given-names></string-name>, <string-name><surname>P&#x00E9;ter</surname> <given-names>I</given-names></string-name>, <string-name><surname>K&#x0151;v&#x00E1;ri</surname> <given-names>B</given-names></string-name>, <string-name><surname>B&#x00E9;csi</surname> <given-names>T</given-names></string-name></person-group>. <article-title>MCTS-based policy improvement for reinforcement learning</article-title>. <source>Mach Learn Know Extr</source>. <year>2025</year>;<volume>7</volume>(<issue>3</issue>):<fpage>98</fpage>. doi:<pub-id pub-id-type="doi">10.3390/make7030098</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Evolving adaptive and interpretable decision trees for cooperative submarine search</article-title>. <source>Def Technol</source>. <year>2025</year>;<volume>48</volume>:<fpage>83</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dt.2025.02.007</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Browne</surname> <given-names>CB</given-names></string-name>, <string-name><surname>Powley</surname> <given-names>E</given-names></string-name>, <string-name><surname>Whitehouse</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lucas</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Cowling</surname> <given-names>PI</given-names></string-name>, <string-name><surname>Rohlfshagen</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey of monte carlo tree search methods</article-title>. <source>IEEE Trans Comput Intell AI Games</source>. <year>2012</year>;<volume>4</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tciaig.2012.2186810</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Trotti</surname> <given-names>F</given-names></string-name>, <string-name><surname>Farinelli</surname> <given-names>A</given-names></string-name>, <string-name><surname>Muradore</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Path re-planning with stochastic obstacle modeling: a monte carlo tree search approach</article-title>. In: <conf-name>Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2024 Oct 14&#x2013;18</conf-name>; <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>. p. <fpage>8017</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sukwadi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Airlangga</surname> <given-names>G</given-names></string-name>, <string-name><surname>Basuki</surname> <given-names>WW</given-names></string-name>, <string-name><surname>Kristian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Rahmananta</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sugianto</surname> <given-names>LF</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Comparative analysis of path planning algorithms for multi-UAV systems in dynamic and cluttered environments: a focus on efficiency, smoothness, and collision avoidance</article-title>. <source>Int J Robot Cont Syst</source>. <year>2024</year>;<volume>4</volume>(<issue>4</issue>):<fpage>1602</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.31763/ijrcs.v4i4.1555</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bonanni</surname> <given-names>L</given-names></string-name>, <string-name><surname>Meli</surname> <given-names>D</given-names></string-name>, <string-name><surname>Castellini</surname> <given-names>A</given-names></string-name>, <string-name><surname>Farinelli</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Monte carlo tree search with velocity obstacles for safe and efficient motion planning in dynamic environments</article-title>. <comment>arXiv:2501.09649. 2025</comment>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Dynamic environment path planning based on hybrid pre-training algorithm</article-title>. <source>Eng Res Exp</source>. <year>2025</year>;<volume>7</volume>(<issue>3</issue>):<fpage>035267</fpage>. doi:<pub-id pub-id-type="doi">10.1088/2631-8695/adf93f</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Dynamic scene path planning of UAVs based on deep reinforcement learning</article-title>. <source>Drones</source>. <year>2024</year>;<volume>8</volume>(<issue>2</issue>):<fpage>60</fpage>. doi:<pub-id pub-id-type="doi">10.3390/drones8020060</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>CK</given-names></string-name></person-group>. <source>Adaptive long-range UAV flight planning using Monte Carlo search trees</source>. <publisher-loc>Vancouver, BC, Canada</publisher-loc>: <publisher-name>University of British Columbia</publisher-name>; <year>2026</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Swarm-based cooperative coverage strategy for amphibious unmanned systems in complex wetland environments</article-title>. <source>IFAC-PapersOnLine</source>. <year>2025</year>;<volume>59</volume>(<issue>20</issue>):<fpage>2388</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ifacol.2025.11.516</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Page</surname> <given-names>BR</given-names></string-name>, <string-name><surname>Moridian</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mahmoudian</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Collaborative mission planning for long-term operation considering energy limitations</article-title>. <source>IEEE Robot Autom Lett</source>. <year>2020</year>;<volume>5</volume>(<issue>3</issue>):<fpage>4751</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/lra.2020.3003881</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A USV-UAV cooperative trajectory planning algorithm with hull dynamic constraints</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>4</issue>):<fpage>1845</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23041845</pub-id>; <pub-id pub-id-type="pmid">36850442</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>M</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhuang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Si</surname> <given-names>J</given-names></string-name>, <string-name><surname>L&#x00FC;</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Collaborative marine targets search algorithm for USVs and AAVs under energy constraint</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>12</volume>(<issue>9</issue>):<fpage>12137</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2024.3520176</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>B</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Multi-AUV dynamic cooperative path planning with hybrid particle swarm and dynamic window algorithm in Three-Dimensional terrain and ocean current environment</article-title>. <source>Biomimetics</source>. <year>2025</year>;<volume>10</volume>(<issue>8</issue>):<fpage>536</fpage>. doi:<pub-id pub-id-type="doi">10.3390/biomimetics10080536</pub-id>; <pub-id pub-id-type="pmid">40862909</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>H&#x00E4;usler</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Saccon</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aguiar</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Hauser</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pascoal</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Cooperative motion planning for multiple autonomous marine vehicles</article-title>. <source>IFAC Proc Vol</source>. <year>2012</year>;<volume>45</volume>(<issue>27</issue>):<fpage>244</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>&#x015A;wiechowski</surname> <given-names>M</given-names></string-name>, <string-name><surname>Godlewski</surname> <given-names>K</given-names></string-name>, <string-name><surname>Sawicki</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ma&#x0144;dziuk</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Monte Carlo tree search: a review of recent modifications and applications</article-title>. <source>Artif Intell Rev</source>. <year>2023</year>;<volume>56</volume>(<issue>3</issue>):<fpage>2497</fpage>&#x2013;<lpage>562</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-022-10228-y</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kemmerling</surname> <given-names>M</given-names></string-name>, <string-name><surname>L&#x00FC;tticke</surname> <given-names>D</given-names></string-name>, <string-name><surname>Schmitt</surname> <given-names>RH</given-names></string-name></person-group>. <article-title>Beyond games: a systematic review of neural Monte Carlo tree search applications</article-title>. <source>Appl Intell</source>. <year>2024</year>;<volume>54</volume>(<issue>1</issue>):<fpage>1020</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-023-05240-w</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Coulom</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Computing &#x201C;elo ratings&#x201D; of move patterns in the game of go</article-title>. <source>ICGA J</source>. <year>2007</year>;<volume>30</volume>(<issue>4</issue>):<fpage>198</fpage>&#x2013;<lpage>208</lpage>. doi:<pub-id pub-id-type="doi">10.3233/icg-2007-30403</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Keller</surname> <given-names>T</given-names></string-name>, <string-name><surname>Eyerich</surname> <given-names>P</given-names></string-name></person-group>. <article-title>PROST: probabilistic planning based on UCT</article-title>. <source>Proc Int Conf Autom Plan Sched</source>. <year>2012</year>;<volume>22</volume>:<fpage>119</fpage>&#x2013;<lpage>27</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Pruning in UCT algorithm</article-title>. In: <conf-name>Proceedings of the 2010 International Conference on Technologies and Applications of Artificial Intelligence; 2010 Nov 18&#x2013;20</conf-name>; <publisher-loc>Hsinchu, Taiwan</publisher-loc>. p. <fpage>177</fpage>&#x2013;<lpage>81</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Baier</surname> <given-names>H</given-names></string-name>, <string-name><surname>Winands</surname> <given-names>MH</given-names></string-name></person-group>. <article-title>Beam monte-carlo tree search</article-title>. In: <conf-name>Proceedings of the 2012 IEEE Conference on Computational Intelligence and Games (CIG); 2012 Sep 11&#x2013;14</conf-name>; <publisher-loc>Granada, Spain</publisher-loc>. p. <fpage>227</fpage>&#x2013;<lpage>33</lpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vien</surname> <given-names>NA</given-names></string-name>, <string-name><surname>Toussaint</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Hierarchical monte-carlo planning</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2015</year>;<volume>29</volume>(<issue>1</issue>):<fpage>3613</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v29i1.9687</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>De Waard</surname> <given-names>M</given-names></string-name>, <string-name><surname>Roijers</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Bakkes</surname> <given-names>SC</given-names></string-name></person-group>. <article-title>Monte carlo tree search with options for general video game playing</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computational Intelligence and Games (CIG); 2016 Sep 20&#x2013;23</conf-name>; <publisher-loc>Santorini, Greece</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Planning of heuristics: strategic planning on large language models with monte carlo tree search for automating heuristic optimization</article-title>. <comment>arXiv:2502.11422. 2025</comment>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Manome</surname> <given-names>N</given-names></string-name>, <string-name><surname>Shinohara</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ui</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Simple modification of the upper confidence bound algorithm by generalized weighted averages</article-title>. <source>PLoS One</source>. <year>2025</year>;<volume>20</volume>(<issue>5</issue>):<fpage>e0322757</fpage>. doi:<pub-id pub-id-type="doi">10.1371/journal.pone.0322757</pub-id>; <pub-id pub-id-type="pmid">40334251</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jeon</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>GH</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>KE</given-names></string-name></person-group>. <article-title>Monte-carlo tree search in continuous action spaces with value gradients</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>:<fpage>4561</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i04.5885</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Finnsson</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bj&#x00F6;rnsson</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Learning simulation control in general game-playing agents</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2010</year>;<volume>24</volume>(<issue>1</issue>):<fpage>954</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v24i1.7651</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lorentz</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Using evaluation functions in Monte-Carlo tree search</article-title>. <source>Theor Comput Sci</source>. <year>2016</year>;<volume>644</volume>(<issue>4</issue>):<fpage>106</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.tcs.2016.06.026</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wal&#x0119;dzik</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ma&#x0144;dziuk</surname> <given-names>J</given-names></string-name></person-group>. <article-title>An automatically generated evaluation function in general game playing</article-title>. <source>IEEE Trans Comput Intell AI Games</source>. <year>2013</year>;<volume>6</volume>(<issue>3</issue>):<fpage>258</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tciaig.2013.2286825</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Painter</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lacerda</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hawes</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Convex hull monte-carlo tree-search</article-title>. <source>Proc Int Conf Autom Plan Sched</source>. <year>2020</year>;<volume>30</volume>:<fpage>217</fpage>&#x2013;<lpage>25</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>F</given-names></string-name>, <string-name><surname>Nakhost</surname> <given-names>H</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A local monte carlo tree search approach in deterministic planning</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2011</year>;<volume>25</volume>:<fpage>1832</fpage>&#x2013;<lpage>3</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rosin</surname> <given-names>CD</given-names></string-name></person-group>. <article-title>Nested rollout policy adaptation for Monte Carlo tree search</article-title>. In: <conf-name>Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence; 2011 Jul 16&#x2013;22</conf-name>; <publisher-loc>Barcelona Catalonia, Spain</publisher-loc>. p. <fpage>649</fpage>&#x2013;<lpage>54</lpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nag&#x00F3;rko</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Parallel nested rollout policy adaptation</article-title>. In: <conf-name>Proceedings of the 2019 IEEE Conference on Games (CoG); 2019 Aug 20&#x2013;23</conf-name>; <publisher-loc>London, UK</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Silver</surname> <given-names>D</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Maddison</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Guez</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sifre</surname> <given-names>L</given-names></string-name>, <string-name><surname>Van Den Driessche</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mastering the game of Go with deep neural networks and tree search</article-title>. <source>Nature</source>. <year>2016</year>;<volume>529</volume>(<issue>7587</issue>):<fpage>484</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1038/nature16961</pub-id>; <pub-id pub-id-type="pmid">26819042</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Goodman</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Re-determinizing MCTS in Hanabi</article-title>. In: <conf-name>Proceedings of the 2019 IEEE Conference on Games (CoG); 2019 Aug 20&#x2013;23</conf-name>; <publisher-loc>London, UK</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Childs</surname> <given-names>BE</given-names></string-name>, <string-name><surname>Brodeur</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Kocsis</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Transpositions and move groups in monte carlo tree search</article-title>. In: <conf-name>Proceedings of the 2008 IEEE Symposium On Computational Intelligence and Games; 2008 Dec 15&#x2013;18</conf-name>; <publisher-loc>Perth, WA, Australia</publisher-loc>. p. <fpage>389</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cowling</surname> <given-names>PI</given-names></string-name>, <string-name><surname>Powley</surname> <given-names>EJ</given-names></string-name>, <string-name><surname>Whitehouse</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Information set monte carlo tree search</article-title>. <source>IEEE Trans Comput Intell AI Games</source>. <year>2012</year>;<volume>4</volume>(<issue>2</issue>):<fpage>120</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hayward</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Three-head neural network architecture for monte carlo tree search</article-title>. In: <conf-name>Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence; 2018 Jul 13&#x2013;19</conf-name>; <publisher-loc>Stockholm, Sweden</publisher-loc>. p. <fpage>3762</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lucas</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Samothrakis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Perez</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Fast evolutionary adaptation for monte carlo tree search</article-title>. In: <conf-name>European Conference on the Applications of Evolutionary Computation</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2014</year>. p. <fpage>349</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Baier</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cowling</surname> <given-names>PI</given-names></string-name></person-group>. <article-title>Evolutionary MCTS for multi-action adversarial games</article-title>. In: <conf-name>Proceedings of the 2018 IEEE Conference on Computational Intelligence and Games (CIG); 2018 Aug 14&#x2013;17</conf-name>; <publisher-loc>Maastricht, the Netherlands</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chaslot</surname> <given-names>GMB</given-names></string-name>, <string-name><surname>Winands</surname> <given-names>MH</given-names></string-name>, <string-name><surname>van Den Herik</surname> <given-names>HJ</given-names></string-name></person-group>. <article-title>Parallel monte-carlo tree search</article-title>. In: <conf-name>Proceedings of the International Conference on Computers and Games</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2008</year>. p. <fpage>60</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Enzenberger</surname> <given-names>M</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>M</given-names></string-name></person-group>. <chapter-title>A lock-free multithreaded Monte-Carlo tree search algorithm</chapter-title>. In: <source>Advances in computer games</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2009</year>. p. <fpage>14</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Barriga</surname> <given-names>NA</given-names></string-name>, <string-name><surname>Stanescu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Buro</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Parallel UCT search on GPUs</article-title>. In: <conf-name>Proceedings of the 2014 IEEE Conference on Computational Intelligence and Games; 2014 Aug 26&#x2013;29</conf-name>; <publisher-loc>Dortmund, Germany</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mirsoleimani</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Plaat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Van Den Herik</surname> <given-names>J</given-names></string-name>, <string-name><surname>Vermaseren</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Scaling monte carlo tree search on intel xeon phi</article-title>. In: <conf-name>Proceedings of the 2015 IEEE 21st International Conference on Parallel and Distributed Systems (ICPADS); 2015 Dec 14&#x2013;17</conf-name>; <publisher-loc>Melbourne, VIC, Australia</publisher-loc>. p. <fpage>666</fpage>&#x2013;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nayak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nissimagoudar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Revankar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Giraddi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hosamani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Iyer</surname> <given-names>NC</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Accelerating decision-making in AI: parallelizing monte carlo tree search for connect 4 using CPU and GPU</article-title>. <source>Proc Comput Sci</source>. <year>2025</year>;<volume>263</volume>:<fpage>82</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2025.07.011</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kocsis</surname> <given-names>L</given-names></string-name>, <string-name><surname>Szepesv&#x00E1;ri</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Bandit based monte-carlo planning</article-title>. In: <conf-name>Proceedings of the European Conference on Machine Learning</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2006</year>. p. <fpage>282</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chaslot</surname> <given-names>GMJ</given-names></string-name>, <string-name><surname>Winands</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Herik</surname> <given-names>HJVD</given-names></string-name>, <string-name><surname>Uiterwijk</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Bouzy</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Progressive strategies for Monte-Carlo tree search</article-title>. <source>New Mathem Nat Comput</source>. <year>2008</year>;<volume>4</volume>(<issue>3</issue>):<fpage>343</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1142/s1793005708001094</pub-id>; <pub-id pub-id-type="pmid">31116912</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weichert</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kister</surname> <given-names>A</given-names></string-name>, <string-name><surname>Volbach</surname> <given-names>P</given-names></string-name>, <string-name><surname>Houben</surname> <given-names>S</given-names></string-name>, <string-name><surname>Trost</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wrobel</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Explainable production planning under partial observability in high-precision manufacturing</article-title>. <source>J Manufact Syst</source>. <year>2023</year>;<volume>70</volume>:<fpage>514</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jmsy.2023.08.009</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>