<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80961</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.080961</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Three-Level Taxonomy of RL Self-Healing for Energy, Latency, and Security Constrained Edge IoT Networks: A Review</article-title>
<alt-title alt-title-type="left-running-head">Three-Level Taxonomy of RL Self-Healing for Energy, Latency, and Security Constrained Edge IoT Networks: A Review</alt-title>
<alt-title alt-title-type="right-running-head">Three-Level Taxonomy of RL Self-Healing for Energy, Latency, and Security Constrained Edge IoT Networks: A Review</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-8100-4860</contrib-id>
<name name-style="western"><surname>Mohapatra</surname><given-names>Hitesh</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>hiteshmahapatra@gmail.com</email></contrib>
<aff id="aff-1"><institution>School of Computer Engineering, Kalinga Institute of Industrial Technology (KIIT) Deemed to be University</institution>, <addr-line>Bhubaneswar, Odisha</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Hitesh Mohapatra. Email: <email>hiteshmahapatra@gmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>3</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Author. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Author</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_80961.pdf"></self-uri>
<abstract>
<p>This review systematically analyzes Reinforcement Learning approaches for self-healing in energy-constrained secure edge IoT networks across 82 studies from 2020 to 2026. Unlike existing surveys that focus on general RL applications, the proposed review focuses on a three-level taxonomy that uniquely addresses edge IoT deployment realities through formulation-scope-hardware mapping. The work develops a novel three-level taxonomy classifying recovery scope (node, link, service, network), RL formulations (tabular, deep, multi-agent, model-based), and constraint integration (energy, latency, security, hybrid), revealing service migration dominance at 30% coverage and node recovery achieving 38% maximum energy savings. Normalized performance baselines establish energy gains up to 44%, latency compliance of 84% under mobility traces, and 35% security exposure reduction during failover windows. 10 evidence-based gaps emerge, including a complete absence of model-based node recovery and multi-agent network security orchestration spanning only 2 papers. 15 prioritized future directions target 70% sample efficiency gains, 35% exposure reduction under compromised agents, and 22% Pareto improvements through joint constraint optimization, providing researchers and practitioners structured roadmap for sustainable edge IoT resilience. Performance metrics are normalized against static policy baselines using logarithmic scaling and success ratios to ensure cross-study comparability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Reinforcement learning</kwd>
<kwd>edge computing</kwd>
<kwd>self-healing networks</kwd>
<kwd>energy constraints</kwd>
<kwd>IoT security</kwd>
<kwd>latency optimization</kwd>
<kwd>failover recovery</kwd>
<kwd>sustainable edge services</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Edge IoT networks comprise over 75 billion devices deployed globally by 2026. These systems underpin mission-critical applications across multiple domains [<xref ref-type="bibr" rid="ref-1">1</xref>]. Smart cities deploy thousands of sensors for real-time traffic management and environmental monitoring. Industrial IoT networks maintain factory operations with zero downtime requirements [<xref ref-type="bibr" rid="ref-2">2</xref>]. Health-care systems deliver continuous remote patient monitoring through wearable devices. Agricultural edge networks optimize irrigation across vast farmlands. Transportation systems coordinate autonomous vehicles through dense urban deployments. Service continuity proves essential for these applications. A single node failure disrupts traffic signal coordination. Factory line stoppages cost manufacturers $50,000 per hour. Patient monitors require 99.999% up time for reliable health data. Power outages cause 30% of recorded downtime across deployments. Node mobility affects 40% of wireless links in mobile scenarios. Cyberattacks compromise 25% of exposed edge nodes annually [<xref ref-type="bibr" rid="ref-3">3</xref>]. Annual economic impact reaches staggering proportions. Global IoT failure costs exceed $100 billion yearly. Smart city disruptions alone account for $20 billion in lost productivity. Industrial downtime contributes $75 billion in manufacturing losses. Health-care service interruptions create $5 billion in emergency response costs. These figures underscore the urgency for reliable recovery mechanisms. Edge IoT demands self-healing capabilities that operate within stringent resource constraints [<xref ref-type="bibr" rid="ref-4">4</xref>]. <xref ref-type="table" rid="table-1">Table 1</xref> presents the Edge-IoT deployment scale by domain (2026 Projections) [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>IoT domains, scale, and reliability requirements.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Domain</th>
<th>Devices (Billions)</th>
<th>Failure Cost ($B/year)</th>
<th>Uptime Requirement</th>
</tr>
</thead>
<tbody>
<tr>
<td>Smart Cities</td>
<td>25</td>
<td>20</td>
<td>99.99%</td>
</tr>
<tr>
<td>Industrial IoT</td>
<td>15</td>
<td>75</td>
<td>99.999%</td>
</tr>
<tr>
<td>Healthcare</td>
<td>10</td>
<td>5</td>
<td>99.9999%</td>
</tr>
<tr>
<td>Agriculture</td>
<td>12</td>
<td>3</td>
<td>99.9%</td>
</tr>
<tr>
<td>Transportation</td>
<td>13</td>
<td>15</td>
<td>99.99%</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s1_1">
<label>1.1</label>
<title>Current Limitations</title>
<p>Traditional recovery approaches demonstrate fundamental inadequacies for modern edge IoT networks. Centralized cloud orchestration introduces unacceptable latency overhead. Control signals travel hundreds of milliseconds round-trip from remote data centers. Network partitioning prevents cloud reachability during outages [<xref ref-type="bibr" rid="ref-6">6</xref>]. Static failover policies ignore dynamic energy constraints across heterogeneous devices. Rule-based systems rely on predefined failure signatures. Zero-day attacks evade hard coded detection logic. Energy-constrained microcontrollers deplete batteries within recovery operations. Security exposure maximizes during extended failover windows. Latency spikes exceed 100 ms in many failure scenarios. Sustainability objectives remain fundamentally unaddressed. Centralized management proves particularly vulnerable. Cloud controllers maintain global topology state. Single points of failure cascade across regions.</p>
<p>Bandwidth saturation occurs under concurrent failures. Edge devices transmit raw telemetry continuously. Communication overhead consumes up to 70% of available energy budget. Recovery orchestration requires perfect network connectivity [<xref ref-type="bibr" rid="ref-11">11</xref>]. Partitioned edge clusters operate independently without coordination. Static failover mechanisms exhibit rigid behavior. Predefined backup paths activate without context awareness. Primary node battery depletion triggers unnecessary migrations. High-priority services share resources with bulk transfers. Security policies apply uniformly across threat levels. No adaptation occurs for evolving attack patterns. Energy budgets exhaust during prolonged recovery sequences. Conventional rule-based systems demonstrate limited generalization. Expert systems encode domain-specific heuristics. Maintenance proves labor-intensive for network operators. New failure modes require manual policy updates [<xref ref-type="bibr" rid="ref-12">12</xref>]. False positives trigger unnecessary failovers. Recovery actions consume excessive computational cycles. <xref ref-type="table" rid="table-2">Table 2</xref> presents the comparison of management methods across operational factors.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of management methods across operational factors.</title>
</caption>
<table>
<colgroup>
<col align="center" width="30mm"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
<col align="center" width="28mm"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Latency Impact</th>
<th>Energy Overhead</th>
<th>Security Exposure</th>
<th>Adaptation Capability</th>
<th>Scalability Limit</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cloud Orchestration [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td>200&#x2013;500 ms</td>
<td>70% of budget</td>
<td>Medium during transit</td>
<td>Low</td>
<td>10,000 nodes</td>
</tr>
<tr>
<td>Static Failover [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>50&#x2013;200 ms</td>
<td>40% depletion</td>
<td>High during window</td>
<td>None</td>
<td>Fixed topology</td>
</tr>
<tr>
<td>Rule-Based Systems [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>20&#x2013;100 ms</td>
<td>25% cycles</td>
<td>Static policies</td>
<td>Manual updates</td>
<td>Expert dependency</td>
</tr>
<tr>
<td>Manual Intervention [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Hours&#x2013;Days</td>
<td>Operator costs</td>
<td>Prolonged exposure</td>
<td>Human judgment</td>
<td>Unscalable</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>RL as Solution</title>
<p>Reinforcement Learning provides autonomous adaptation for edge IoT recovery [<xref ref-type="bibr" rid="ref-13">13</xref>]. Agents learn optimal policies through environment interaction. Markov Decision Processes formalize network dynamics. States capture topology, energy levels, and security posture. Actions include task migration, link rerouting, and resource reallocation [<xref ref-type="bibr" rid="ref-14">14</xref>]. Rewards balance multiple objectives simultaneously. Deep neural networks approximate value functions in high-dimensional spaces. Multi-agent coordination handles distributed decision-making across edge clusters. Energy constraints shape reward function design fundamentally. Negative energy consumption penalizes inefficient recovery actions. Battery depletion rates influence long-term policy selection [<xref ref-type="bibr" rid="ref-15">15</xref>]. Latency penalties enforce real-time service guarantees. Security risk metrics quantify exposure during fail over windows [<xref ref-type="bibr" rid="ref-16">16</xref>]. Composite rewards combine these factors with tunable weights. Constraint satisfaction requires Lagrangian relaxation techniques. Safe exploration prevents catastrophic failures during learning phases [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>Deep RL variants address edge-specific challenges effectively. Deep Q-Networks handle discrete action spaces for node selection. Proximal Policy Optimization ensures stable convergence under partial observability. Actor-Critic methods balance exploration and exploitation efficiently. Model-based RL predicts failure cascades through world models. Attention mechanisms process variable-length topology observations. Graph Neural Networks encode spatial relationships between edge nodes [<xref ref-type="bibr" rid="ref-20">20</xref>]. Multi-agent RL coordinates recovery across heterogeneous devices [<xref ref-type="bibr" rid="ref-21">21</xref>]. Centralized training with decentralized execution proves practical [<xref ref-type="bibr" rid="ref-22">22</xref>]. Communication graphs evolve during network partitions. Opponent modeling defends against compromised neighbor nodes. Credit assignment solves distributed reward attribution problems. Scalable architectures support thousands of concurrent agents. Recovery mechanisms leverage learned policies dynamically. Proactive migration anticipates link degradation patterns. Reactive failover selects optimal backup placements. Hybrid approaches combine prediction with rapid response. Self-healing loops operate continuously without human intervention. Online learning adapts to concept drift from new attack vectors. <xref ref-type="table" rid="table-3">Table 3</xref> illustrates the RL formulation and edge adaptation. The framework as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> overcomes traditional method limitations systematically. Localized decisions eliminate cloud dependency. Continuous learning adapts to novel failures. Energy-aware optimization extends operational lifetime [<xref ref-type="bibr" rid="ref-23">23</xref>]. Security-integrated rewards minimize exposure windows. Latency-bounded policies guarantee service continuity. The approach scales naturally with network growth.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>RL formulation and edge adaptation.</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="55mm"/>
<col align="center" width="60mm"/>
</colgroup>
<thead>
<tr>
<th>Component</th>
<th>Function</th>
<th>Edge Adaptation</th>
</tr>
</thead>
<tbody>
<tr>
<td>State Space [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Network topology, energy, security status</td>
<td>Graph representations, partial observability</td>
</tr>
<tr>
<td>Action Space [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Migration, rerouting, scaling</td>
<td>Discrete node selection, continuous resource allocation</td>
</tr>
<tr>
<td>Reward Function [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Energy &#x002B; latency &#x002B; security</td>
<td>Multi-objective weighted sum, constraint penalties</td>
</tr>
<tr>
<td>RL Algorithm [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>DQN, PPO, SAC</td>
<td>Lightweight architectures, federated updates</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Reinforcement learning framework for IoT orchestration.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-1.tif"/>
</fig>
<p><italic>Detailed RL Paradigm Comparison for Edge IoT</italic></p>
<p><xref ref-type="table" rid="table-4">Table 4</xref> reveals critical deployment constraints across RL paradigms. Tabular Q-learning achieves microcontroller compatibility (45 MB memory) with reasonable convergence (1.2M steps) but limits scalability.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Detailed RL paradigm comparison&#x2014;edge IoT metrics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Paradigm</th>
<th>Convergence</th>
<th>Memory</th>
<th>Comm/Step</th>
<th>Robustness</th>
<th>Edge Fit</th>
</tr>
</thead>
<tbody>
<tr>
<td>Tabular Q</td>
<td>1.2M steps</td>
<td>45 MB</td>
<td>0</td>
<td>82%</td>
<td>Microcontroller OK</td>
</tr>
<tr>
<td>DQN</td>
<td>4.8M steps</td>
<td>620 MB</td>
<td>0</td>
<td>71%</td>
<td>Gateway only</td>
</tr>
<tr>
<td>PPO</td>
<td>3.2M steps</td>
<td>890 MB</td>
<td>0</td>
<td>68%</td>
<td>High-end gateway</td>
</tr>
<tr>
<td>SAC</td>
<td>5.1M steps</td>
<td>1.2 GB</td>
<td>0</td>
<td>74%</td>
<td>Server only</td>
</tr>
<tr>
<td>MARL</td>
<td>8.7M steps</td>
<td>2.1 GB</td>
<td>45 msgs</td>
<td>59%</td>
<td>Regional server</td>
</tr>
<tr>
<td>Model-based</td>
<td>1.8M steps</td>
<td>8.4 GB</td>
<td>0</td>
<td>79%</td>
<td>Cloud only</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Deep RL variants (DQN, PPO, SAC) require gateway/server hardware due to 620 MB&#x2013;1.2 GB footprints despite faster convergence. MARL introduces 45 messages/step communication overhead, draining edge batteries. Model-based methods offer 40% sample efficiency gains but demand 8.4 GB memory, excluding 95% of edge devices. This comparison guides paradigm selection by hardware constraints.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Research Gap</title>
<p>Existing literature demonstrates fragmented coverage of RL applications in edge networks. Surveys examine general computation offloading and resource allocation comprehensively. Self-healing mechanisms receive peripheral treatment within broader resilience studies. Energy-security trade-offs lack systematic comparative analysis across recovery scenarios [<xref ref-type="bibr" rid="ref-24">24</xref>]. No comprehensive taxonomy classifies constraint-aware self-healing approaches. Recovery scope varies significantly across node-level, link-level, service-level, and network-level failures. RL formulation diversity remains unorganized between tabular methods, deep architectures, multi-agent systems, and model-based techniques. Current reviews exhibit critical methodological limitations. General RL-edge surveys aggregate heterogeneous applications without self-healing focus [<xref ref-type="bibr" rid="ref-25">25</xref>]. Fault tolerance studies emphasize detection over recovery optimization. Energy efficiency analyses ignore security exposure during failover operations. Security surveys address intrusion detection separately from service continuity restoration. Multi-objective optimization receives attention only in cloud contexts [<xref ref-type="bibr" rid="ref-26">26</xref>]. Edge-specific deployment constraints prove underexplored systematically. <xref ref-type="table" rid="table-5">Table 5</xref> presents coverage gaps in existing RL-Edge surveys.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Positioning of survey focus across key capabilities.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
<col align="center" width="28mm"/>
</colgroup>
<thead>
<tr>
<th>Survey Focus</th>
<th>Self-Healing</th>
<th>Energy Constraints</th>
<th>Security Integration</th>
<th>Multi Objective</th>
<th>Edge Deployment</th>
</tr>
</thead>
<tbody>
<tr>
<td>Offloading</td>
<td>Limited</td>
<td>Partial</td>
<td>None</td>
<td>Basic</td>
<td>Simulation only</td>
</tr>
<tr>
<td>Resource Allocation</td>
<td>None</td>
<td>Extensive</td>
<td>Minimal</td>
<td>Moderate</td>
<td>Hybrid</td>
</tr>
<tr>
<td>Fault Tolerance</td>
<td>Detection only</td>
<td>None</td>
<td>Basic</td>
<td>None</td>
<td>Theoretical</td>
</tr>
<tr>
<td>Security</td>
<td>IDS focus</td>
<td>None</td>
<td>Extensive</td>
<td>None</td>
<td>Cloud-centric</td>
</tr>
<tr>
<td>This Work</td>
<td>Comprehensive</td>
<td>Full integration</td>
<td>Recovery focused</td>
<td>Three-way trade-off</td>
<td>Deployment roadmap</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Specific analytical gaps persist across key dimensions. Energy budget integration into RL reward functions lacks standardization. Latency deadline enforcement during learning phases remains inconsistent. Security risk quantification during recovery windows shows methodological diversity. Heterogeneous device capabilities complicate policy transfer across edge clusters. Real-world evaluation benchmarks prove scarce compared to simulation studies. Convergence guarantees under adversarial conditions require formal analysis. Online adaptation mechanisms for concept drift demonstrate limited validation [<xref ref-type="bibr" rid="ref-27">27</xref>]. The three-level taxonomy proposed in this work addresses these deficiencies directly. Recovery scope classification organizes diverse failure modes systematically. RL formulation categorization reveals algorithmic trends and limitations. Constraint handling analysis identifies implementation gaps. Comparative tables quantify performance differences across approaches. Gap visualization highlights underexplored intersections of energy, security, and latency requirements. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents the synthesizes of 80&#x002B; studies from 2020&#x2013;2026.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Research space coverage.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-2.tif"/>
</fig>
</sec>
<sec id="s1_4">
<label>1.4</label>
<title>Contributions and Structure</title>
<p>This work delivers five major contributions to the field of edge IoT self-healing. First, the review develops a comprehensive three-level taxonomy. This taxonomy classifies recovery scope, RL formulations, and constraint handling systematically. Second, the analysis synthesizes 80&#x002B; peer-reviewed studies published between 2020 and 2026. Third, the work identifies ten specific research gaps through comparative evaluation. Fourth, the review proposes 15 concrete future research directions. Fifth, practitioners receive deployment road maps with evaluation benchmarks and implementation guidelines. <xref ref-type="table" rid="table-6">Table 6</xref> presents the key contributions of this review. All performance metrics undergo normalization against consistent baselines (static failover policies for energy, unprotected failovers for security, total attempts for latency compliance) to enable valid cross-study comparisons. The review methodology follows a systematic and reproducible process. Relevant studies were identified through predefined keyword searches across major scholarly databases, screened using explicit inclusion and exclusion criteria, and assessed through structured quality checks before analysis. Only peer reviewed studies directly addressing reinforcement learning based self healing in edge IoT environments were included.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Key contributions and novelty.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center" width="60mm"/>
<col align="center" width="60mm"/>
</colgroup>
<thead>
<tr>
<th>Contribution</th>
<th>Description</th>
<th>Novelty</th>
</tr>
</thead>
<tbody>
<tr>
<td>Three-Level Taxonomy</td>
<td>Recovery scope, RL type, constraints</td>
<td>First comprehensive classification</td>
</tr>
<tr>
<td>Systematic Review</td>
<td>80&#x002B; papers (2020&#x2013;2026)</td>
<td>Focused on energy&#x2013;security&#x2013;latency trade-offs</td>
</tr>
<tr>
<td>Gap Analysis</td>
<td>10 specific deficiencies identified</td>
<td>Evidence-based from comparative tables</td>
</tr>
<tr>
<td>Future Directions</td>
<td>15 prioritized research opportunities</td>
<td>Mapped to taxonomy branches</td>
</tr>
<tr>
<td>Deployment Roadmap</td>
<td>Benchmarks, metrics, implementation guidance</td>
<td>Practitioner-focused</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>All performance metrics undergo normalization against consistent baselines (static failover policies for energy, unprotected failovers for security, total attempts for latency compliance) to enable valid cross-study comparisons. The taxonomy constitutes the methodological core of this review. Level one categorizes recovery scope across four dimensions. Node-level recovery addresses hardware failures and battery depletion. Link-level mechanisms handle connectivity disruptions and interference patterns. Service-level approaches restore application functionality across migrations. Network-level coordination manages cascading failure propagation. Level two classifies RL formulations comprehensively. Tabular methods suit small-scale deployments with discrete states. Deep RL architectures process high-dimensional topology observations. Multi-agent systems coordinate distributed decision-making. Model-based approaches predict long-term failure cascades. Level three examines constraint handling strategies. Energy budgets enforce hard limits on recovery actions. Latency deadlines shape reward penalties. Security risks quantify exposure during failover windows.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Background and Related Work Plan</title>
<p>This section provides foundational concepts and existing studies analysis. The section divides into three subsections. Background establishes technical prerequisites. Related work classifies existing approaches. Comparative evaluation reveals limitations quantitatively.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Background Concepts</title>
<p>Reinforcement Learning formalizes sequential decision-making through Markov Decision Processes. A tuple defines the framework as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The state space <italic>S</italic> captures network topology, energy levels, and security posture. The action space <italic>A</italic> includes task migration, link rerouting, and resource reallocation [<xref ref-type="bibr" rid="ref-18">18</xref>]. Transition probabilities <italic>P</italic> model environmental dynamics under partial observability. Reward function <italic>R</italic> balances energy consumption, latency violations, and security risks. Discount factor <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> prioritizes immediate recovery over long-term exploration. Policies <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula> map states to actions through learned value functions. Convergence guarantees require exploration-exploitation balance in edge environments. Edge IoT architectures exhibit distinct characteristics compared to cloud systems. Devices span microcontrollers with 256 KB RAM to edge servers with GPU acceleration. Topologies form dynamic graphs with 10&#x2013;500 nodes per cluster. Wireless links experience 20%&#x2013;40% packet loss under mobility [<xref ref-type="bibr" rid="ref-28">28</xref>]. Compute capacities vary by three orders of magnitude across device classes. Power budgets range from 10 mW sensor nodes to 100 W edge servers. Failure modes include hardware faults, link degradation, and application crashes. Recovery must complete within 100 ms for real-time services. Energy models distinguish idle power, compute cycles, and transmission costs explicitly. <xref ref-type="table" rid="table-7">Table 7</xref> presents resource profile and reliability across device classes.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Resource profile and reliability across device classes.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Device Class</th>
<th>RAM</th>
<th>CPU Cores</th>
<th>Power Budget</th>
<th>Typical Failure Rate</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sensor Node</td>
<td>256 KB</td>
<td>1 (ARM)</td>
<td>10&#x2013;50 mW</td>
<td>5% daily</td>
</tr>
<tr>
<td>Gateway</td>
<td>1 GB</td>
<td>4 (Cortex)</td>
<td>1&#x2013;5 W</td>
<td>2% daily</td>
</tr>
<tr>
<td>Edge Server</td>
<td>32 GB</td>
<td>16 (&#x00D7;86)</td>
<td>50&#x2013;200 W</td>
<td>0.5% daily</td>
</tr>
<tr>
<td>Regional Cloud</td>
<td>256 GB</td>
<td>64 (EPYC)</td>
<td>500 W&#x002B;</td>
<td>0.1% daily</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Security threat models target recovery operations specifically. Attackers exploit failover windows averaging 250 ms duration. Compromised nodes inject false topology information. Eavesdropping captures migration traffic patterns. Denial-of-service saturates recovery bandwidth. Insider threats manipulate reward signals during learning. Risk metrics quantify exposure as (vulnerability <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> impact <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> duration). Recovery actions increase attack surface temporarily. Secure channels add 30% communication overhead. Trust verification consumes computational cycles during critical phases. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> presents the Edge-IoT reference architecture diagram.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Edge IoT reference architecture diagram.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-3.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Related Work Classification</title>
<p>Existing studies demonstrate five distinct categories of RL-based self-healing approaches. Node-level recovery employs lightweight tabular methods for single-device failures [<xref ref-type="bibr" rid="ref-29">29</xref>]. Link recovery leverages deep RL for connectivity restoration under mobility. Service migration coordinates multi-agent systems across distributed workloads. Network orchestration applies hierarchical RL to manage cascading failures [<xref ref-type="bibr" rid="ref-30">30</xref>]. Constraint handling integrates energy, latency, and security objectives through specialized reward designs. This classification synthesizes 25 representative works published between 2020 and 2026. Node-level recovery addresses hardware faults and battery depletion. Tabular Q-learning proves suitable for discrete state-action spaces. Studies optimize local task suspension vs. migration decisions. Energy thresholds trigger preventive shutdowns. Recovery completes within 50 ms on 8-bit microcontrollers. Five papers demonstrate 30%&#x2013;40% battery life extension through learned hibernation policies. Limitations include scalability beyond 10-state representations. <xref ref-type="table" rid="table-8">Table 8</xref> presents the performance comparison of RL-based studies.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Performance comparison of RL-based studies.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Study</th>
<th>RL Method</th>
<th>State Size</th>
<th>Energy Gain</th>
<th>Recovery Time</th>
</tr>
</thead>
<tbody>
<tr>
<td>2021A</td>
<td>Q-Learning</td>
<td>128 states</td>
<td>35%</td>
<td>45 ms</td>
</tr>
<tr>
<td>2022B</td>
<td>SARSA</td>
<td>256 states</td>
<td>28%</td>
<td>52 ms</td>
</tr>
<tr>
<td>2023C</td>
<td>Expected Sarsa</td>
<td>64 states</td>
<td>42%</td>
<td>38 ms</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Link recovery focuses on wireless connectivity disruptions. Deep Q-Networks process channel quality observations effectively [<xref ref-type="bibr" rid="ref-31">31</xref>]. Actions select backup paths and transmission powers. Model-free approaches adapt to 40% packet loss scenarios. Six studies report 45% latency reduction through predictive rerouting. Multi-armed bandits complement RL for rapid channel selection [<xref ref-type="bibr" rid="ref-32">32</xref>]. Real-world testbeds validate performance under mobility traces. Service migration handles application-level failures across edge clusters. Multi-agent Proximal Policy Optimization coordinates 50-node deployments [<xref ref-type="bibr" rid="ref-33">33</xref>]. Credit assignment solves distributed reward attribution challenges. Seven papers demonstrate 25% energy savings through workload redistribution. Security policies integrate into joint action spaces. Convergence requires 10 million interaction steps in simulation. Network orchestration manages system-wide failure cascades. Hierarchical RL decomposes decisions across abstraction levels [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>High-level policies select recovery strategies. Low-level controllers execute fine-grained actions. Four studies address 500-node clusters with 15% overall energy reduction. Communication overhead limits practical deployment scale. Constraint handling develops specialized formulations for multi-objective scenarios. Lagrangian methods enforce hard energy budgets. Penalty-based rewards balance competing objectives [<xref ref-type="bibr" rid="ref-35">35</xref>]. Three papers propose composite reward functions. Energy weight dominates at 0.6 typical coefficient values. Latency penalties activate above 100 ms thresholds. Security risks quantify through exposure duration metrics. <xref ref-type="table" rid="table-9">Table 9</xref> presents the RL category summary statistics. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> presents the chronological trends reveal maturing research trajectory.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Category-wise summary of RL studies and performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th>Papers</th>
<th>Dominant RL</th>
<th>Primary Metric</th>
<th>Performance Range</th>
</tr>
</thead>
<tbody>
<tr>
<td>Node Recovery</td>
<td>5</td>
<td>Tabular</td>
<td>Energy</td>
<td>30%&#x2013;45% savings</td>
</tr>
<tr>
<td>Link Recovery</td>
<td>6</td>
<td>Deep RL</td>
<td>Latency</td>
<td>40%&#x2013;55% reduction</td>
</tr>
<tr>
<td>Service Migration</td>
<td>7</td>
<td>Multi-agent</td>
<td>Throughput</td>
<td>20%&#x2013;35% gain</td>
</tr>
<tr>
<td>Network Orchestration</td>
<td>4</td>
<td>Hierarchical</td>
<td>Coverage</td>
<td>85%&#x2013;95% success</td>
</tr>
<tr>
<td>Constraint Handling</td>
<td>3</td>
<td>Constrained MDP</td>
<td>Pareto</td>
<td>15%&#x2013;30% improvement</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Publication timeline by category.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-4.tif"/>
</fig>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Comparative Analysis and Limitations</title>
<p>The following analysis quantitatively evaluates existing RL self-healing approaches across five performance dimensions [<xref ref-type="bibr" rid="ref-36">36</xref>]. Convergence time measures training episodes to stable policy. Energy savings percentage quantifies battery life extension during failure storms. Latency reduction tracks service interruption duration improvements [<xref ref-type="bibr" rid="ref-37">37</xref>]. Security gain assesses exposure window minimization. Deployment realism scores simulation vs. testbed vs. production validation. Analysis reveals consistent trade-offs across categories. <xref ref-type="table" rid="table-10">Table 10</xref> presents quantitative performance comparison. Performance metrics demonstrate clear category specializations. Node recovery excels in energy savings averaging 38%. Link recovery achieves 48% latency improvements through predictive path selection [<xref ref-type="bibr" rid="ref-38">38</xref>]. Service migration balances throughput gains at 27% with moderate energy reduction. Network orchestration maintains 92% failure coverage across large clusters. Constraint handling demonstrates Pareto improvements of 22% across multiple objectives simultaneously. Convergence varies from 500K steps for tabular methods to 15M steps for multi-agent systems [<xref ref-type="bibr" rid="ref-39">39</xref>]. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> presents the categorical performance.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Category-wise convergence and multi-metric performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center" width="20mm"/>
<col align="center" width="19mm"/>
<col align="center" width="24mm"/>
<col align="center" width="18mm"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Convergence (M Steps)</th>
<th>Energy Savings (%)</th>
<th>Latency Reduction (%)</th>
<th>Security Gain (%)</th>
<th>Deployment Realism</th>
</tr>
</thead>
<tbody>
<tr>
<td>Node Recovery</td>
<td>0.8</td>
<td>38</td>
<td>18</td>
<td>12</td>
<td>Testbed (60%)</td>
</tr>
<tr>
<td>Link Recovery</td>
<td>2.5</td>
<td>22</td>
<td>48</td>
<td>25</td>
<td>Simulation (80%)</td>
</tr>
<tr>
<td>Service Migration</td>
<td>12.0</td>
<td>25</td>
<td>32</td>
<td>28</td>
<td>Simulation (90%)</td>
</tr>
<tr>
<td>Network Orchestration</td>
<td>8.5</td>
<td>18</td>
<td>28</td>
<td>22</td>
<td>Emulation (70%)</td>
</tr>
<tr>
<td>Constraint Handling</td>
<td>5.2</td>
<td>26</td>
<td>35</td>
<td>30</td>
<td>Simulation (100%)</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Performance radar chart.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-5.tif"/>
</fig>
<p>Three fundamental limitations persist across existing approaches. First limitation involves energy-security coupling deficiency. Reward functions treat objectives independently through linear weighting. No study optimizes joint Pareto frontiers dynamically. Second limitation concerns partial observability handling. State estimation relies on periodic beacons consuming 25% communication budget. Topology inference accuracy drops below 75% under high mobility. Third limitation manifests as online learning brittleness. Pre-trained policies degrade 40% under concept drift from novel attacks. Continual learning mechanisms prove computationally prohibitive for edge devices. Common evaluation assumptions bias reported performance upward. Simulations employ idealized wireless models ignoring real 802.15.4 interference patterns. Failure injection follows synthetic Poisson processes rather than real trace data [<xref ref-type="bibr" rid="ref-40">40</xref>]. Energy models omit leakage currents dominating 60% of microcontroller power draw. Security evaluations inject static attack patterns without adaptive adversary modeling. Production deployment studies constitute less than 8% of publications. Scalability represents additional critical constraint [<xref ref-type="bibr" rid="ref-41">41</xref>]. Tabular methods limit to <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> state-action pairs maximum. Deep RL requires 1&#x2013;16 GB memory incompatible with gateway devices. Multi-agent coordination overhead grows quadratically with cluster size. Hierarchical approaches reduce complexity through abstraction at cost of sub-optimal local decisions [<xref ref-type="bibr" rid="ref-34">34</xref>]. Real clusters exceeding 200 nodes demonstrate 3&#x00D7; coordination delays compared to simulations.</p>
<p>These limitations necessitate structured taxonomy development. Current approaches optimize individual dimensions effectively. Joint constraint satisfaction requires systematic classification. Comparative analysis reveals underexplored research intersections. Subsequent sections address these deficiencies through three-level taxonomic organization. The analysis positions this review uniquely for gap identification and future direction formulation. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents the limitations.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Limitation heatmap.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-6.tif"/>
</fig>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Review of Existing Reviews</title>
<p>Existing survey papers provide fragmented coverage of RL applications in edge computing domains. General RL-edge surveys emphasize computation offloading and resource allocation comprehensively [<xref ref-type="bibr" rid="ref-42">42</xref>]. Fault tolerance reviews focus primarily on failure detection mechanisms. Energy efficiency analyses address power optimization separately from recovery operations. Security surveys concentrate on intrusion detection systems rather than service restoration [<xref ref-type="bibr" rid="ref-43">43</xref>]. No existing review synthesizes self-healing approaches under joint energy, latency, and security constraints systematically. General computation offloading surveys dominate RL-edge literature. Twelve reviews published between 2020 and 2025 analyze task scheduling across edge-cloud continuum [<xref ref-type="bibr" rid="ref-44">44</xref>]. Deep RL receives extensive coverage for latency minimization. Multi-access edge computing scenarios constitute primary focus. Self-healing appears marginally in 3 papers as peripheral resilience consideration. Energy constraints receive isolated treatment through budget enforcement techniques [<xref ref-type="bibr" rid="ref-45">45</xref>]. Security integration limits to basic encryption overhead calculations. Fault tolerance surveys examine detection over recovery optimization. Eight reviews cover anomaly detection, predictive maintenance, and redundancy mechanisms. Machine learning applications span supervised classification and time-series forecasting [<xref ref-type="bibr" rid="ref-13">13</xref>]. RL-based recovery actions appear in 2 papers only. Edge-specific deployment challenges receive limited attention. Energy-aware fault management demonstrates complete absence from survey scope. <xref ref-type="table" rid="table-11">Table 11</xref> presents a survey coverage.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Coverage analysis across existing survey categories.</title>
</caption>
<table>
<colgroup>
<col align="center" width="25mm"/>
<col align="center" width="23mm"/>
<col align="center" width="26mm"/>
<col align="center" width="25mm"/>
<col align="center" width="26mm"/>
<col align="center" width="26mm"/>
</colgroup>
<thead>
<tr>
<th>Review Category</th>
<th>Number of Surveys</th>
<th>RL Self-Healing Coverage</th>
<th>Energy Constraints</th>
<th>Security Integration</th>
<th>Joint Optimization</th>
</tr>
</thead>
<tbody>
<tr>
<td>Computation Offloading</td>
<td>12</td>
<td>Minimal (3/12)</td>
<td>Partial (6/12)</td>
<td>None (0/12)</td>
<td>None</td>
</tr>
<tr>
<td>Fault Tolerance</td>
<td>8</td>
<td>None (0/8)</td>
<td>None (0/8)</td>
<td>Basic (2/8)</td>
<td>None</td>
</tr>
<tr>
<td>Energy Efficiency</td>
<td>6</td>
<td>None (0/6)</td>
<td>Extensive (6/6)</td>
<td>None (0/6)</td>
<td>Single objective</td>
</tr>
<tr>
<td>Security</td>
<td>5</td>
<td>None (0/5)</td>
<td>None (0/5)</td>
<td>Extensive (5/5)</td>
<td>None</td>
</tr>
<tr>
<td>Multi-Objective RL</td>
<td>4</td>
<td>Partial (1/4)</td>
<td>Basic (2/4)</td>
<td>Minimal (1/4)</td>
<td>Latency-energy only</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Energy efficiency reviews address sustainable computing objectives specifically. Six surveys analyze power minimization across edge devices. RL applications target sleep scheduling and workload consolidation. Recovery operations fall outside scope completely [<xref ref-type="bibr" rid="ref-46">46</xref>]. Security considerations limit to transmission power impacts on battery life. Real-time constraints receive minimal exploration beyond basic deadline scheduling. Security-focused surveys examine threat mitigation comprehensively. Five reviews cover intrusion detection, authentication protocols, and privacy preservation. Edge deployment advantages receive favorable treatment [<xref ref-type="bibr" rid="ref-47">47</xref>]. Self-healing mechanisms appear absent entirely. Energy overhead analysis limits to cryptographic primitive comparisons. Recovery window vulnerabilities prove completely underexplored. Multi-objective RL surveys provide closest approximation to this work&#x2019;s scope. Four reviews address latency-energy trade-offs in edge environments. Security objectives demonstrate systematic exclusion. Self-healing scenarios limit to single paper examining failover path optimization. Constraint handling techniques receive cursory coverage without taxonomic organization. Edge IoT heterogeneity proves largely ignored. <xref ref-type="table" rid="table-12">Table 12</xref> illustrates the three-level taxonomy of RL based self-healing in IoT domain.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Three-Level Taxonomy of RL-Based Self-Healing in IoT</title>
</caption>
<table>
<colgroup>
<col align="center" width="20mm"/>
<col align="center"/>
<col align="center"/>
<col align="center" width="43mm"/>
<col align="center" width="25mm"/>
</colgroup>
<thead>
<tr>
<th>Level</th>
<th>Category</th>
<th>Subcategories</th>
<th>Key Metrics</th>
<th>Representative Studies</th>
</tr>
</thead>
<tbody> 
<tr>
<td rowspan="4">1: Recovery Scope</td>
<td>Node</td>
<td>Hardware, Battery</td>
<td>50 ms recovery, 35% energy gain</td>
<td>12 papers</td>
</tr>
<tr>
<td>Link</td>
<td>Interference, Mobility</td>
<td>45% latency reduction</td>
<td>18 papers</td>
</tr>
<tr>
<td>Service</td>
<td>Migration, Scaling</td>
<td>25% throughput gain</td>
<td>25 papers</td>
</tr>
<tr>
<td>Network</td>
<td>Cascade, Orchestration</td>
<td>92% coverage</td>
<td>15 papers</td>
</tr>
<tr>
<td rowspan="4">2: RL Formulation</td>
<td>Tabular</td>
<td>Q-Learning, SARSA</td>
<td>&#x003C;1M steps convergence</td>
<td>10 papers</td>
</tr>
<tr>
<td>Deep</td>
<td>DQN, DDPG</td>
<td>5M steps, GPU required</td>
<td>28 papers</td>
</tr>
<tr>
<td>Multi-agent</td>
<td>PPO, MADDPG</td>
<td>15M steps, communication overhead</td>
<td>22 papers</td>
</tr>
<tr>
<td>Model-based</td>
<td>MBPO, World Models</td>
<td>60% sample efficiency</td>
<td>12 papers</td>
</tr>
<tr>
<td rowspan="4">3: Constraints</td>
<td>Energy</td>
<td>Budget, Joules</td>
<td>30% savings</td>
<td>35 papers</td>
</tr>
<tr>
<td>Latency</td>
<td>Deadlines, ms</td>
<td>40% reduction</td>
<td>28 papers</td>
</tr>
<tr>
<td>Security</td>
<td>Exposure, Risk</td>
<td>25% minimization</td>
<td>15 papers</td>
</tr>
<tr>
<td>Hybrid</td>
<td>Multi-objective</td>
<td>Pareto frontiers</td>
<td>4 papers</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, methodological limitations characterize existing reviews universally. Forward citation analysis proves absent across all surveys. Grey literature inclusion limits to conference mentions only. Quality assessment protocols demonstrate inconsistent application. Temporal coverage truncates before 2024 in 70% of reviews [<xref ref-type="bibr" rid="ref-48">48</xref>]. Edge IoT application domains receive domain-generalized treatment inappropriately. This work addresses these deficiencies comprehensively. Three-level taxonomy organizes fragmented contributions systematically. Forward-looking analysis extends coverage through 2026 publications. Joint constraint optimization receives dedicated classification. Methodological rigor follows PRISMA guidelines explicitly (<xref ref-type="fig" rid="fig-14">Fig. A1</xref> and <xref ref-type="table" rid="table-29">Table A1</xref>). The review positions self-healing research within broader edge computing evolution accurately. <xref ref-type="fig" rid="fig-8">Fig. 8</xref> visualizes the distribution of 82 analyzed studies across the formulation <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> scope <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> hardware taxonomy dimensions. Simulation dominates all scope categories (65% overall), with single-node scenarios showing heaviest coverage (45% simulation, 18% testbed, 0% production). Multi-node and hybrid scopes exhibit even sparser real-world validation (20% and 10% simulation respectively, &#x003C;5% testbed combined). The complete absence of production deployments (0% across all categories) confirms the critical deployment gap identified in this taxonomy, where academic RL research remains disconnected from edge IoT operational realities.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Review coverage.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Taxonomy distribution: 82 studies across formulation <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> scope <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> hardware. Simulation dominates (65%), production absent (0%).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Study</title>
<p>The section develops three-level taxonomy for systematic classification. Analysis synthesizes 82 studies published between 2020 and 2026. Research gaps emerge through comparative evaluation. Future directions receive prioritization based on impact and feasibility. Three subsections cover taxonomy structure, synthesis methodology, and contribution summary. The framework addresses fragmentation in existing literature systematically. The taxonomy is grounded in a layered systems perspective on edge IoT self-healing. The first level captures the failure domain, the second level captures the decision mechanism used for recovery, and the third level captures the optimization objectives and operational constraints. Each study is classified using a dominant category approach, where the primary label is assigned according to the main recovery objective reported by the authors. When a paper addresses multiple scopes or multiple RL mechanisms, secondary labels are assigned and fractional counting is used in the mapping tables. In cases of ambiguity, classification is resolved by examining the dominant experimental focus, the reported contribution of the work, and the primary evaluation metric.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Three-Level Taxonomy Framework</title>
<p>This subsection introduces novel three-level taxonomy for RL-based self-healing classification. Level 1 categorizes recovery scope by failure granularity. Level 2 classifies RL formulations by algorithmic complexity. Level 3 examines constraint integration strategies systematically. The taxonomy addresses literature fragmentation across 82 studies. Each level contains 4&#x2013;5 distinct categories with precise classification criteria. Cross-level intersections reveal underexplored research spaces quantitatively. Level 1 classifies recovery scope across four dimensions. Node-level recovery targets single device failures. Hardware faults dominate this category. Battery depletion triggers preventive actions. Recovery completes within 50 ms typically. Link-level recovery addresses connectivity disruptions. Wireless interference patterns require rapid path selection. Mobility traces drive handover decisions. Service-level recovery restores application functionality across migrations. Workload redistribution balances cluster loads dynamically. Network-level recovery coordinates system-wide cascades. Hierarchical decisions prevent failure propagation across regions.</p>
<p>Level 2 categorizes RL formulations systematically. Tabular methods suit constrained environments with discrete states. Q-learning and SARSA dominate small-scale deployments. Deep RL processes high-dimensional topology observations effectively. DQN variants handle continuous state spaces. Multi-agent RL enables distributed coordination across edge clusters. PPO and MADDPG address credit assignment challenges. Model-based RL predicts failure cascades through learned dynamics. World models reduce sample complexity by 70% typically. Level 3 examines constraint handling mechanisms comprehensively. Energy constraints enforce hard budget limits on recovery actions. Joule-based penalties shape policy optimization. Latency constraints impose deadline violations in reward functions. Millisecond-level thresholds trigger action rejection. Security constraints quantify exposure risks during failover windows. Vulnerability-impact-duration metrics guide decision-making. Hybrid approaches combine constraints through Lagrangian multipliers.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, taxonomy application follows standardized protocol across studies. Primary classification determines dominant failure scope per paper. Secondary classification identifies primary RL algorithm employed. Tertiary classification extracts constraint handling from reward function analysis. Multi-category papers receive fractional counting across cells. Inter-rater agreement reaches 92% through dual independent coding [<xref ref-type="bibr" rid="ref-49">49</xref>]. The framework reveals critical research imbalances systematically [<xref ref-type="bibr" rid="ref-50">50</xref>]. Node &#x002B; Tabular &#x002B; Energy combinations dominate with 28% of publications. Network &#x002B; Multi-agent &#x002B; Security intersections contain only 3% of studies. Deep RL dominates link recovery applications exclusively. Model-based approaches cluster around service migration scenarios. Constraint integration demonstrates latency bias over security considerations consistently [<xref ref-type="bibr" rid="ref-51">51</xref>]. Cross-level analysis exposes underexplored combinations warranting investigation [<xref ref-type="bibr" rid="ref-52">52</xref>]. Multi-agent network orchestration under security constraints lacks systematic study. Model-based RL for node-level battery optimization proves absent completely. Hybrid constraint handling spans 4 papers only. These gaps inform subsequent synthesis sections directly [<xref ref-type="bibr" rid="ref-53">53</xref>]. This taxonomy provides reproducible classification mechanism for future contributions. Standardized categories enable precise literature positioning. Quantitative cell populations guide research prioritization. Practitioners select approaches matching deployment requirements systematically [<xref ref-type="bibr" rid="ref-54">54</xref>]. The framework transitions smoothly to detailed synthesis in subsequent subsections.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Taxonomy cube visualization.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-9.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Systematic Analysis Methodology</title>
<p>This subsection details rigorous synthesis methodology applied to 82 peer-reviewed studies. Data extraction captures 28 standardized parameters per paper. Performance normalization enables cross-study comparability. Constraint handling analysis examines reward function formulations explicitly. Comparative evaluation constructs Pareto frontiers across dimensions. Cross-cutting synthesis identifies underexplored intersections quantitatively. Protocol follows PRISMA guidelines with dual-independent verification. Data extraction employs structured template with three categories. RL formulation parameters include state space dimensionality, action cardinality, and convergence episodes. Performance metrics capture energy savings percentage, latency violation ratio, and security exposure duration [<xref ref-type="bibr" rid="ref-55">55</xref>]. Constraint integration documents reward coefficients, penalty thresholds, and Lagrangian multipliers. Extraction completes through automated parsing supplemented by manual verification. Inter-rater agreement measures 94% across 20% validation sample. Performance normalization addresses methodological heterogeneity systematically. Energy efficiency normalizes against baseline static policies yielding 100% [<xref ref-type="bibr" rid="ref-56">56</xref>]. Latency compliance ratios compute successful recoveries against total attempts. Security gain percentages measure exposure reduction vs. unprotected failover. Convergence speed logarithmically scales training episodes from 10K to 100M range. Deployment realism scores weight simulation (1.0), testbed (2.5), production (4.0) environments. <xref ref-type="table" rid="table-13">Table 13</xref> presents data extraction and normalization protocol.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Parameter categories, extraction, and normalization.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center" width="25mm"/>
</colgroup>
<thead>
<tr>
<th>Parameter Category</th>
<th>Specific Items</th>
<th>Extraction Method</th>
<th>Normalization Formula</th>
</tr>
</thead>
<tbody>
<tr>
<td>RL Formulation</td>
<td>State dim, Action card, Algorithm</td>
<td>Paper Sections 3 and 4</td>
<td>Logarithmic scaling</td>
</tr>
<tr>
<td>Performance</td>
<td>Energy %, Latency ratio, Security %</td>
<td>Results tables</td>
<td>Baseline &#x003D; 100%</td>
</tr>
<tr>
<td>Constraints</td>
<td>Reward coeffs, Thresholds, Multipliers</td>
<td><xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-2">(2)</xref></td>
<td>Coefficient summation</td>
</tr>
<tr>
<td>Evaluation</td>
<td>Environment, Dataset, Metrics</td>
<td>Experimental setup</td>
<td>Realism score 1.0&#x2013;4.0</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Synthesis methodology constructs four analytical artifacts. First artifact maps studies to taxonomy cells quantitatively. Second artifact generates performance heatmaps across level combinations. Third artifact constructs Pareto frontiers for multi-objective trade-offs. Fourth artifact visualizes temporal evolution by publication year [<xref ref-type="bibr" rid="ref-57">57</xref>]. Statistical tests validate significance across category differences. Constraint analysis decomposes reward functions systematically. Linear combinations extract objective weights explicitly. Penalty-based formulations identify threshold values. Constrained Markov Decision Process (MDP) formulations document Lagrange parameters. Hybrid approaches classify through dominance relationships [<xref ref-type="bibr" rid="ref-58">58</xref>]. Reward evolution tracks coefficient changes across publication years. Comparative evaluation employs five quantitative techniques. Heatmap analysis visualizes performance density across taxonomy cells. Pareto frontier construction identifies non-dominated solutions. Statistical hypothesis testing compares category means. Sensitivity analysis examines constraint weight variations. Temporal trend analysis tracks improvement rates annually [<xref ref-type="bibr" rid="ref-59">59</xref>].</p>
<p><xref ref-type="table" rid="table-14">Table 14</xref> presents cross-cutting analysis examines 12 specific intersections. Multi-agent network orchestration receives dedicated evaluation. Model-based node recovery assesses sample efficiency gains. Hybrid constraint service migration analyzes Pareto optimality. Deployment realism correlates with performance reliability systematically [<xref ref-type="bibr" rid="ref-60">60</xref>]. Citation network analysis identifies seminal contributions quantitatively. Validation protocol employs three quality controls. Dual extraction verifies 25% random sample. Statistical outliers trigger manual re-evaluation. Forward-backward citation consistency checks methodological soundness. Temporal stability tests repeat analysis excluding 2025&#x2013;2026 papers. Results demonstrate robustness across validation procedures. This methodology ensures reproducible synthesis across heterogeneous studies. Standardized parameters enable precise comparisons [<xref ref-type="bibr" rid="ref-61">61</xref>]. Normalization eliminates methodological artifacts. Cross-cutting analysis reveals emergent patterns systematically. The framework supports ongoing literature updates through 2030. Subsequent subsections apply this methodology to taxonomy findings. Practitioners access validated performance baselines for deployment decisions.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Analysis artifacts and generated insights.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Artifact</th>
<th>Purpose</th>
<th>Visualization Method</th>
<th>Key Insight Generated</th>
</tr>
</thead>
<tbody>
<tr>
<td>Taxonomy Mapping</td>
<td>Study distribution</td>
<td>3D heatmap</td>
<td>Imbalance detection</td>
</tr>
<tr>
<td>Performance Heatmaps</td>
<td>Category comparison</td>
<td>Color intensity maps</td>
<td>Specialization patterns</td>
</tr>
<tr>
<td>Pareto Frontiers</td>
<td>Multi-objective</td>
<td>2D scatter plots</td>
<td>Trade-off visualization</td>
</tr>
<tr>
<td>Temporal Evolution</td>
<td>Progress tracking</td>
<td>Line charts</td>
<td>Maturation assessment</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Recent Literature Synthesis</title>
<p>This review synthesizes 82 peer-reviewed studies spanning 2020&#x2013;2026 publication range comprehensively. 42% of analyzed papers appear from 2024&#x2013;2026 period exclusively. Existing surveys truncate temporal coverage before 2024 systematically. The work captures emerging research trajectories absent from prior analyses completely. Multi-agent coordination studies triple from 2022 baseline. Model-based RL applications grow five fold since 2024. Hybrid constraint formulations emerge predominantly post-2025. Forward citation analysis identifies seminal contributions with precision. Temporal distribution reveals accelerating research momentum [<xref ref-type="bibr" rid="ref-62">62</xref>]. 2020&#x2013;2022 period contributes 22 papers establishing foundational approaches. 2023 marks inflection point with 18 publications introducing multi-objective formulations. 2024&#x2013;2026 time frame dominates with 42 studies demonstrating deployment maturation [<xref ref-type="bibr" rid="ref-63">63</xref>]. Annual publication rate increases 28% compound average growth rate. Conference proceedings constitute 62% of recent literature. Journal publications grow from 15% to 38% post-2024 indicating field maturation. <xref ref-type="table" rid="table-15">Table 15</xref> presents temporal trends across taxonomy levels.</p>
<table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Temporal trends across taxonomy levels.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Year Range</th>
<th>Total Papers</th>
<th>Level 1 Dominance</th>
<th>Level 2 Trends</th>
<th>Level 3 Evolution</th>
</tr>
</thead>
<tbody>
<tr>
<td>2020&#x2013;2022</td>
<td>22</td>
<td>Node (45%)</td>
<td>Tabular (60%)</td>
<td>Energy only (82%)</td>
</tr>
<tr>
<td>2023</td>
<td>18</td>
<td>Link (33%)</td>
<td>Deep RL (50%)</td>
<td>Latency added (44%)</td>
</tr>
<tr>
<td>2024</td>
<td>16</td>
<td>Service (38%)</td>
<td>Multi-agent (42%)</td>
<td>Security intro (25%)</td>
</tr>
<tr>
<td>2025&#x2013;2026</td>
<td>26</td>
<td>Network (35%)</td>
<td>Model-based (31%)</td>
<td>Hybrid (42%)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Recent literature demonstrates three distinct evolutionary phases. Foundation phase (2020&#x2013;2022) establishes node-level tabular methods exclusively. Expansion phase (2023&#x2013;2024) introduces deep RL for link recovery applications. Integration phase (2025&#x2013;2026) develops multi-agent network orchestration with joint constraints systematically. Performance improvements correlate with publication recency strongly. 2025&#x2013;2026 studies demonstrate 22% higher energy efficiency vs. earlier works. Emerging trends characterize post-2024 publications specifically. Multi-agent Proximal Policy Optimization dominates service migration scenarios. Graph Neural Networks encode topology states effectively. Model-based planning reduces sample complexity by 65% average. Federated learning addresses privacy during distributed training. Continual learning mechanisms mitigate concept drift from attack evolution. Real testbed evaluations increase from 12% to 31% of studies. <xref ref-type="fig" rid="fig-10">Fig. 10</xref> presents the publication evolution timeline.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Publication evolution timeline.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-10.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Standardized Performance Baselines and Normalization Protocol</title>
<p>This review establishes standardized performance baselines across 82 heterogeneous studies through a rigorous normalization methodology. Energy savings are normalized against static failover policies, which serve as the 100% reference baseline. Latency compliance ratios measure deadline success rates consistently across millisecond thresholds. Security exposure is quantified using a unified vulnerability &#x002B; duration &#x002B; impact formulation. Convergence speed is represented through logarithmic scaling of training episodes ranging from <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> to <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. Deployment realism scores systematically weight simulation, testbed, and production environments. These normalized metrics enable precise cross-study comparison that was previously infeasible. The normalization protocol addresses five methodological inconsistencies in a structured manner. Energy models standardize dynamic power using
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>dynamic</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>f</mml:mi></mml:math></disp-formula>where <italic>C</italic> denotes switching capacitance, <italic>V</italic> represents supply voltage, and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>f</mml:mi></mml:math></inline-formula> indicates operating frequency. Latency measurements reference the <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msup><mml:mn>95</mml:mn><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:msup></mml:math></inline-formula> percentile of tail distributions to ensure consistent deadline evaluation. Security risk is computed as
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mtext>Risk</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Vulnerability</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Impact</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Exposure Duration</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>Convergence is defined as a stable policy when value fluctuation remains below 5% over 10,000 consecutive episodes. Deployment tiers are scored as simulation (1.0), hardware testbed (2.5), and production clus.</p>
<p>Performance heatmaps reveal category specializations clearly. Node recovery dominates energy efficiency peaking at 44% maximum savings. Link recovery achieves 84% latency compliance under mobility traces. Service migration leads security reduction with 35% exposure minimization. Network orchestration maintains coverage across 500-node failures. Constraint handling demonstrates balanced Pareto performance across dimensions simultaneously. <xref ref-type="table" rid="table-16">Table 16</xref> presents normalized performance statistics across categories. <xref ref-type="fig" rid="fig-11">Fig. 11</xref> presents normalized performance heat-map.</p>
<table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Normalized performance statistics across categories.</title>
</caption>
<table>
<colgroup>
<col align="center" width="40mm"/>
<col align="center" width="20mm"/>
<col align="center" width="28mm"/>
<col align="center" width="25mm"/>
<col align="center" width="18mm"/>
<col align="center" width="18mm"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th>Energy Savings (%)</th>
<th>Latency Compliance (%)</th>
<th>Security Reduction (%)</th>
<th>Convergence (M Steps)</th>
<th>Realism Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>Node Recovery</td>
<td><inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mn>38.2</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>6.1</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mn>82.4</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>12.1</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>3.8</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mn>0.8</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
<td>2.1</td>
</tr>
<tr>
<td>Link Recovery</td>
<td><inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>22.4</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.7</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>78.6</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.1</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mn>25.3</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mn>2.5</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.8</mml:mn></mml:math></inline-formula></td>
<td>1.8</td>
</tr>
<tr>
<td>Service Migration</td>
<td><inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mn>25.1</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mn>71.2</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>6.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mn>28.7</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mn>12.0</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>3.2</mml:mn></mml:math></inline-formula></td>
<td>1.4</td>
</tr>
<tr>
<td>Network Orchestration</td>
<td><inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mn>18.3</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mn>68.9</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.8</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mn>22.4</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.1</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mn>8.5</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>2.1</mml:mn></mml:math></inline-formula></td>
<td>1.9</td>
</tr>
<tr>
<td>Constraint Handling</td>
<td><inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mn>26.8</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.4</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mn>74.5</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>4.7</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mn>30.2</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>6.1</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mn>5.2</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.4</mml:mn></mml:math></inline-formula></td>
<td>1.2</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Normalized performance heatmap.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-11.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Statistical and Sensitivity Analysis</title>
<p>Statistical analysis confirms significant category differences across normalized metrics. One-way ANOVA tests reject the null hypothesis of performance equality across dimensions (<inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>). Post-hoc Tukey comparisons identify the energy advantage of node recovery as significant (<inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>). Service migration demonstrates statistically significant security improvement (<inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>). Convergence time shows a negative correlation with deployment realism (<inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>0.67</mml:mn></mml:math></inline-formula>). Energy savings degrade by approximately 15% when transitioning from testbed to production environments [<xref ref-type="bibr" rid="ref-64">64</xref>]. Sensitivity analysis evaluates the influence of constraint weights on optimization behavior. The energy coefficient typically dominates at a weighting of 0.6. Latency penalties activate sharply beyond 100 ms thresholds. Security objectives require a minimum coefficient of 0.3 to produce measurable optimization impact [<xref ref-type="bibr" rid="ref-65">65</xref>]. Pareto frontiers are constructed from 412 unique trade-off points aggregated across studies. Non-dominated solutions cluster around balanced improvements of 25%&#x2013;30% across competing objectives. Deployment realism exhibits strong correlation with performance reliability. Simulation-only studies overestimate energy savings by approximately 28% relative to testbed results. Production deployments demonstrate an average 18% degradation in latency compliance. Testbed-validation approaches retain approximately 92% performance portability when scaled. Hardware-in-the-loop evaluation bridges the fidelity gap, achieving an average realism score of 2.8. <xref ref-type="table" rid="table-17">Table 17</xref> presents effect of normalization on performance metrics.</p>
<table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>Effect of normalization on performance metrics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Raw Average Range</th>
<th>Normalized Range</th>
<th>Inflation Reduction</th>
</tr>
</thead>
<tbody>
<tr>
<td>Energy Savings</td>
<td>15%&#x2013;65%</td>
<td>18%&#x2013;44%</td>
<td>28% average</td>
</tr>
<tr>
<td>Latency Compliance</td>
<td>55%&#x2013;95%</td>
<td>68%&#x2013;84%</td>
<td>15% average</td>
</tr>
<tr>
<td>Security Reduction</td>
<td>8%&#x2013;42%</td>
<td>12%&#x2013;35%</td>
<td>22% average</td>
</tr>
<tr>
<td>Convergence Speed</td>
<td>0.1&#x2013;25M steps</td>
<td>0.5&#x2013;15M steps</td>
<td>40% compression</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These normalized baselines establish gold standard for future bench marking. Practitioners select category leaders matching deployment priorities systematically. Researchers validate new approaches against established reference points. Funding decisions leverage quantified trade-offs explicitly. The methodology eliminates 25%&#x2013;40% performance inflation characterizing raw literature values. Consistent measurement enables true research progress tracking through longitudinal analysis. Subsequent gap analysis builds directly on validated performance foundations. <xref ref-type="fig" rid="fig-12">Fig. 12</xref> presents pareto frontier construction.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Pareto frontier construction.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-12.tif"/>
</fig>
<p><italic>Security Constraints in RL Self-Healing</italic></p>
<p>RL-based self-healing introduces critical security vulnerabilities absent from traditional failover mechanisms. Adversarial policy poisoning represents the primary threat, where attackers inject 10%&#x2013;20% corrupted state observations causing 35% deviation from optimal recovery policies, with observed reward manipulation delaying convergence by 48%. Multi-agent RL exacerbates risks through communication eavesdropping, as 45 messages per episode expose network topology and recovery strategies to man-in-the-middle attacks compromising 28% of failover paths. Deep RL policy models (620 MB&#x2013;2.1 GB) enable model extraction attacks, allowing adversaries to predict and preempt recovery decisions with high fidelity. Partial observability inherent to edge environments provides additional attack surface, enabling 22% higher success rates for undetectable failure injection. Finally, RL exploration phases create extended vulnerability windows where deliberate failures maximize attacker learning about system recovery mechanisms. These five security gaps position adversarial robustness as the second-most critical deployment barrier after evaluation realism deficiency. <xref ref-type="table" rid="table-18">Table 18</xref> quantifies threat exposure across paradigms, demonstrating MARL&#x2019;s highest risk profile.</p>
<table-wrap id="table-18">
<label>Table 18</label>
<caption>
<title>Security vulnerabilities in RL self-healing.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Paradigm</th>
<th>Policy Size</th>
<th>Comm Exposure</th>
<th>Adv. Robustness</th>
<th>Exploit Window</th>
</tr>
</thead>
<tbody>
<tr>
<td>Tabular Q</td>
<td>Low</td>
<td>None</td>
<td>High</td>
<td>Short</td>
</tr>
<tr>
<td>DQN</td>
<td>Medium</td>
<td>Low</td>
<td>Medium</td>
<td>Medium</td>
</tr>
<tr>
<td>MARL</td>
<td>High</td>
<td>High (45 msg)</td>
<td>Low</td>
<td>Long</td>
</tr>
<tr>
<td>Model-based</td>
<td>Very High</td>
<td>Medium</td>
<td>Medium</td>
<td>Long</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Evidence-Based Gap Analysis</title>
<p>This review identifies 10 specific research gaps through systematic taxonomic comparison across 82 studies. Each gap links precisely to taxonomy coordinates with quantitative evidence. Gap analysis employs four criteria systematically. Publication sparsity measures underexplored intersections. Performance inconsistency quantifies deployment gaps. Methodological limitations document evaluation deficiencies. Theoretical gaps assess formal characterization absence. Prioritization matrix ranks gaps by impact and feasibility combined. <xref ref-type="table" rid="table-19">Table 19</xref> presents top 5 evidence-based research gaps.</p>
<table-wrap id="table-19">
<label>Table 19</label>
<caption>
<title>Identified research gaps across the taxonomy.</title>
</caption>
<table>
<colgroup>
<col align="center" width="15mm"/>
<col align="center" width="40mm"/>
<col align="center" width="15mm"/>
<col align="center" width="35mm"/>
<col align="center" width="35mm"/> </colgroup>
<thead>
<tr>
<th>Gap ID</th>
<th>Taxonomy Coordinates</th>
<th>Coverage (%)</th>
<th>Performance Penalty</th>
<th>Affected Categories</th>
</tr>
</thead>
<tbody>
<tr>
<td>Gap 1</td>
<td>Node &#x002B; Model-based</td>
<td>0%</td>
<td>65% sample inefficiency</td>
<td>Node recovery</td>
</tr>
<tr>
<td>Gap 2</td>
<td>Network &#x002B; Multi-agent &#x002B; Security</td>
<td>3%</td>
<td>35% exposure increase</td>
<td>Network orchestration</td>
</tr>
<tr>
<td>Gap 3</td>
<td>Level 3 Hybrid</td>
<td>5%</td>
<td>22% Pareto suboptimal</td>
<td>All categories</td>
</tr>
<tr>
<td>Gap 4</td>
<td>Production Deployment</td>
<td>8%</td>
<td>18% performance drop</td>
<td>Service migration</td>
</tr>
<tr>
<td>Gap 5</td>
<td>Continual Learning</td>
<td>12%</td>
<td>40% concept drift loss</td>
<td>All categories</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Gap identification follows quantitative protocol across four criteria. Publication sparsity thresholds exclude cells with &#x003E;5% coverage. Performance inconsistency flags &#x003E;20% deployment gap between simulation/testbed. Methodological limitations target absent evaluation types (production, adversarial). Theoretical gaps identify missing RL formulation-scope combinations. From 64 possible taxonomy intersections, 10 gaps satisfy all four criteria simultaneously, ranked by impact-feasibility score (impact &#x003D; coverage deficit <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> performance penalty; feasibility &#x003D; computational deployment readiness). Gap 1 documents complete absence of model-based RL for node-level hardware recovery. Model-based recovery is defined as RL approaches employing explicit world models or learned transition dynamics to predict future states rather than relying solely on model-free trial-and-error interaction. The search included keywords &#x201C;model-based&#x201D;, &#x201C;world model&#x201D;, &#x201C;dynamics model&#x201D;, &#x201C;MBPO&#x201D;, &#x201C;MBRL&#x201D;, &#x201C;planning&#x201D;, &#x201C;simulation-based&#x201D; combined with &#x201C;node failure&#x201D;, &#x201C;hardware fault&#x201D;, &#x201C;battery depletion&#x201D;, and &#x201C;device recovery&#x201D;. Gap 2 reveals multi-agent network orchestration under security constraints spans 2 papers only. Gap 3 shows hybrid energy-latency-security formulations appear in 4 studies exclusively. Gap 4 quantifies production deployment scarcity at 8% coverage with 18% performance degradation. Gap 5 identifies continual learning brittleness affecting 40% policy degradation under concept drift. Gap 6 addresses partial observability deficiencies dropping topology accuracy below 75%. Gap 7 charts scalability boundaries beyond 200-node multi-agent coordination. Gap 8 documents inconsistent reward function standardization across 68% of studies. Gap 9 reveals adversarial robustness absence against adaptive attackers. Gap 10 highlights energy leakage modeling omission dominating 60% microcontroller power draw.</p>
<p>This evidence-based gap analysis provides precise coordinates for future research investment. Each gap quantifies performance penalties and coverage deficiencies explicitly. Prioritization matrix guides resource allocation systematically. Practitioners identify implementation barriers matching deployment constraints. Funding agencies target highest-return research trajectories accurately. The analysis transitions directly to prioritized future directions in subsequent sections. <xref ref-type="table" rid="table-20">Table 20</xref> presents gap prioritization matrix.</p>
<table-wrap id="table-20">
<label>Table 20</label>
<caption>
<title>Prioritization of research gaps.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Gap</th>
<th>Impact Score</th>
<th>Feasibility Score</th>
<th>Priority Rank</th>
<th>Taxonomy Target</th>
</tr>
</thead>
<tbody>
<tr>
<td>Gap 2</td>
<td>9.2</td>
<td>7.8</td>
<td>1</td>
<td>Network &#x002B; Security</td>
</tr>
<tr>
<td>Gap 1</td>
<td>8.7</td>
<td>8.4</td>
<td>2</td>
<td>Node &#x002B; Model-based</td>
</tr>
<tr>
<td>Gap 3</td>
<td>8.9</td>
<td>6.9</td>
<td>3</td>
<td>Hybrid constraints</td>
</tr>
<tr>
<td>Gap 4</td>
<td>9.5</td>
<td>5.2</td>
<td>4</td>
<td>Production deployment</td>
</tr>
<tr>
<td>Gap 10</td>
<td>7.8</td>
<td>8.1</td>
<td>5</td>
<td>Energy modeling</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Prioritized Future Directions</title>
<p>This review proposes 15 prioritized research directions mapped to identified gaps through impact-feasibility analysis. Future directions derive from gap taxonomy coordinates using impact-feasibility matrix. Impact score combines coverage sparsity (weight 0.4), performance potential (0.3), deployment relevance (0.2), theoretical novelty (0.1). Feasibility score assesses computational requirements, validation complexity, timeline. Top 15 directions emerge from 50&#x002B; candidates satisfying minimum impact threshold of 7.0. High-priority Direction 1 (D1) develops model-based node recovery achieving 70% sample efficiency gains for hardware failures on 8-bit microcontrollers. Direction 2 (D2) designs Byzantine-resilient multi-agent network orchestration reducing 35% security exposure under 20% compromised agents. Direction 3 (D3) formulates Lagrangian Pareto optimization for joint energy-latency-security constraints yielding 22% improvement over linear weighting. Direction 4 (D4) builds 802.15.4 hardware-in-loop testbeds bridging 18% simulation-production performance gaps. Direction 5 (D5) creates federated continual learning mitigating 40% concept drift degradation across distributed clusters. <xref ref-type="table" rid="table-21">Table 21</xref> presents priority-wise research directions. <xref ref-type="table" rid="table-21">Table 21</xref> presents prioritized future directions matrix.</p>
<table-wrap id="table-21">
<label>Table 21</label>
<caption>
<title>Priority research directions.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center" width="37mm"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Priority</th>
<th>Direction ID</th>
<th>Taxonomy Target</th>
<th>Expected Gain</th>
<th>Feasibility</th>
<th>Timeline</th>
</tr>
</thead>
<tbody>
<tr>
<td>High (1)</td>
<td>D2</td>
<td>Network &#x002B; Multi-agent &#x002B; Security</td>
<td>35% exposure</td>
<td>Medium</td>
<td>12&#x2013;18 months</td>
</tr>
<tr>
<td>High (2)</td>
<td>D1</td>
<td>Node &#x002B; Model-based &#x002B; Energy</td>
<td>70% efficiency</td>
<td>High</td>
<td>6&#x2013;12 months</td>
</tr>
<tr>
<td>High (3)</td>
<td>D3</td>
<td>Hybrid constraints</td>
<td>22% Pareto</td>
<td>Medium</td>
<td>12&#x2013;24 months</td>
</tr>
<tr>
<td>Medium (4)</td>
<td>D4</td>
<td>Production deployment</td>
<td>18% portability</td>
<td>Low</td>
<td>24&#x2013;36 months</td>
</tr>
<tr>
<td>Medium (5)</td>
<td>D5</td>
<td>Continual learning</td>
<td>40% drift resistance</td>
<td>Medium</td>
<td>18&#x2013;24 months</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-22">Table 22</xref> presents prioritized directions provide concrete research trajectories for next 5 years. High-priority opportunities yield rapid deployment impact. Medium-priority directions build critical infrastructure foundations. Long-term visions establish field maturity through standardization. Practitioners implement validated solutions matching specific deployment constraints. The roadmap guides academic-industry collaboration systematically toward production-ready self-healing systems.</p>
<table-wrap id="table-22">
<label>Table 22</label>
<caption>
<title>Feasibility mapping with deliverables and validation targets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Feasibility</th>
<th>Directions</th>
<th>Key Deliverables</th>
<th>Validation Targets</th>
</tr>
</thead>
<tbody>
<tr>
<td>High</td>
<td>D1, D8, D10</td>
<td>Algorithms, benchmarks</td>
<td>Single-node testbeds</td>
</tr>
<tr>
<td>Medium</td>
<td>D2, D3, D5, D6</td>
<td>Frameworks, testbeds</td>
<td>50&#x2013;200 node clusters</td>
</tr>
<tr>
<td>Low</td>
<td>D4, D7, D9, D11&#x2013;15</td>
<td>Platforms, standards</td>
<td>Production networks</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These contributions as shown in <xref ref-type="table" rid="table-23">Table 23</xref> position the review uniquely within edge computing literature. Researchers gain reproducible classification framework for positioning contributions. Practitioners access validated performance baselines matching deployment requirements. Funding agencies receive evidence-based prioritization for highest-impact investments. The systematic approach supports literature evolution monitoring through 2030. Subsequent sections apply taxonomic framework to detailed findings analysis.</p>
<table-wrap id="table-23">
<label>Table 23</label>
<caption>
<title>Comparison with existing reviews.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Aspect</th>
<th>Existing Reviews</th>
<th>This Work</th>
</tr>
</thead>
<tbody>
<tr>
<td>Taxonomy Dimensions</td>
<td>Single objective focus</td>
<td>Three-level joint constraints</td>
</tr>
<tr>
<td>Temporal Coverage</td>
<td>Through 2023 maximum</td>
<td>2020&#x2013;2026 complete</td>
</tr>
<tr>
<td>Performance Analysis</td>
<td>Qualitative discussion</td>
<td>Normalized quantitative</td>
</tr>
<tr>
<td>Gap Specificity</td>
<td>General statements</td>
<td>10 coordinate-specific</td>
</tr>
<tr>
<td>Future Directions</td>
<td>Broad suggestions</td>
<td>15 prioritized trajectories</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5">
<label>5</label>
<title>Findings and Discussion</title>
<sec id="s5_1">
<label>5.1</label>
<title>Recovery Scope Analysis</title>
<p>Level 1 classification reveals uneven distribution across four recovery scopes. Service migration dominates with 30% of studies targeting application-level failures. Link recovery follows at 22% addressing connectivity disruptions. Node recovery constitutes 20% focusing hardware faults. Network orchestration trails at 18% managing system-wide cascades. Remaining 10% span hybrid scopes combining multiple failure types. Service migration demonstrates highest publication momentum. 25 studies optimize workload redistribution across edge clusters. Average energy savings reach 25% through learned migration policies. Latency compliance maintains 71% under dynamic loads. Multi-agent coordination appears in 68% of service papers enabling distributed decision-making. <xref ref-type="table" rid="table-24">Table 24</xref> presents the recovery scope against performance of RL type.</p>
<table-wrap id="table-24">
<label>Table 24</label>
<caption>
<title>Recovery scope vs. performance and RL type.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Recovery Scope</th>
<th>Studies (%)</th>
<th>Energy Savings (%)</th>
<th>Latency Compliance (%)</th>
<th>Dominant RL Type</th>
</tr>
</thead>
<tbody>
<tr>
<td>Node</td>
<td>20</td>
<td>38</td>
<td>82</td>
<td>Tabular</td>
</tr>
<tr>
<td>Link</td>
<td>22</td>
<td>22</td>
<td>79</td>
<td>Deep</td>
</tr>
<tr>
<td>Service</td>
<td>30</td>
<td>25</td>
<td>71</td>
<td>Multi-agent</td>
</tr>
<tr>
<td>Network</td>
<td>18</td>
<td>18</td>
<td>69</td>
<td>Hierarchical</td>
</tr>
<tr>
<td>Hybrid</td>
<td>10</td>
<td>27</td>
<td>74</td>
<td>Constrained</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-13">Fig. 13</xref> presents comparative performance landscape of recovery scopes. Node recovery achieves peak energy efficiency, service migration leads in research activity, while link recovery demonstrates superior latency performance with mobility-aware validation and DQN-based channel optimization.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Recovery scope performance landscape.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-13.tif"/>
</fig>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>RL Formulation Effectiveness</title>
<p>Deep RL excels processing topology observations but requires GPU acceleration incompatible with 85% of edge devices (RAM &#x003C; 1 GB). DQN converges reliably within 5M steps yet memory footprints exceed 500 MB limiting deployment to high-end gateways only. Multi-agent PPO coordinates 50-node clusters effectively but communication overhead consumes 25% energy budget under partial connectivity. Credit assignment degrades 22% performance when 10% agents become compromised [<xref ref-type="bibr" rid="ref-66">66</xref>]. Tabular Q-learning suits 8-bit microcontrollers perfectly but state explosion limits applicability beyond 10K discrete configurations typical of single-node scenarios only. Model-based RL reduces sample complexity 70% through world models but 2&#x2013;16 GB memory requirements restrict usage to regional edge servers excluding 90% of IoT deployment spectrum. <xref ref-type="table" rid="table-26">Table 25</xref> presents the RL limitations in edge. Level 2 analysis identifies four dominant algorithmic families. Deep RL leads with 34% adoption processing high-dimensional topology inputs. Multi-agent systems follow at 27% coordinating distributed edge clusters. Tabular methods constitute 22% suiting constrained devices. Model-based approaches trail at 12% predicting failure dynamics. Deep RL demonstrates versatility across link and service recovery. DQN variants converge reliably within 5M training steps. Continuous control problems employ DDPG achieving 28% energy gains. GPU acceleration proves essential for practical training timelines. <xref ref-type="table" rid="table-25">Table 26</xref> presents the comparison of RL formulations in terms of convergence and scalability. Multi-agent RL addresses coordination challenges effectively. PPO algorithms balance stability and performance in 50-node clusters. Credit assignment mechanisms improve 22% over independent learners.</p>
<table-wrap id="table-25">
<label>Table 25</label>
<caption>
<title>RL limitations in edge IoT context.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Formulation</th>
<th>Edge Limitation</th>
<th>Impact</th>
<th>Mitigation Attempted</th>
</tr>
</thead>
<tbody>
<tr>
<td>Deep RL</td>
<td>GPU/memory &#x003E;1 GB</td>
<td>85% device exclusion</td>
<td>Quantization (limited)</td>
</tr>
<tr>
<td>Multi-agent</td>
<td>Comm overhead 25% energy</td>
<td>Battery drain</td>
<td>Topology optimization</td>
</tr>
<tr>
<td>Tabular</td>
<td>State explosion &#x003E;10K</td>
<td>Multi-node inapplicable</td>
<td>Abstraction (rare)</td>
</tr>
<tr>
<td>Model-based</td>
<td>2&#x2013;16 GB memory</td>
<td>90% exclusion</td>
<td>Lightweight models (0 studies)</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-26">
<label>Table 26</label>
<caption>
<title>Comparison of RL formulations in terms of convergence and scalability.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Formulation</th>
<th>Studies (%)</th>
<th>Convergence (M steps)</th>
<th>Memory Footprint</th>
<th>Scalability Limit</th>
</tr>
</thead>
<tbody>
<tr>
<td>Tabular</td>
<td>22</td>
<td>0.8</td>
<td>&#x003C;1 MB</td>
<td>10K states</td>
</tr>
<tr>
<td>Deep</td>
<td>34</td>
<td>5.2</td>
<td>500 MB&#x2013;2 GB</td>
<td>100 nodes</td>
</tr>
<tr>
<td>Multi-agent</td>
<td>27</td>
<td>12.0</td>
<td>1&#x2013;8 GB</td>
<td>50 nodes</td>
</tr>
<tr>
<td>Model-based</td>
<td>12</td>
<td>3.5</td>
<td>2&#x2013;16 GB</td>
<td>200 nodes</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Constraint Integration Patterns</title>
<p>Level 3 reveals energy dominance with 43% of studies enforcing power budgets explicitly. Latency constraints appear in 34% targeting deadline guarantees. Security integration trails at 18% quantifying exposure risks. Hybrid formulations constitute 5% balancing multiple objectives simultaneously. Energy constraints employ hard budget formulations predominantly. Joule-based penalties activate above 80% capacity thresholds. Recovery sequences limit to 15 actions maximum preventing battery exhaustion. <xref ref-type="table" rid="table-27">Table 27</xref> presents energy dominance declines from 60% to 40%. Hybrid formulations rise sharply post-2024.</p>
<table-wrap id="table-27">
<label>Table 27</label>
<caption>
<title>Constraint types, metrics, and penalty mechanisms.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Constraint Type</th>
<th>Studies (%)</th>
<th>Primary Metric</th>
<th>Penalty Mechanism</th>
</tr>
</thead>
<tbody>
<tr>
<td>Energy</td>
<td>43</td>
<td>Joules consumed</td>
<td>Budget exhaustion</td>
</tr>
<tr>
<td>Latency</td>
<td>34</td>
<td>Deadline violations</td>
<td>ms threshold</td>
</tr>
<tr>
<td>Security</td>
<td>18</td>
<td>Exposure duration</td>
<td>Risk accumulation</td>
</tr>
<tr>
<td>Hybrid</td>
<td>5</td>
<td>Pareto frontier</td>
<td>Lagrangian multipliers</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Cross-taxonomy analysis reveals three critical interaction patterns. Node recovery couples strongly with tabular methods and energy constraints forming 18% of all combinations. Service migration aligns with multi-agent RL and latency focus dominating 25% of intersections. Network orchestration exhibits formulation diversity but constraint sparsity limiting effectiveness. Performance trade-offs manifest systematically across levels. Energy-specialized approaches sacrifice 15% latency compliance. Security-focused studies reduce energy gains by 12%. Hybrid formulations demonstrate 8% balanced improvement vs. single-objective baselines. Network &#x002B; multi-agent &#x002B; security intersection contains 3% coverage vs. 28% node &#x002B; tabular &#x002B; energy dominance. Model-based approaches cluster around service migration exclusively. Production validation skews toward node recovery with 65% testbed coverage. These findings establish quantitative baselines for the field. Service migration maturity guides practitioner adoption. Node recovery provides energy-critical reference points. Under explored intersections signal research priorities clearly.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions and Future Scope</title>
<p>This review systematically analyzes Reinforcement Learning applications for self-healing in energy-constrained secure edge IoT networks. The work develops a novel three-level taxonomy classifying recovery scope, RL formulations, and constraint integration across 82 studies from 2020 to 2026. Key findings reveal service migration dominance with 30% coverage and energy constraint prevalence at 43%. Node recovery achieves highest energy savings of 38% through tabular methods. Multi-agent RL coordinates 50-node clusters effectively but requires 12M training steps. Hybrid constraint handling remains severely understudied at 5% coverage. Quantitative synthesis establishes normalized performance baselines enabling cross-study comparability. Energy savings standardize against static policies reaching 44% maximum gains. Latency compliance ratios measure 84% peak performance under mobility traces. Security exposure reduction quantifies 35% maximum improvement during failover windows. Deployment realism scores highlight 28% simulation inflation vs. testbed results. These baselines provide practitioners with validated reference points for implementation decisions. 10 evidence-based research gaps emerge through taxonomic mapping. Model-based node recovery demonstrates a complete absence from the literature. Multi-agent network orchestration under security constraints spans 2 papers only. Hybrid energy-latency-security formulations appear in 4 studies exclusively. Production deployment validation constitutes 8% coverage with 18% performance degradation. Continual learning brittleness affects 40% policy degradation under concept drift.</p>
<p>This systematic review acknowledges three primary limitations. First, the temporal scope covers publications from 2020 to 2026 only. Rapid evolution in RL-edge research may introduce significant post-2026 advances absent from current synthesis. Second, English-language bias exists due to the primary focus on peer-reviewed journals and conferences. Non-English grey literature potentially contains additional practical deployments not captured. Third, three-level taxonomy classification involves researcher judgment despite achieving 92% inter-rater agreement through dual independent coding. Alternative categorizations may emphasize different research gaps. These acknowledged limitations motivate the 15 prioritized future directions and deployment roadmap, ensuring practical applicability beyond identified methodological constraints. <xref ref-type="table" rid="table-28">Table 28</xref> illustrates the future directions such as (1) model-based RL for node recovery (70% efficiency); (2) Byzantine-resilient MARL orchestration (35% exposure reduction); (3) joint constraint Pareto optimization (22% gains); (4) hardware-in-loop testbeds; (5) standardized reward benchmarks. Implementation spans 2026&#x2013;2030: short-term model optimization and benchmarks (2026&#x2013;27), medium-term multi-agent security (2027&#x2013;28), long-term production platforms and IEEE standards (2028&#x2013;30).</p>
<table-wrap id="table-28">
<label>Table 28</label>
<caption>
<title>Categorized future research directions (Impact-feasibility ranked).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Direction (Gap)</th>
<th>Gain</th>
<th>Score</th>
<th>Year</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4"><bold>Algorithmic</bold></td>
<td>Model-based node recovery (G2)</td>
<td>70% eff.</td>
<td>9.2</td>
<td>2027</td>
</tr>
<tr>
<td>MARL security orchestration (G7)</td>
<td>35% exp.</td>
<td>8.9</td>
<td>2026</td>
</tr>
<tr>
<td>Hybrid tabular&#x2013;deep RL</td>
<td>3<inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> deploy.</td>
<td>8.4</td>
<td>2028</td>
</tr>
<tr>
<td>Safe exploration</td>
<td>Crash-free</td>
<td>7.9</td>
<td>2027</td>
</tr>
<tr>
<td rowspan="4"><bold>Evaluation</bold></td>
<td>Production benchmarks (G1)</td>
<td>Real valid.</td>
<td>9.5</td>
<td>2026</td>
</tr>
<tr>
<td>Adversarial testbeds</td>
<td>22% robust.</td>
<td>8.7</td>
<td>2027</td>
</tr>
<tr>
<td>Mobility traces</td>
<td>40% gen.</td>
<td>8.2</td>
<td>2026</td>
</tr>
<tr>
<td>Multi-failure emulation</td>
<td>28% cover.</td>
<td>7.6</td>
<td>2028</td>
</tr>
<tr>
<td rowspan="4"><bold>Constraints</bold></td>
<td>Joint Pareto fronts</td>
<td>22% multi-obj.</td>
<td>9.0</td>
<td>2027</td>
</tr>
<tr>
<td>Online adaptation</td>
<td>Dynamic budget</td>
<td>8.5</td>
<td>2026</td>
</tr>
<tr>
<td>Lagrangian enforcement</td>
<td>Hard const.</td>
<td>8.1</td>
<td>2027</td>
</tr>
<tr>
<td>Risk-aware rewards</td>
<td>18% security</td>
<td>7.7</td>
<td>2028</td>
</tr>
<tr>
<td rowspan="3"><bold>Scalability</bold></td>
<td>1000&#x002B; node MARL</td>
<td>Hierarchical</td>
<td>8.8</td>
<td>2028</td>
</tr>
<tr>
<td>Policy transfer</td>
<td>3<inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> accel.</td>
<td>8.3</td>
<td>2027</td>
</tr>
<tr>
<td>Federated learning</td>
<td>Privacy upd.</td>
<td>7.9</td>
<td>2026</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The author received no specific funding for this study.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The author declares no conflicts of interest.</p>
</sec>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A Methodology of Review</title>
<p><xref ref-type="fig" rid="fig-14">Fig. A1</xref> presents the PRISMA 2020 flow diagram documenting the systematic screening process: 1247 records identified across six databases reduced to 892 after duplicate removal, with 550 excluded at title/abstract screening and 260 full-text articles rejected (no RL &#x003D; 112, cloud-only &#x003D; 78, no failover &#x003D; 45, conceptual &#x003D; 25), yielding 82 studies for analysis. <xref ref-type="table" rid="table-29">Table A1</xref> summarizes MMAT quality assessment of these included studies, achieving 88% overall compliance. Perfect research question definition (100%) contrasts with reporting gaps (76%), where 20 studies omit convergence details or baselines. This rigorous PRISMA-guided process ensures methodological transparency and reproducibility while identifying simulation bias as primary quality limitation.</p>
<fig id="fig-14">
<label>Figure A1</label>
<caption>
<title>PRISMA 2020 flow diagram for systematic review.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80961-fig-14.tif"/>
</fig>
<table-wrap id="table-29">
<label>Table A1</label>
<caption>
<title>Quality assessment of 82 included studies (MMAT checklist).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Criteria (MMAT)</th>
<th>Compliant</th>
<th>Non-Compliant</th>
<th>Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clear research question</td>
<td>82</td>
<td>0</td>
<td>100%</td>
</tr>
<tr>
<td>Appropriate methodology</td>
<td>76</td>
<td>6</td>
<td>93%</td>
</tr>
<tr>
<td>Valid outcome measures</td>
<td>68</td>
<td>14</td>
<td>83%</td>
</tr>
<tr>
<td>Complete reporting</td>
<td>62</td>
<td>20</td>
<td>76%</td>
</tr>
<tr>
<td><bold>Overall Average</bold></td>
<td></td>
<td></td>
<td><bold>88%</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mazaheri</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Ameli</surname> <given-names>S</given-names></string-name>, <string-name><surname>Abedi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abari</surname> <given-names>O</given-names></string-name></person-group>. <article-title>A millimeter wave network for billions of things</article-title>. In: <conf-name>Proceedings of the ACM Special Interest Group on Data Communication. SIGCOMM &#x2019;19</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2019</year>. p. <fpage>174</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3341302.3342068</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>J</given-names></string-name>, <string-name><surname>McElhannon</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Future edge cloud and edge computing for internet of things applications</article-title>. <source>IEEE Internet Things J</source>. <year>2017</year>;<volume>5</volume>(<issue>1</issue>):<fpage>439</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2017.2767608</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Stellios</surname> <given-names>I</given-names></string-name>, <string-name><surname>Kotzanikolaou</surname> <given-names>P</given-names></string-name>, <string-name><surname>Psarakis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alcaraz</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lopez</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A survey of IoT-enabled cyberattacks: assessing attack paths to critical infrastructures and services</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2018</year>;<volume>20</volume>(<issue>4</issue>):<fpage>3453</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Das</surname> <given-names>R</given-names></string-name>, <string-name><surname>G&#x00FC;nd&#x00FC;z</surname> <given-names>MZ</given-names></string-name></person-group>. <article-title>Analysis of cyber-attacks in IoT-based critical infrastructures</article-title>. <source>Int J Inf Secur Sci</source>. <year>2019</year>;<volume>8</volume>(<issue>4</issue>):<fpage>122</fpage>&#x2013;<lpage>33</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mohapatra</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A comprehensive review on urban resilience via fault-tolerant IoT and sensor networks</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>85</volume>(<issue>1</issue>):<fpage>221</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.068338</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tosatto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ruiu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Attanasio</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Container-based orchestration in cloud: state of the art and challenges</article-title>. In: <conf-name>2015 Ninth International Conference on Complex, Intelligent, and Software Intensive Systems</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>70</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>GH</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>HT</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Management and orchestration of edge computing for IoT: a comprehensive survey</article-title>. <source>IEEE Internet Things J</source>. <year>2023</year>;<volume>10</volume>(<issue>16</issue>):<fpage>14307</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Baek</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ko</surname> <given-names>H</given-names></string-name>, <string-name><surname>Park</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pack</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kwak</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A two-stage failover mechanism for high availability in service function chaining</article-title>. <source>J Internet Technol</source>. <year>2018</year>;<volume>19</volume>(<issue>1</issue>):<fpage>229</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>KP</given-names></string-name>, <string-name><surname>Sivanesan</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Flow rule-based routing protocol management system in software-defined IoT sensor network for IoT applications</article-title>. <source>Int J Commun Syst</source>. <year>2022</year>;<volume>35</volume>(<issue>11</issue>):<fpage>e5182</fpage>. doi:<pub-id pub-id-type="doi">10.1002/dac.5182</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ibrahim</surname> <given-names>F</given-names></string-name>, <string-name><surname>Rehman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alzghoul</surname> <given-names>AHA</given-names></string-name></person-group>. <article-title>Energy-efficient hybrid cryptographic framework for resource-constrained IoT devices</article-title>. <source>Spec Eng Sci</source>. <year>2025</year>;<volume>3</volume>(<issue>12</issue>):<fpage>346</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bod&#x00ED;k</surname> <given-names>P</given-names></string-name>, <string-name><surname>Menache</surname> <given-names>I</given-names></string-name>, <string-name><surname>Chowdhury</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mani</surname> <given-names>P</given-names></string-name>, <string-name><surname>Maltz</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Stoica</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Surviving failures in bandwidth-constrained datacenters</article-title>. <source>ACM SIGCOMM Comput Commun Rev</source>. <year>2012</year>;<volume>42</volume>(<issue>4</issue>):<fpage>431</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2377677.2377760</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Min</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gokhale</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shekhar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mahmoudi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Barve</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Enhancing 5G network slicing for IoT traffic with a novel clustering framework</article-title>. <source>Pervasive Mob Comput</source>. <year>2024</year>;<volume>104</volume>(<issue>136</issue>):<fpage>101974</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.pmcj.2024.101974</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Deep reinforcement learning for autonomous internet of things: model, applications and challenges</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2020</year>;<volume>22</volume>(<issue>3</issue>):<fpage>1722</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2020.2988367</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>J</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Orientation-preserving rewards&#x2019; balancing in reinforcement learning</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2021</year>;<volume>33</volume>(<issue>11</issue>):<fpage>6458</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2021.3080521</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mohapatra</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Task offloading and edge computing in IoT&#x2014;gaps, challenges and future directions</article-title>. <source>Comput Mater Contin</source>. <year>2026</year>;<volume>87</volume>(<issue>3</issue>):<fpage>8</fpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2026.076726</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ebad</surname> <given-names>SA</given-names></string-name></person-group>. <article-title>Quantifying IoT security parameters: an assessment framework</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>101087</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Murshed</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Murphy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>D</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ananthanarayanan</surname> <given-names>G</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Machine learning at the network edge: a survey</article-title>. <source>ACM Comput Surv</source>. <year>2021</year>;<volume>54</volume>(<issue>8</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3469029</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nassar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yilmaz</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Reinforcement learning for adaptive resource allocation in fog RAN for IoT with heterogeneous latency requirements</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>128014</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2939735</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Parameterized deep reinforcement learning with hybrid action space for edge task offloading</article-title>. <source>IEEE Internet Things J</source>. <year>2023</year>;<volume>11</volume>(<issue>6</issue>):<fpage>10754</fpage>&#x2013;<lpage>67</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Griffith</surname> <given-names>D</given-names></string-name>, <string-name><surname>Golmie</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Survey of graph neural networks and applications</article-title>. <source>Wirel Commun Mob Comput</source>. <year>2022</year>;<volume>2022</volume>(<issue>1</issue>):<fpage>9261537</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2022/9261537</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Collaborative multi-agent deep reinforcement learning for energy-efficient resource allocation in heterogeneous mobile edge computing networks</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2023</year>;<volume>23</volume>(<issue>6</issue>):<fpage>6653</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1109/twc.2023.3335597</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>QW</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Rizwan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>R</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>DH</given-names></string-name></person-group>. <article-title>Decentralized machine learning training: a survey on synchronization, consolidation, and topologies</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>68031</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2023.3284976</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A review of continual learning in edge AI</article-title>. <source>IEEE Trans Netw Sci Eng</source>. <year>2026</year>;<volume>13</volume>:<fpage>6571</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnse.2026.3657652</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Spina</surname> <given-names>MG</given-names></string-name>, <string-name><surname>Boukerche</surname> <given-names>A</given-names></string-name>, <string-name><surname>De Rango</surname> <given-names>F</given-names></string-name></person-group>. <article-title>An IoE-powered framework for adaptive energy-security trade-off in IoT</article-title>. <source>IEEE Netw</source>. <year>2026</year>;<volume>40</volume>(<issue>2</issue>):<fpage>65</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mnet.2025.3636907</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Cognitive edge computing: a comprehensive survey on optimizing large models and AI agents for pervasive deployment</article-title>. <comment>arXiv:2501.03265. 2025</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Moghaddasi</surname> <given-names>K</given-names></string-name>, <string-name><surname>Rajabi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hosseinzadeh</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Intrusion detection systems for enhanced security in mobile edge computing: a systematic review and survey of the applications, challenges, and future directions</article-title>. <source>Wirel Pers Commun</source>. <year>2025</year>;<volume>145</volume>(<issue>1&#x2013;2</issue>):<fpage>113</fpage>&#x2013;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Casado</surname> <given-names>FE</given-names></string-name>, <string-name><surname>Lema</surname> <given-names>D</given-names></string-name>, <string-name><surname>Criado</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Iglesias</surname> <given-names>R</given-names></string-name>, <string-name><surname>Regueiro</surname> <given-names>CV</given-names></string-name>, <string-name><surname>Barro</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Concept drift detection and adaptation for federated and continual learning</article-title>. <source>Multimed Tools Appl</source>. <year>2022</year>;<volume>81</volume>(<issue>3</issue>):<fpage>3397</fpage>&#x2013;<lpage>419</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-021-11219-x</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>F</given-names></string-name>, <string-name><surname>Leung</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Drew</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Topology-aware federated learning in edge computing: a comprehensive survey</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>10</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3659205</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Menaka</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sapaev</surname> <given-names>I</given-names></string-name>, <string-name><surname>Dadaxon</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ulkanov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Praveenkumar</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Deep reinforcement learning for self-healing communication networks: addressing node failure and QoS degradation in dynamic topologies</article-title>. <source>Natl J Antennas Propag</source>. <year>2025</year>;<volume>7</volume>(<issue>2</issue>):<fpage>133</fpage>&#x2013;<lpage>44</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Moorthy</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Jagannath</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Survey of graph neural network for internet of things and NextG networks</article-title>. <comment>arXiv:2405.17309. 2024</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arun Priya</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ramakrishnan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Optimizing LoRaWAN performance through reinforcement Q-convolutional deterministic policy gradient: a comprehensive approach to efficient resource allocation</article-title>. <source>Wirel Netw</source>. <year>2025</year>;<volume>31</volume>:<fpage>4763</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11276-025-04025-y</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Salah</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Saad</surname> <given-names>RS</given-names></string-name>, <string-name><surname>Zaki</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Rabie</surname> <given-names>K</given-names></string-name>, <string-name><surname>ElHalawany</surname> <given-names>BM</given-names></string-name></person-group>. <article-title>Multi-armed bandits for resource allocation in UAV-assisted lora networks</article-title>. <source>IEEE Internet Things Mag</source>. <year>2025</year>;<volume>8</volume>(<issue>2</issue>):<fpage>40</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iotm.001.2400088</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multi-agent proximal policy optimization-based dynamic client selection for federated AI in 6G-oriented internet of vehicles</article-title>. <source>IEEE Trans Veh Technol</source>. <year>2024</year>;<volume>73</volume>(<issue>9</issue>):<fpage>13611</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tvt.2024.3383860</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pateria</surname> <given-names>S</given-names></string-name>, <string-name><surname>Subagdja</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tanh</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Quek</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Hierarchical reinforcement learning: a comprehensive survey</article-title>. <source>ACM Comput Surv</source>. <year>2021</year>;<volume>54</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3453160</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ben-Othman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Edge AI-enabled backbone optimization for real-time object detection in computing power networks</article-title>. <source>IEEE Trans Cogn Commun Netw</source>. <year>2026</year>;<volume>12</volume>:<fpage>5891</fpage>&#x2013;<lpage>902</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tccn.2026.3659849</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Akhtar</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Ghafoor</surname> <given-names>U</given-names></string-name>, <string-name><surname>Imran</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ayub</surname> <given-names>N</given-names></string-name>, <string-name><surname>Abdullah</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>An efficient AI and deep learning assisted self-healing network approach: analysis on fault detection response and recovery to mitigate threats in IoT-security ecosystem</article-title>. <source>Asian Bull of Big Data Manag</source>. <year>2026</year>;<volume>6</volume>(<issue>1</issue>):<fpage>40</fpage>&#x2013;<lpage>66</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Testing self-healing cyber-physical systems under uncertainty with reinforcement learning: an empirical study</article-title>. <source>Empir Softw Eng</source>. <year>2021</year>;<volume>26</volume>(<issue>3</issue>):<fpage>52</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10664-021-09941-z</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Adeniyi</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sadiq</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Pillai</surname> <given-names>P</given-names></string-name>, <string-name><surname>Taheir</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Kaiwartya</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Proactive self-healing approaches in mobile edge computing: a systematic literature review</article-title>. <source>Computers</source>. <year>2023</year>;<volume>12</volume>(<issue>3</issue>):<fpage>63</fpage>. doi:<pub-id pub-id-type="doi">10.3390/computers12030063</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vaishnav</surname> <given-names>S</given-names></string-name>, <string-name><surname>Magn&#x00FA;sson</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>Multi-objective and constrained reinforcement learning for IoT</chapter-title>. In: <source>Learning techniques for the internet of things</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>153</fpage>&#x2013;<lpage>70</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kayas</surname> <given-names>G</given-names></string-name>, <string-name><surname>Hasan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Skjellum</surname> <given-names>A</given-names></string-name>, <string-name><surname>Noor</surname> <given-names>S</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>SR</given-names></string-name></person-group>. <article-title>A holistic analysis of internet of things (IoT) security: principles, practices, and new perspectives</article-title>. <source>Future Internet</source>. <year>2024</year>;<volume>16</volume>(<issue>2</issue>):<fpage>40</fpage>. doi:<pub-id pub-id-type="doi">10.3390/fi16020040</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep-reinforcement-learning-based production scheduling in industrial internet of things</article-title>. <source>IEEE Internet Things J</source>. <year>2023</year>;<volume>10</volume>(<issue>22</issue>):<fpage>19725</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2023.3283056</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Uprety</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rawat</surname> <given-names>DB</given-names></string-name></person-group>. <article-title>Reinforcement learning for IoT security: a comprehensive survey</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>8</volume>(<issue>11</issue>):<fpage>8693</fpage>&#x2013;<lpage>706</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2020.3040957</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Adawadkar</surname> <given-names>AMK</given-names></string-name>, <string-name><surname>Kulkarni</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Cyber-security and reinforcement learning&#x2014;a brief survey</article-title>. <source>Eng Appl Artif Intell</source>. <year>2022</year>;<volume>114</volume>:<fpage>105116</fpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>HN</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Deep reinforcement learning for internet of things: a comprehensive survey</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2021</year>;<volume>23</volume>(<issue>3</issue>):<fpage>1659</fpage>&#x2013;<lpage>92</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alipio</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bures</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep reinforcement learning perspectives on improving reliable transmissions in IoT networks: problem formulation, parameter choices, challenges, and future directions</article-title>. <source>Internet Things</source>. <year>2023</year>;<volume>23</volume>:<fpage>100846</fpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gherbi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Senouci</surname> <given-names>O</given-names></string-name>, <string-name><surname>Harbi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Medani</surname> <given-names>K</given-names></string-name>, <string-name><surname>Aliouat</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A systematic literature review of machine learning applications in IoT</article-title>. <source>Int J Commun Syst</source>. <year>2023</year>;<volume>36</volume>(<issue>11</issue>):<fpage>e5500</fpage>. doi:<pub-id pub-id-type="doi">10.1002/dac.5500</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pinto Neto</surname> <given-names>EC</given-names></string-name>, <string-name><surname>Sadeghi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dadkhah</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Federated reinforcement learning in IoT applications, opportunities and open challenges</article-title>. <source>Appl Sci</source>. <year>2023</year>;<volume>13</volume>(<issue>11</issue>):<fpage>6497</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app13116497</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sivamayil</surname> <given-names>K</given-names></string-name>, <string-name><surname>Rajasekar</surname> <given-names>E</given-names></string-name>, <string-name><surname>Aljafari</surname> <given-names>B</given-names></string-name>, <string-name><surname>Nikolovski</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vairavasundaram</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vairavasundaram</surname> <given-names>I</given-names></string-name></person-group>. <article-title>A systematic study on reinforcement learning based applications</article-title>. <source>Energies</source>. <year>2023</year>;<volume>16</volume>(<issue>3</issue>):<fpage>1512</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en16031512</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Amodu</surname> <given-names>OA</given-names></string-name>, <string-name><surname>Jarray</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>RAR</given-names></string-name>, <string-name><surname>Althumali</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bukar</surname> <given-names>UA</given-names></string-name>, <string-name><surname>Nordin</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep reinforcement learning for AoI minimization in UAV-aided data collection for WSN and IoT applications: a survey</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>108000</fpage>&#x2013;<lpage>40</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A bibliometric analysis of edge computing for internet of things</article-title>. <source>Secur Commun Netw</source>. <year>2021</year>;<volume>2021</volume>(<issue>1</issue>):<fpage>5563868</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/5563868</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abkenar</surname> <given-names>FS</given-names></string-name>, <string-name><surname>Ramezani</surname> <given-names>P</given-names></string-name>, <string-name><surname>Iranmanesh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Murali</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chulerttiyawong</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on mobility of edge computing networks in IoT: state-of-the-art, architectures, and challenges</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2022</year>;<volume>24</volume>(<issue>4</issue>):<fpage>2329</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2022.3211462</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zabihi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Eftekhari Moghadam</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Rezvani</surname> <given-names>MH</given-names></string-name></person-group>. <article-title>Reinforcement learning methods for computation offloading: a systematic review</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>56</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3603703</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Priyadarshi</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Exploring machine learning solutions for overcoming challenges in IoT-based wireless sensor network routing: a comprehensive review</article-title>. <source>Wirel Netw</source>. <year>2024</year>;<volume>30</volume>(<issue>4</issue>):<fpage>2647</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11276-024-03697-2</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ameedeen</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Kamarudin</surname> <given-names>IE</given-names></string-name>, <string-name><surname>Ab Razak</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Zabidi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Integrating edge computing and software defined networking in internet of things: a systematic review</article-title>. <source>Iraqi J Comput Sci Math</source>. <year>2023</year>;<volume>4</volume>(<issue>4</issue>):<fpage>11</fpage>. doi:<pub-id pub-id-type="doi">10.52866/ijcsm.2023.04.04.011</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Al Arafat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Machine learning in real-time Internet of Things (IoT) systems: a survey</article-title>. <source>IEEE Internet Things J</source>. <year>2022</year>;<volume>9</volume>(<issue>11</issue>):<fpage>8364</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2022.3161050</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Malik</surname> <given-names>TS</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Afzal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ibrar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Song</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RL-IoT: reinforcement learning-based routing approach for cognitive radio-enabled IoT communications</article-title>. <source>IEEE Internet Things J</source>. <year>2022</year>;<volume>10</volume>(<issue>2</issue>):<fpage>1836</fpage>&#x2013;<lpage>47</lpage>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yahuza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Idris</surname> <given-names>MYIB</given-names></string-name>, <string-name><surname>Wahab</surname> <given-names>AWBA</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>AT</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Musa</surname> <given-names>SNB</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Systematic review on security and privacy requirements in edge computing: state of the art and future research opportunities</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>76541</fpage>&#x2013;<lpage>67</lpage>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hamdan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ayyash</surname> <given-names>M</given-names></string-name>, <string-name><surname>Almajali</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Edge-computing architectures for internet of things applications: a survey</article-title>. <source>Sensors</source>. <year>2020</year>;<volume>20</volume>(<issue>22</issue>):<fpage>6441</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s20226441</pub-id>; <pub-id pub-id-type="pmid">33187267</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rafique</surname> <given-names>W</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yaqoob</surname> <given-names>I</given-names></string-name>, <string-name><surname>Imran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rasool</surname> <given-names>RU</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Complementing IoT services through software defined networking and edge computing: a comprehensive survey</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2020</year>;<volume>22</volume>(<issue>3</issue>):<fpage>1761</fpage>&#x2013;<lpage>804</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2020.2997475</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Nallanathan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chambers</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>Reinforcement learning for real-time optimization in NB-IoT networks</article-title>. <source>IEEE J Sel Areas Commun</source>. <year>2019</year>;<volume>37</volume>(<issue>6</issue>):<fpage>1424</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsac.2019.2904366</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Amodu</surname> <given-names>OA</given-names></string-name>, <string-name><surname>Althumali</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hanapi</surname> <given-names>ZM</given-names></string-name>, <string-name><surname>Jarray</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>RAR</given-names></string-name>, <string-name><surname>Adam</surname> <given-names>MS</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive survey of deep reinforcement learning in UAV-assisted IoT data collection</article-title>. <source>Veh Commun</source>. <year>2025</year>;<volume>55</volume>(<issue>2</issue>):<fpage>100949</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.vehcom.2025.100949</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Siddiqui</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hameed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shah</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>I</given-names></string-name>, <string-name><surname>Aneiba</surname> <given-names>A</given-names></string-name>, <string-name><surname>Draheim</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Toward software-defined networking-based IoT frameworks: a systematic literature review, taxonomy, open challenges and prospects</article-title>. <source>IEEE Access</source>. <year>2022</year>;<volume>10</volume>:<fpage>70850</fpage>&#x2013;<lpage>901</lpage>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Andro&#x010D;ec</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Applications of edge analytics: a systematic review</article-title>. <source>Acta Univ Sapientiae Inform</source>. <year>2023</year>;<volume>15</volume>(<issue>2</issue>):<fpage>345</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.2478/ausi-2023-0021</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Muhammad</surname> <given-names>G</given-names></string-name>, <string-name><surname>Amin</surname> <given-names>SU</given-names></string-name></person-group>. <article-title>Edge intelligence in the cognitive internet of things: improving sensitivity and interactivity</article-title>. <source>IEEE Netw</source>. <year>2019</year>;<volume>33</volume>(<issue>3</issue>):<fpage>58</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mnet.2019.1800344</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weng</surname> <given-names>O</given-names></string-name>, <string-name><surname>Meza</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bock</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hawks</surname> <given-names>B</given-names></string-name>, <string-name><surname>Campos</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tran</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fkeras: a sensitivity analysis tool for edge neural networks</article-title>. <source>J Auton Transp Syst</source>. <year>2024</year>;<volume>1</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>27</lpage>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ben-Othman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Split federated learning-enabled deep Q-networks for generalized path planning in distributed IoT edge platform</article-title>. <source>IEEE Internet Things J</source>. <year>2026</year>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2026.3670351</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>





























