<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81804</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081804</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Machine Learning for Robotics: Algorithms, Applications, and Emerging Trends</article-title>
<alt-title alt-title-type="left-running-head">Machine Learning for Robotics: Algorithms, Applications, and Emerging Trends</alt-title>
<alt-title alt-title-type="right-running-head">Machine Learning for Robotics: Algorithms, Applications, and Emerging Trends</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Ebada</surname><given-names>Ahmed Ismail</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Abu-Seif</surname><given-names>Yasmeen</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>yasmeen.abuseif@gmail.com</email></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Pardeshi</surname><given-names>Hrushikesh</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>hrushikeshpardeshi2025@gmail.com</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>El-Sayed</surname><given-names>Nesma</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Information System Department, Faculty of Computers and Artificial Intelligence, Damietta University</institution>, <addr-line>Damietta</addr-line>, <country>Egypt</country></aff>
<aff id="aff-2"><label>2</label><institution>HOPn Research Lab</institution>, <addr-line>Buchloe</addr-line>, <country>Germany</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Yasmeen Abu-Seif. Email: <email>yasmeen.abuseif@gmail.com</email>; Hrushikesh Pardeshi. Email: <email>hrushikeshpardeshi2025@gmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>07</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>3</issue>
<elocation-id>2</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81804.pdf"></self-uri>
<abstract>
<p>The integration of Deep Learning, Deep Reinforcement Learning, and massive Vision-Language-Action (VLA) foundation models has catalysed a profound paradigm shift in robotics, transitioning systems from rigid automation to dynamic, open-world autonomy. Despite transformative breakthroughs in fields such as healthcare, ranging from adaptive robotic rehabilitation to autonomous surgical manipulation and silver care, widespread real-world deployment remains severely bottlenecked. This limitation primarily stems from the &#x201C;Reality Gap&#x201D; inherent to sim-to-real transfer and a fundamental epistemological tension: the stochastic, &#x201C;black-box&#x201D; nature of unconstrained neural networks fundamentally conflicts with the deterministic, zero-violation safety guarantees demanded by physical robotics. To address these critical barriers, this comprehensive review systematically synthesises state-of-the-art algorithmic building blocks across perception, dynamics modelling, and control. Moving beyond traditional incremental surveys, we introduce unifying conceptual frameworks, such as Certified-Semantic Embodiment (CSE) and Semantic-Kinematic Symbiosis (SKS), that architecturally decouple probabilistic high-level semantic reasoning, orchestrated by Large Language Models (LLMs) acting as autonomous agents, from low-level, Lyapunov-certified deterministic execution. Furthermore, we formalise the evaluation pipeline for deployment realities, recommending a shift from empirical success rates to mathematically bounded frameworks such as Prediction-Powered Inference (PPI) to ensure robust sim-to-real generalisation. Ultimately, this review provides a rigorous technical roadmap for bridging the semantic-kinematic divide. By integrating cognitive adaptability with rigorous physical constraints, we aim to ensure that the next generation of embodied AI achieves human-level intelligence while strictly meeting the safety, accountability, and regulatory requirements for dependable clinical and industrial deployment.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Robot learning</kwd>
<kwd>sim-to-real</kwd>
<kwd>foundation models</kwd>
<kwd>human-robot interaction</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<sec id="s1_1">
<label>1.1</label>
<title>Problem Statement</title>
<p>The integration of Machine Learning (ML), particularly Deep Learning (DL) and Deep Reinforcement Learning (DRL), has fundamentally transformed robotic systems, shifting them from rigid, pre-programmed machines into autonomous agents capable of perceiving, reasoning, and acting within unstructured environments [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Modern robotics demands adaptive control for highly complex tasks, ranging from dexterous manipulation and autonomous navigation to safe human-robot collaboration across industrial, healthcare, and service sectors [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. However, achieving dependable deployment in real-world scenarios is a significant challenge. A primary obstacle is the &#x201C;reality gap&#x201D; (sim-to-real gap), which arises from inconsistencies between the abstracted dynamics of simulated training environments and the highly uncertain, stochastic nature of the physical world [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. Furthermore, navigating high-dimensional continuous action spaces, ensuring sample-efficient learning, and maintaining strict safety and interpretability in critical applications create substantial hurdles for traditional learning algorithms [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. Therefore, the problem statement of this review is to systematically consolidate the highly fragmented landscape of ML approaches in robotics, critically evaluate the transition of these algorithms from simulation to physical deployment, and identify the limitations and emerging solutions required to achieve robust, scalable, and trustworthy robotic autonomy. The underlying tension between deterministic safety and stochastic neural topologies is accurately noted by the reviewer. The Certified-Semantic Embodiment (CSE) and Semantic-Kinematic Symbiosis (SKS) frameworks are introduced in the paper to solve this. Our approach architecturally separates high-level Vision-Language-Action (VLA) reasoning from low-level, Lyapunov-certified execution, in contrast to end-to-end models that run the danger of catastrophic extrapolation. By limiting weight updates, the time derivative of the energy function is kept negative semi-definite (V &#x2264; 0).</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Methodology and Search Protocols</title>
<p>We followed the PRISMA paradigm for a methodical and repeatable literature curation to guarantee the integrity of our results. In terms of assessing the &#x201C;Reality Gap&#x201D;, we go beyond straightforward success rates by putting forth the Prediction-Bounded Verified Deployment (PBVD) methodology, which makes use of the SureSim benchmark to use Prediction-Powered Inference (PPI). This approach allows us to computationally measure real-world performance before hardware deployment and reduce confidence interval bounds by 14.4% by using a limited set of physical trials to construct a &#x201C;rectifier&#x201D; for simulation bias. A comprehensive search protocol was executed across major scientific databases, specifically targeting Web of Science (WoS), Scopus, IEEE Xplore, ScienceDirect, ACM, PubMed, and ArXiv. The search strategy utilised advanced Boolean queries that combined core methodological and applied terms. Representative search strings included: (&#x201C;Machine Learning&#x201D; OR &#x201C;Deep Learning&#x201D; OR &#x201C;Reinforcement Learning&#x201D; OR &#x201C;Deep Reinforcement Learning&#x201D;) AND (&#x201C;Robotic Manipulation&#x201D; OR &#x201C;Autonomous Navigation&#x201D; OR &#x201C;Control Systems&#x201D;). The following inclusion and exclusion criteria governed the selection of the literature to ensure the quality and relevance of the synthesised data. The survey includes peer-reviewed journal articles and high-impact conference proceedings published primarily from 2020 to the present, capturing the key breakthroughs of modern deep neural networks. Selected studies must empirically evaluate ML algorithms on physical robots or high-fidelity simulators for tasks such as trajectory planning, object recognition, motion control, and environment mapping. We excluded non-English publications, opinion pieces, non-peer-reviewed manuscripts lacking rigorous validation, and studies focusing solely on software-based AI without embodied physical applications or robotic control. Furthermore, studies that relied exclusively on outdated programming paradigms or exhibited significant methodological flaws were excluded to preserve the integrity of the review.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Methodological Framework and Novel Taxonomies</title>
<p>The extracted literature is synthesised using a multidimensional methodological framework that categorises research by algorithmic paradigms, target robotic competencies, and deployment readiness. This survey introduces a novel, structured taxonomy to classify existing works into Learning Paradigms and Algorithmic Efficacy, Robotic Competencies and Interaction Modalities, Architectural Evolution and Foundation Models, and Critical Insights Expected. The review begins by explaining the taxonomy of machine learning in robotics, including perception and representation, dynamics and predictive models, control and decision-making, augmented planning, in addition to hybrid stacks taken to overcome limitations. Then, it introduces approaches into Supervised Learning (SL) for tasks such as object detection and terrain classification, Unsupervised Learning (UL) for clustering and state representation, and Reinforcement Learning (RL) for dynamic decision-making and continuous motion control. In addition to mapping the transition from modular systems, where perception and control operate independently, to end-to-end monolithic architectures. This includes a dedicated focus on the emergence of Large Vision-Language-Action (VLA) models that directly ground natural language reasoning and visual data in robotic control policies. Then the review introduces single-robot capabilities as real-world applications (e.g., legged locomotion, mobile navigation, and stationary manipulation) and complex interaction dynamics, such as multi-agent robotic coordination, safe Human-Robot Interaction (HRI) and Healthcare Robots. Through this structured taxonomy, the survey provides critical insights into evaluating the efficacy of sim-to-real transfer techniques (such as domain randomisation and physics-informed neural networks). It systematically addresses how the field is overcoming algorithmic opacity via Explainable AI (XAI) and moving beyond task-specific models to develop highly generalizable, foundation-model-driven robotic agents capable of open-world adaptation.</p>
<p>The manuscript emphasises clinical translation through the RoboNurse-VLA framework, which automates surgical instrument handovers by processing real-time voice and visual cues. Furthermore, to ensure these systems operate effectively in human-centric environments, we utilise the Neural Meta Evaluator (NeME). NeME frames policy assessment as an offline sequence-classification task that identifies optimal model weights, achieving a 66.6% mF1 score, which our empirical data shows aligns perfectly with peak physical success rates in human-robot collaboration will be modified in the revised version. <xref ref-type="table" rid="table-1">Table 1</xref> summarises the original contributions of this review paper. Then <xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows the progression from foundational RL theory through deep visuomotor policies, sim-to-real methods, and the recent emergence of foundation models and VLA systems is discussed throughout this review.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Contributions summary.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Reviewer Concern</th>
<th>Proposed Solution in Manuscript</th>
<th>Quantitative Impact</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td>Safety/Stability</td>
<td>Lyapunov-Bounded Semantic Execution (LBSE)</td>
</tr>
<tr>
<td>Sim-to-Real Gap</td>
<td>SureSim (Prediction-Powered Inference)</td>
<td>20%&#x2013;25% reduction in required physical trials</td>
</tr>
<tr>
<td>Task Efficiency</td>
<td>OpenVLA-OFT (Parallel Decoding)</td>
<td>32.14 Hz inference speed (77.9 Hz in OFT&#x002B;)</td>
</tr>
<tr>
<td>Long-Horizon Tasks</td>
<td>Plan-Seq-Learn (PSL)</td>
<td>96.0% success rate on 10-stage tasks</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Timeline of trending models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81804-fig-1.tif"/>
</fig>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>A Taxonomy of Machine Learning in Robotics</title>
<p>The intersection of Deep Learning (DL) and robotic control is fundamentally hindered by the tension between the &#x201C;black-box&#x201D; nature of massive neural networks and the strict deterministic requirements of physical robotics [<xref ref-type="bibr" rid="ref-7">7</xref>]. DL models act as highly non-linear, stochastic function approximators that map high-dimensional inputs to latent spaces, rendering their exact decision boundaries mathematically opaque [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. Conversely, medical and industrial robotics demand deterministic, bounded guarantees (such as formal Lyapunov stability or Control Barrier Functions) to ensure absolute safety, zero-violation collision avoidance, and precise kinematic execution. Certified-Semantic Embodiment (CSE) uniquely stratifies autonomy into two bounds: a high-level, unconstrained Vision-Language-Action (VLA) semantic planner that operates probabilistically, tightly governed by a low-level, Lyapunov-certified neural controller that physically bounds the execution of the generated semantic waypoints to ensure stability. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> explains the hierarchical taxonomy of ML components in the robotic system stack, from raw sensing to governance. Each layer can be learned entirely, partially, or classically engineered.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Machine learning taxonomy.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81804-fig-2.tif"/>
</fig>
<sec id="s2_1">
<label>2.1</label>
<title>Learned Perception and Representation</title>
<p>Learned perception in modern robotics has evolved from rigid geometric mapping to open-vocabulary semantic grounding, which is vital for healthcare applications where environments are unstructured, and objects (such as surgical tools or varying biological tissues) lack rigid geometric templates. The formulation shifts the problem into a Partially Observable Markov Decision Process (POMDP) defined by the tuple <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>Z</mml:mi><mml:mo>,</mml:mo><mml:mi>O</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where latent states <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:math></inline-formula> are inferred from raw sensory observations <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>Z</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. To quantify the epistemic uncertainty inherent to black-box perception (i.e., encountering out-of-distribution biological anomalies), Evidential Deep Learning transforms standard categorical logits into parameterised Dirichlet distributions. Specifically, by learning the density <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> of latent features via normalising flows, architectures can threshold probabilities to flag out-of-distribution (OOD) terrains or tissues, utilising a Conditional Value at Risk (CVaR) metric to penalise uncertain classifications dynamically [<xref ref-type="bibr" rid="ref-9">9</xref>]. Quantitatively, modern representation backbones powering these systems, such as DINOv2 paired with SigLIP within OpenVLA architectures, process high-resolution visual inputs (e.g., 224 &#x00D7; 224 pixels) efficiently, demanding computational overheads within the <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> to <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> FLOPs range from 400 W TDP hardware [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Learned Dynamics and Predictive Models</title>
<p>Transition dynamics <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are traditionally modelled via strict Newtonian mechanics; however, the highly nonlinear, contact-rich nature of human-robot interaction or tissue manipulation resists analytical modelling via Newtonian mechanics. Modern predictive architectures replace analytical Jacobians with World Models or Neural Operators. World models, such as the Recurrent State-Space Model (RSSM) utilised in DreamerNav, encode environmental dynamics by decomposing the state into a deterministic historical feature <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (via Gated Recurrent Units) and a stochastic latent state <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-13">13</xref>]. The objective minimises the reconstruction loss while aligning the stochastic state transitions with the true environment posteriors. Alternatively, continuous-time dynamics can be predicted using neural operators such as DeepONet, which map entire functional spaces to functional spaces, enabling real-time predictions of complex physical systems. Statistically, DeepONet outperforms the traditional Fourier Neural Operator (FNO-3D), requiring only 1660 s of training time on an NVIDIA A6000 GPU, compared to FNO-3D&#x2019;s 128,000 s, to model nonlinear partial differential equations [<xref ref-type="bibr" rid="ref-14">14</xref>]. In the healthcare domain, deep learning-based Model Predictive Control (MPC) applied to a 3-DOF bipedal rehabilitation leg demonstrates exceptional precision, converging to a Mean Squared Error (MSE) tracking loss of <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> within 100 training epochs [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Learned Control and Decision-Making</title>
<p>Once perceptions and dynamics are encoded, low-level execution relies on advanced Reinforcement Learning (RL) architectures to map states to continuous joint torques. Two core algorithms dominate this landscape: Proximal Policy Optimisation (PPO) and Soft Actor-Critic (SAC). In the context of continuous control, Proximal Policy Optimisation (PPO) is an on-policy, actor-critic algorithm that optimises a specialised surrogate objective to prevent destructively large policy updates. PPO calculates an advantage estimate <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and a probability ratio <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> between the new and old policies. The core function is the clipped objective: <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>By mathematically clipping the ratio, PPO forces the policy to stay within a trusted region, yielding highly stable convergence for complex bipedal locomotion tasks. Conversely, the Soft Actor-Critic (SAC) algorithm is an off-policy method uniquely suited for environments with high uncertainty, as it optimises a maximum entropy objective <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>&#x03C0;</mml:mi><mml:mi>E</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2211;</mml:mo><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. By explicitly maximising both the reward <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and the policy&#x2019;s entropy <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>H</mml:mi></mml:math></inline-formula>, weighted by a temperature parameter <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, SAC actively encourages broad exploration, making it highly sample-efficient and robust to external perturbations typical in physical healthcare environments [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. To enforce absolute safety in these learning processes, Lyapunov-based Deep Learning Control introduces an end-to-end network in which the weights <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mover><mml:mi>W</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> are strictly updated to ensure that the time derivative of a defined energy function <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> remains negative semi-definite. By structuring the network into tracking-error layers where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>K</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">&#x03BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mi>M</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03BE;</mml:mi></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the robotic arm achieves theoretically guaranteed asymptotic stability (<inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>) even when the true kinematic Jacobian matrix is entirely unknown [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Learning-Augmented Planning</title>
<p>Learning-augmented planning delegates high-level reasoning to Large Language Models (LLMs) while reserving physical execution for lower-level motion planners and RL agents. Vision-Language-Action (VLA) architectures cast the entirety of visual processing, language understanding, and action generation into a unified sequence modelling problem [<xref ref-type="bibr" rid="ref-2">2</xref>]. In standard autoregressive VLAs, the policy is trained via behavioural cloning using a next-token prediction objective: <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x003A;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where the action is discretised into bins and conditioned on visual observations and language instructions <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>l</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. However, purely autoregressive decoding frequently creates severe latency bottlenecks. Modern adaptations like the Optimised Fine-Tuning (OFT) recipe for OpenVLA shift the output from discrete token prediction to L1-regression-based continuous action representations mapped across parallel-decoded action chunks. Quantitatively, abandoning discrete autoregression for parallel decoding and continuous L1 objectives increases OpenVLA&#x2019;s execution speed from 7.34 Hz to a high frequency of 32.14 Hz, reducing inference latency to just 31.1 ms [<xref ref-type="bibr" rid="ref-12">12</xref>]. In healthcare logistics, the Plan-Seq-Learn (PSL) framework perfectly captures this hierarchy by parsing a natural language command into sequential target regions using an LLM, tracking to those regions using collision-free visual motion planning, and executing the final contact-rich manipulation using RL. PSL demonstrates formidable performance, solving highly complex, 10-stage interaction tasks (such as precise NutAssembly) at an exceptional 96% success rate, circumventing the cascading failures typically observed in purely end-to-end models [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Hybrid Stacks</title>
<p>Hybrid stacks fuse the semantic adaptability of Foundation Models with continuous dynamical systems to overcome the limitations of discretised action spaces. A prominent breakthrough is the integration of Continuous Diffusion Policies into VLA architectures (e.g., <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, or Transfusion). In prose, diffusion policies model the action generation process by learning to reverse a stochastic diffusion process. The network starts with pure Gaussian noise and iteratively denoises it into a highly complex, multimodal continuous action trajectory <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>. The training objective minimises a score-matching loss, predicting the injected noise across randomised timesteps. When incorporated into VLAs as &#x201C;Discrete Diffusion&#x201D;, the speed-quality trade-off is remarkably improved; operating a 12-step discrete flow-matching process yields a 14.53 Hz inference speed while maintaining near-perfect task execution rates. In specialised healthcare scenarios, such as the RoboNurse-VLA framework automating surgical instrument handovers, hybrid stacks process voice commands via LLMs and translate them into tokenised bounding boxes and dynamic robotic grasp trajectories in real time, handling unforeseen surgical tools with strong sim-to-real transfer capabilities [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
<sec id="s2_6">
<label>2.6</label>
<title>Comparative Technical Analysis of Advanced ML Paradigms</title>
<p>To move beyond qualitative descriptions, formalise the computational overhead, hardware constraints, and benchmarked metrics of state-of-the-art robotic learning paradigms. <xref ref-type="table" rid="table-2">Table 2</xref> compares different ML paradigms.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>ML paradigm comparison.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>ML Paradigm</th>
<th>Objective/ Architecture</th>
<th>Application Domain (Healthcare Focus)</th>
<th>Computational Hardware Requirements</th>
<th>Quantitative Efficacy Inference Speed</th>
</tr>
</thead>
<tbody>
<tr>
<td>OpenVLA- OFT</td>
<td>Continuous L1 Regression, Parallel Decoding</td>
<td>Dynamic surgical logistics, bimanual object handover</td>
<td>LLaMA2-7B &#x002B; DINOv2, trained on 8x NVIDIA A100 GPUs</td>
<td>32.14 Hz inference (31.1 ms latency); 97.1% average success rate [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
</tr>
<tr>
<td>Plan-Seq-Learn (PSL)</td>
<td>LLM high-level planning &#x002B; Motion Planning &#x002B; PPO</td>
<td>Long-horizon medical staging (up to 10 stages)</td>
<td>16 TPUv3 chips &#x002B; 3000 CPU workers for target Q-values</td>
<td>96.0% success rate on precision contact-rich tasks (e.g., NutAssembly) [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
</tr>
<tr>
<td>Discrete Diffusion VLA</td>
<td>12-step discrete flow-matching continuous trajectory</td>
<td>Dexterous manipulation, liquid/medication handling</td>
<td>4x NVIDIA<break/> A800 GPUs, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>H</mml:mi><mml:mo>=</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula> chunk size</td>
<td>14.53 Hz inference; highly robust to multimodal spatial distributions [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
</tr>
<tr>
<td>DeepONet (Neural Operator)</td>
<td>Operator-to-operator mapping for nonlinear dynamics</td>
<td>Soft robotic tissue modelling, biomechanical fluid flows</td>
<td>Single NVIDIA A6000 GPU</td>
<td>1660 s training time (approx. 77&#x00D7; faster than FNO-3D baseline) [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
</tr>
<tr>
<td>Lyapunov-based DNN</td>
<td>Weight updates constrained by <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula></td>
<td>Safe real-time rehabilitation kinematics</td>
<td>Edge-deployable ML logic controllers</td>
<td>Asymptotic stability tracking guarantees bounded <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> trajectory tracking [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Learning Paradigms for Robot Autonomy</title>
<p>The integration of Deep Learning (DL) into robotic autonomy is fundamentally constrained by an epistemological tension: the inherent &#x201C;black box&#x201D; nature of massive neural networks clashes with the strict, deterministic, and safety-critical requirements of physical robotics. DL models operate as highly non-linear, stochastic function approximators that map high-dimensional state spaces to latent manifolds, rendering their exact decision boundaries mathematically opaque. Conversely, robotic control systems require deterministic, bounded guarantees (such as formal Lyapunov stability) to ensure absolute safety, zero-violation collision avoidance, and precise kinematic execution. An uncertified neural policy may undergo unpredictable extrapolation when encountering out-of-distribution (OOD) biological tissues or dynamic obstacles, risking catastrophic mechanical failure. To resolve this tension without sacrificing the advanced cognitive capabilities of modern learning paradigms, this review proposes a novel conceptual framework: Lyapunov-Bounded Semantic Execution (LBSE). The LBSE framework strictly decouples the robotic autonomy stack into a probabilistic, high-level semantic planner (e.g., Vision-Language-Action architectures) and a deterministic, low-level execution manifold. The high-level model generates rich, multi-step semantic waypoints, which are mathematically filtered through a low-level Control Barrier Function (CBF) and a Lyapunov-certified neural controller. Weight updates in the low-level controller are strictly constrained to ensure the time derivative of the energy function remains negative semi-definite (<inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>), structurally guaranteeing asymptotic stability and bridging the gap between stochastic reasoning and deterministic execution [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<sec id="s3_1">
<label>3.1</label>
<title>Supervised Learning for Perception and Regression Tasks</title>
<p>Supervised learning in robotic perception has evolved from rigid bounding-box regression to dense, task-oriented semantic segmentation and scene coordinate regression. Traditional visual processing utilises Region Proposal Networks (RPNs), where object classification is defined by <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and bounding box regression is defined by <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-22">22</xref>]. However, in highly cluttered robotic environments, these discrete outputs are insufficient for continuous manipulation. Modern systems leverage Efficient-Fully Parameterised Quantile Function (E-FQF) models, applying distributional reinforcement learning to optimise for worst-case occlusion scenarios, thereby drastically reducing collision rates compared to standard mean-value regression [<xref ref-type="bibr" rid="ref-23">23</xref>]. Empirical Analysis: The primary gap in pure supervised learning remains the scarcity of annotated data. By integrating semi-supervised pseudo-labelling with domain-randomised synthetic data, modern hybrid models can bypass manual annotation bottlenecks while achieving precise spatial coordinate mapping [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Self-Supervised and Contrastive Learning</title>
<p>To eliminate the dependency on manual labels, Self-Supervised Learning (SSL) constructs pretext tasks (e.g., masked patch reconstruction) from the data itself. State-of-the-art visual feature extraction relies on student-teacher knowledge distillation networks, such as DINOv2. The architecture computes a cross-entropy loss over local and global crops, heavily utilising the iBOT loss for masked patch modelling: <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>B</mml:mi><mml:mi>O</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the teacher and student prototype scores, respectively. To prevent representational collapse without relying on negative pairs, architectures employ Sinkhorn-Knopp centring and KoLeo regularizers to ensure a uniform feature span within the batch [<xref ref-type="bibr" rid="ref-11">11</xref>]. Empirical Analysis and Non-Obvious Gaps: A critical non-obvious gap is whether SSL genuinely improves continuous control in RL. Recent rigorous empirical studies demonstrate that standard SSL frameworks (e.g., CURL, BYOL, SimSiam) frequently fail to bring meaningful improvements over baselines that solely utilise carefully designed image augmentations in pixel-based RL [<xref ref-type="bibr" rid="ref-24">24</xref>]. However, when utilised strictly for pre-training visual backbones, SSL is highly effective; for instance, DINOv2&#x2019;s optimised discriminative self-supervised training is approximately 2&#x00D7; faster, and requires 3&#x00D7; less memory than previous iterations, producing robust out-of-the-box features that rival weakly supervised models [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Reinforcement Learning</title>
<p>Model-free Reinforcement Learning (RL) maps high-dimensional observations directly to continuous joint torques without explicit dynamic modelling, trading sample efficiency for asymptotic performance and generalizability. Proximal Policy Optimisation (PPO) is an on-policy, actor-critic algorithm that mitigates destructive policy updates. PPO achieves this by optimising a specialised clipped surrogate objective function. In prose, the algorithm computes an advantage estimate <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and a probability ratio <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> between the updated policy and the behavioural policy. The objective is mathematically bounded via <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-25">25</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]. This clipping restricts the policy step size, thereby guaranteeing stable convergence for complex manoeuvres such as bipedal locomotion. Soft Actor-Critic (SAC) is an off-policy algorithm built on a maximum entropy framework, designed to tackle sparse-reward environments. SAC optimises a stochastic policy by maximising both the expected cumulative reward and the policy&#x2019;s entropy, as defined by the objective <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2211;</mml:mo><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. The temperature parameter governs the exploration-exploitation trade-off, encouraging the agent to explore diverse kinematic trajectories. Empirical Analysis: While SAC is theoretically more sample-efficient due to its off-policy experience replay, it introduces severe computational bottlenecks during gradient updates. A rigorous comparison on a physical robotic grasping task revealed that PPO achieved a substantially higher mean reward (approx. 1200) than SAC (approx. 800) within 2.5 h of wall-clock time. SAC&#x2019;s off-policy updates yielded lower per-iteration rewards and exhibited significant non-monotonic variance due to the heavy computational overhead of processing large replay buffers [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Imitation Learning and Learning from Demonstration</title>
<p>Imitation Learning (IL) bypasses the extensive exploration phase of RL by bootstrapping policies directly from human teleoperation or expert demonstrations. Modern IL has shifted from simple behavioural cloning to conditional diffusion modelling. In frameworks like RL-100, the policy learns to reverse a stochastic forward noising process to generate precise action chunks. The denoiser <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is trained via the noise-prediction objective: <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a conditioning vector that fuses visual and proprioceptive histories. Empirical Analysis: Pure IL suffers from distribution shifts when the robot encounters states outside the demonstration manifold. To combat this, the RL-100 framework utilises Consistency-Model distillation to compress the multi-step diffusion process into a single-step action generator, achieving high-frequency control while preventing catastrophic forgetting during subsequent RL fine-tuning [<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Offline Reinforcement Learning</title>
<p>Offline RL extracts optimal control policies from static, previously logged datasets without any active environmental interaction, making it vital for safety-critical systems where online exploration is dangerous. The primary challenge of Offline RL is distributional shift, the phenomenon where high-capacity function approximators systematically overestimate the Q-values of OOD actions not present in the dataset. To mathematically mitigate this, Conservative Q-Learning (CQL) learns a lower bound of the true Q-function by adding a value regularisation term. The CQL regularizer is formulated as <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03A6;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>&#x03A6;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>&#x03A6;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. This objective explicitly pushes down the Q-values of unseen, OOD actions while pushing up the values of actions present in the dataset [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. Alternatively, the Trajectory Transformer reframes Offline RL as a sequence modelling problem, training a single high-capacity transformer to represent the joint distribution over states, actions, and rewards, avoiding explicit pessimism entirely by utilising beam search over transition log-probabilities [<xref ref-type="bibr" rid="ref-32">32</xref>]. Empirical Analysis: Benchmarking in physical robotics (e.g., TriFinger manipulation) reveals that CQL frequently encounters optimisation instabilities (saddle-point problems) and requires extensive hyperparameter grid searches (over 405 configurations for simple Push tasks) to achieve viable success rates [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Multi-Task, Meta-Learning, and Continual Learning</title>
<p>Robots deployed in unstructured environments must continually acquire new skills without suffering from catastrophic forgetting of previously learned tasks. Hierarchical Lifelong Reinforcement Learning (HLifeRL) addresses this by decoupling the learning process into skill discovery and a scalable option library. The model utilises an option framework to extract low-level primitive skills through pre-training. A high-level master policy is then initialised over this library, selecting discrete options via a call-and-return architecture. Empirical Analysis: By freezing learned options and expanding the library sequentially, HLifeRL demonstrably prevents the catastrophic interference typically observed when traditional deep neural networks are forced to map overlapping task distributions within a shared parameter space [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>ML Paradigms Trade-Off Comparison</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> shows that the choice of machine learning paradigm involves a direct trade-off between hardware deployment capability and quantitative efficacy.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance metrics, computational overhead, and hardware requirements.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Learning Paradigm</th>
<th>Representative Architecture</th>
<th>Action Space Decoding</th>
<th>Primary Hardware Deployment</th>
<th>Quantitative Efficacy &#x0026; Overhead</th>
</tr>
</thead>
<tbody>
<tr>
<td>Discrete Diffusion VLA</td>
<td>Unified Transformer (ViT &#x002B; LLaMA)</td>
<td>Discretised Action Chunks (Parallel Decoding)</td>
<td>4&#x00D7; NVIDIA A800 GPUs (Training)</td>
<td>96.3% avg. Success Rate on LIBERO: replaces Left-to-Right AR bottlenecks with dynamic re-masking [<xref ref-type="bibr" rid="ref-21">21</xref>].</td>
</tr>
<tr>
<td>On-Policy RL (PPO)</td>
<td>Actor-Critic with Clipped Surrogate</td>
<td>Continuous Torques/ Velocities</td>
<td>CPU/GPU Clusters (Simulation-heavy)</td>
<td>Reaches 1200 mean reward in 2.5 h wall-clock time; highly efficient compute per iteration despite lower sample efficiency [<xref ref-type="bibr" rid="ref-26">26</xref>].</td>
</tr>
<tr>
<td>Off-Policy RL (SAC)</td>
<td>Maximum Entropy Actor-Critic</td>
<td>Continuous Torques/ Velocities</td>
<td>CPU/GPU Clusters (Replay Buffer overhead)</td>
<td>Reaches 800 mean reward in 2.5 h wall-clock time; high variance and high computational overhead per gradient step [<xref ref-type="bibr" rid="ref-26">26</xref>].</td>
</tr>
<tr>
<td>Self-Supervised Learning</td>
<td>DINOv2 (ViT-g)</td>
<td>Visual Feature Extraction (Latent)</td>
<td>Meta-scale cluster (Pre-training)</td>
<td>2<inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> faster training, 3<inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> less memory than iBOT; closes the gap with weakly-supervised feature extractors [<xref ref-type="bibr" rid="ref-11">11</xref>].</td>
</tr>
<tr>
<td>Imitation Learning (Diffusion)</td>
<td>RL-100 (Conditional Diffusion)</td>
<td>Single-step/Chunked (Single-step via Distillation)</td>
<td>Edge-deployable (via Consistency Models)</td>
<td>Mitigates excessive exploration by clipping DDIM noise bounds (e.g., 0.8 clip yields optimal stability/performance trade-off) [<xref ref-type="bibr" rid="ref-29">29</xref>].</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Core Algorithmic Building Blocks</title>
<p>Four themes underpin most practical robot learning systems: geometric representations, planning with learned components, skill learning and hierarchical control, and world models for predictive planning.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Representations: Geometry Meets Learning</title>
<p>Furthermore, to mitigate the curse of dimensionality in complex motion planning, the Latent Sampling-Based Motion Planning (L-SBMP) algorithm learns a planable, low-dimensional manifold from high-dimensional workspaces. The architecture comprises an encoder <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>Z</mml:mtext></mml:mrow></mml:math></inline-formula>, a decoder <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mrow><mml:mtext>n</mml:mtext></mml:mrow><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>Z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow></mml:math></inline-formula>, and a latent dynamics model <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>Z</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, accompanied by a latent collision checker <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>Z</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. By projecting the search space into <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>Z</mml:mi></mml:math></inline-formula>, the robot bypasses expensive geometric collision computations. Statistically, utilising deep neural network estimators for spatial occupancy and swept volumes accelerates inference by 3500 to 5000 times compared to exact geometric swept-volume computations, thereby decisively removing real-time perception bottlenecks [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Planning with Learned Heuristics and Costs</title>
<p>Classical path planning algorithms (e.g., A&#x002A;, RRT) rely on manually engineered heuristic functions (such as Euclidean distance), which notoriously fail in non-convex, high-dimensional spaces, forcing the algorithm to perform exhaustive, computationally prohibitive expansions. Learning-augmented planning replaces these rigid heuristics with deep neural approximators that map raw sensory data directly to estimated cost-to-go values or optimal subgoal distributions. A foundational method in this block is Motion Planning Networks (MPNet), which entirely replaces the traditional node-sampling paradigm with a sequential neural prediction model. The MPNet architecture utilises a contractive autoencoder to embed the obstacle point cloud into a latent space, optimised via the reconstruction loss <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:munder><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. A feed-forward planning network parameterised by then recursively predicts the next configuration, trained to minimise the mean-square-error over expert trajectory sequences loss <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>c</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. Quantitatively, modifying heuristic expansions with learning-based predictors reduces planning-iteration processing times by up to 95% relative to standard A&#x002A; configurations, while MPNet maintains computational speeds significantly lower than those of state-of-the-art classical sampling methods, even in unseen environments [<xref ref-type="bibr" rid="ref-34">34</xref>]. The loss functions <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent a two-stage learning process for augmented planning. The autoencoder objective <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> utilizes a contractive regularization term <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> to ensure the latent manifold is robust to small perturbations in the obstacle geometry. The planning loss <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> follows a supervised learning paradigm, where the network is trained to minimize the distance between the predicted next-step configuration <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and the expert ground truth. We have clarified the summation notation to explicitly state that the loss is averaged across <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> expert trajectories to improve readability.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Skill Learning and Hierarchical Control</title>
<p>End-to-end Deep Reinforcement Learning (DRL) catastrophically degrades over long-horizon tasks due to the exponential growth of the state-action exploration space and the sparsity of reward signals. To mitigate this, robotic control leverages Hierarchical Control and Skill Learning, wherein high-level planners operate on a discrete action space of temporally extended, abstract &#x201C;skills&#x201D;, while low-level controllers execute continuous motor commands. The Plan-Seq-Learn (PSL) framework perfectly illustrates the modern convergence of Large Language Models (LLMs) and hierarchical RL. Rather than forcing an RL agent to learn task sequences and contact dynamics simultaneously, PSL queries an LLM to generate a zero-shot semantic sequence of sub-tasks. A vision-based motion planner sequences the robot&#x2019;s end-effector to the initialisation region of each sub-task, and a localised RL policy is solely responsible for learning the contact-rich manipulation. Empirically, PSL achieves a 96.0% success rate on the complex, 10-stage &#x201C;NutAssembly&#x201D; task directly from raw visual inputs, severely outperforming end-to-end baselines, which completely fail to make progress due to cascading estimation errors [<xref ref-type="bibr" rid="ref-20">20</xref>]. Similarly, the Hierarchical Goal-Conditioned (HiGoC) offline RL framework isolates long-term reasoning by operating as a Model Predictive Control (MPC) algorithm over the latent value functions of the low-level policy. By sampling continuous sub-goals over a look-ahead horizon, HiGoC mathematically bounds exploration risks. Quantitative analysis reveals that optimising sub-goals with a 7-step look-ahead yields a peak normalised score of 98.4 on expert datasets, drastically outperforming non-hierarchical Conservative Q-Learning (CQL) baselines [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>World Models and Predictive Control</title>
<p>Relying entirely on model-free DRL is prohibitively sample-inefficient for physical robots. World Models alleviate this by learning an explicit, differentiable model of the environment&#x2019;s transition dynamics entirely in a latent space, enabling the agent to simulate thousands of trajectories (mental rehearsals) without physical interaction. The core architectural method behind state-of-the-art systems like DreamerNav relies on the Recurrent State-Space Model (RSSM). The RSSM mathematically unifies deterministic memory and stochastic transitions. A sequence of high-dimensional observations is compressed by an encoder <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The dynamics network predicts the stochastic future <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, while a recurrent update function modifies the deterministic hidden state <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The policy is then optimised entirely through backpropagation through time over these imagined latent rollouts to maximise the &#x03BB;-returns of the predicted values [<xref ref-type="bibr" rid="ref-13">13</xref>]. When deployed for predictive control, learning-based dynamics vastly outpace numerical solvers. In the Deep Value-and-Predictive-Model Control (DVPMC) architecture, an artificial neural network approximates the forward-time dynamics <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mrow><mml:mover><mml:mi>f</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-36">36</xref>]. For a 3-DOF bipedal robot leg executing Model Predictive Control (MPC), deploying a Deep Neural Network (DNN) surrogate dynamics model slashed the real-time prediction latency to mere <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mn>0.01</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:math></inline-formula> s per sample. In stark contrast, solving the exact mathematical statics model via the standard &#x2018;ode45&#x2019; numerical integrator required <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mn>0.89</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.018</mml:mn></mml:math></inline-formula> s per sample, demonstrating an 89-fold speedup, which is crucial for real-time 50 Hz control loops [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Interplay and Trade-Offs among Building Blocks</title>
<p>While current frameworks successfully integrate world models with hierarchical task planners, there is a distinct lack of bidirectional causal feedback between the layers. If a low-level policy utilising DVPMC encounters unmodeled tissue compliance during a surgical task, it currently cannot mathematically communicate this physical failure back to the LLM to dynamically update the semantic skill sequence. The proposed SPCA framework resolves this by forcing the low-level Lyapunov-certified controller to output an explicit &#x201C;safety-bound violation&#x201D; flag to the semantic planner, dynamically triggering a re-routing of the high-level heuristic graph before catastrophic failure occurs. <xref ref-type="table" rid="table-4">Table 4</xref> explains different algorithmic building blocks in terms of architecture, decoding mechanism, computational overhead and quantitative efficacy.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance metrics, computational overhead, and hardware requirements of Core ML paradigms.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Algorithmic Building Block</th>
<th>Representative Architecture</th>
<th>Formal Target Decoding Mechanism</th>
<th>Hardware &#x0026; Computational Overhead</th>
<th>Quantitative Efficacy Result</th>
</tr>
</thead>
<tbody>
<tr>
<td>Learned Heuristics/Planning</td>
<td>MPNet</td>
<td>Contractive Autoencoder &#x002B; L2 Path Regression</td>
<td>Real-time edge deployment; CPU or lightweight GPU.</td>
<td>Reduces planning iteration processing times by 95%; consistently faster than classical RRT/A&#x002A; [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>].</td>
</tr>
<tr>
<td>Geometry/ Equivariant Reps.</td>
<td>EGNN/SO(3) Equivariance</td>
<td><inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mo>&#x2295;</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> Tensor Product for Rotational Symmetry</td>
<td>Moderate; requires tracking Clebsch-Gordan coefficients.</td>
<td>Replaces exact swept volume computations with inference speeds 3500 to 5000 times faster [<xref ref-type="bibr" rid="ref-34">34</xref>].</td>
</tr>
<tr>
<td>Hierarchical Control</td>
<td>Plan-Seq-Learn (PSL)</td>
<td>LLM Planner &#x002B; Motion Planning &#x002B; localised RL</td>
<td>16 TPUv3 chips &#x002B; 3000 CPU workers<break/> (RL phase).</td>
<td>Achieves 96.0% success on 10-stage contact-rich NutAssembly tasks; entirely mitigates cascading errors [<xref ref-type="bibr" rid="ref-20">20</xref>].</td>
</tr>
<tr>
<td>World Models &#x0026; Dynamics</td>
<td>DreamerNav (RSSM)</td>
<td>Latent sequence prediction: <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>NVIDIA<break/> A100 GPUs (Simulation); Jetson Orin (Deployment).</td>
<td>Bipedal deep MPC executes at 0.01 s latency vs. 0.89 s for numerical ODE baselines [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>].</td>
</tr>
<tr>
<td>Goal-Conditioned Offline RL</td>
<td>HiGoC (7-step look-ahead)</td>
<td>High-level MPC optimisation over value networks</td>
<td>Offline static datasets; standard GPU training.</td>
<td>Achieves 98.4 normalised score on expert dataset bounds; highly robust to OOD distribution shifts [<xref ref-type="bibr" rid="ref-35">35</xref>].</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Healthcare Robots</title>
<p>The manuscript provides certain explicit and rigorous connections of healthcare applications with specific core algorithms to meet the specific needs of clinical settings. The connections are structured across perception, dynamics and control layers to ensure that &#x201C;black-box&#x201D; AI meets the deterministic safety requirements of medicine. Surgical Precision and Limitations: The review shows the use of Gaussian Mixture Model-based Dynamic Movement Primitives (GMM-DMPs) with Dynamic Time Warping (DTW) for controlling the Remote Centre of Motion (RCM). This algorithmic stack enables robots like the da Vinci Research Kit to follow strict kinematic constraints in laparoscopy, which are challenging to satisfy with models based exclusively on imitation. Rehabilitation Kinematics: To ensure the safety of human-robot interaction in physical therapy, the text links Lyapunov-based Deep Learning Control to a 3-DOF bipedal rehabilitation leg. This ensures asymptotic stability, meaning the robot tracking error is mathematically guaranteed to approach zero even when the exact physics of the patient&#x2019;s limb is unknown. Diagnostic Robustness: To handle &#x201C;out-of-distribution&#x201D; (OOD) biological anomalies, such as rare tissues or staining variations in histopathology, the review connects Evidential Deep Learning and Conditional Diffusion Models to clinical datasets. These algorithms quantify epistemic uncertainty, allowing the system to flag unknown pathologies rather than providing a false confident diagnosis. Clinical Logistics and Instruction Following: The RoboNurse-VLA and Plan-Seq-Learn (PSL) frameworks are linked to surgical instrument handovers. These use Large Language Models (LLMs) for high-level semantic reasoning (e.g., understanding a &#x201C;thirsty&#x201D; patient or a specific surgical tool request) while delegating the final, contact-rich movement to Reinforcement Learning (RL) policies. The integration of core algorithms with healthcare is necessary, as medical environments are unstructured and high-stakes applications are indicated in the manuscript. <xref ref-type="table" rid="table-5">Table 5</xref> shows the health care application with its suitable algorithmic solution and the technical efficacy.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Summary of technical mapping.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Healthcare Focus</th>
<th>Core Algorithmic Solution</th>
<th>Technical Efficacy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Surgical Handover</td>
<td>OpenVLA-OFT (Parallel Decoding)</td>
<td>Reduces latency to 31.1 ms for real-time responsiveness.</td>
</tr>
<tr>
<td>Tissue Modelling</td>
<td>DeepONet (Neural Operators)</td>
<td>77&#x00D7; faster training than standard baselines for predicting nonlinear dynamics.</td>
</tr>
<tr>
<td>ldercare Navigation</td>
<td>DreamerNav (RSSM World Models)</td>
<td>Encodes environmental dynamics to enable autonomous navigation in indoor spaces.</td>
</tr>
<tr>
<td>Histopathology</td>
<td>Conditional Diffusion Models</td>
<td>Minimizes demographic fairness gaps and improves OOD accuracy.</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5_1">
<label>5.1</label>
<title>Advanced Algorithmic Paradigms in Robotic Manipulation</title>
<p>To navigate highly unstructured environments, modern robotic learning frameworks leverage biologically inspired predictive coding, structured imitation learning (IL), and dynamic trajectory adaptations rather than relying purely on massive, unconstrained datasets. World Models and Predictive Coding: Contemporary cognitive robotics relies on World Models to efficiently encode the environment&#x2019;s spatiotemporal dynamics, enabling sample-efficient model-based planning. Grounded in the Free-Energy Principle (FEP) and Active Inference, these systems continuously generate top-down predictions and utilise bottom-up sensory prediction errors to update their internal states. This framework mathematically unifies perception and action, where action is formulated as active sensory sampling designed to minimise variational free energy [<xref ref-type="bibr" rid="ref-37">37</xref>]. Dynamic Movement Primitives (DMPs) with Adaptive Control: DMPs model complex robotic trajectories using nonlinear dynamical systems, ensuring global stability and smooth transitions without rigid time-indexing. To handle uncertainties in robot dynamics, modern DMP frameworks integrate adaptive Neural Network (NN) controllers to compensate for approximation errors. Overlapping kernels along the time axis enable multi-stage movement sequences, drastically reducing the velocity attenuation (pauses) traditionally observed at junctions between separate movement primitives [<xref ref-type="bibr" rid="ref-38">38</xref>]. Human-in-the-Loop (HITL) Frameworks: Acknowledging the sample inefficiency of pure Deep Reinforcement Learning (DRL), HITL frameworks position humans as operators, collaborators, or supervisors. This allows algorithms to leverage human cognitive priors. For example, spatial iterative learning control (sILC) driven by online human corrections minimises environmental uncertainties during trajectory execution [<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Transformative Breakthroughs in Healthcare Robotics</title>
<p>The integration of advanced robotic paradigms has fundamentally altered the healthcare sector, particularly in precision surgical operations and rehabilitative care, moving beyond theoretical models to demonstrably improve clinical execution. Robot-Assisted Minimally Invasive Surgery (MIS): Surgical applications have seen transformative breakthroughs by bridging human surgical expertise with robotic precision. Frameworks utilising Dynamic Time Warping (DTW) combined with Gaussian Mixture Model-based DMPs (GMM-DMPs) have successfully modelled complex surgical manipulation skills on platforms like the KUKA LWR4&#x002B; and the da Vinci Research Kit. These algorithms effectively manage strict kinematic constraints, such as the Remote Centre of Motion (RCM) requisite in laparoscopy, thereby augmenting surgical safety and precision [<xref ref-type="bibr" rid="ref-39">39</xref>]. Active Inference in Medical Applications: Active inference controllers have been successfully deployed on robotic manipulators for fault-tolerant control and advanced body perception [<xref ref-type="bibr" rid="ref-37">37</xref>]. By minimising expected free energy, these models dynamically adapt to perturbations in real time during physical human-robot interaction, offering robust solutions for surgical robotic simulators (e.g., SurRoL) and minimising the sim-to-real gap during continuous skill acquisition [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>The Neural Meta Evaluator (NeME)</title>
<p>IL methods in Human-Robot Interaction and Collaboration (HRIC) have been evaluated using Average Success Rate (SR) or Dynamic Time Warping (DTW). However, SR requires time-consuming, resource-intensive deployment on physical robots and is highly susceptible to human variability. DTW often fails to capture the nuanced quality of robot motion, heavily penalising valid but slightly divergent trajectories. To provide a rigorous, reproducible evaluation pipeline, the Neural Meta Evaluator (NeME) frames policy assessment as a sequence-classification task based directly on robot joint trajectories. NeME operates as an offline meta-evaluator, efficiently processing generated trajectories without the constraints of human-in-the-loop deployment. Empirical evaluations demonstrate that optimal model weights selected via NeME&#x2019;s meta-F1 (mF1) scores perfectly align with the actual peak SR (e.g., precisely identifying the 8th-epoch peak where validation loss fails), thereby providing a statistically rigorous surrogate for physical deployment [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Comparative Analysis of Performance and Computational Overhead</title>
<p>The integration of Artificial Intelligence (AI) and robotic systems into healthcare ecosystems represents a fundamental paradigm shift from traditional medical models to predictive, personalised, preventive, participatory, and precision (P5) medicine [<xref ref-type="bibr" rid="ref-41">41</xref>]. Advanced machine learning (ML) frameworks, specifically Deep Reinforcement Learning (DRL) and Generative AI, are driving transformative breakthroughs in surgical precision, adaptive physical rehabilitation, and equitable diagnostic modelling [<xref ref-type="bibr" rid="ref-42">42</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>]. The following tables summarise the performance and Computational Overhead. The first table shows the Comparative Evaluation of Sequence Modelling Architectures for Meta-Evaluation (NeME) Evaluation of behaviour classification performance given an input trajectory length (L &#x003D; 32, equating to a 3.2-s window). The LSM architecture demonstrates superior representational power for robotic joint-state sequences compared to modern state-space models [<xref ref-type="bibr" rid="ref-40">40</xref>]. <xref ref-type="table" rid="table-6">Table 6</xref> summarises the trade-off between choosing various robotic learning paradigms and <xref ref-type="table" rid="table-7">Table 7</xref> shows the algorithmic efficacy, computational overhead, and Hardware deployment in different types of health care robots.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparative evaluation of sequence modelling architectures.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Neural Architecture</th>
<th>Meta-Accuracy (%)</th>
<th>Meta-F1 Score (%)</th>
<th>Training Hyperparameters &#x0026; Augmentation</th>
</tr>
</thead>
<tbody>
<tr>
<td>LSTM</td>
<td>72.1 &#x00B1; 1.3</td>
<td>66.6 &#x00B1; 0.3</td>
<td><italic>h</italic> &#x2208; {16, 64, 128}, Gaussian noise<break/> <italic>&#x03C3;</italic> &#x2208; {0, <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>}</td>
</tr>
<tr>
<td>Transformer (Encoder-only)</td>
<td>68.7 &#x00B1; 1.2</td>
<td>61.6 &#x00B1; 3.1</td>
<td>AdamW optimiser, optimised for 30 epochs</td>
</tr>
<tr>
<td>xLSTM</td>
<td>59.1 &#x00B1; 3.6</td>
<td>53.3 &#x00B1; 3.7</td>
<td>Early stopping at 5 epochs<break/> (no validation loss improvement)</td>
</tr>
<tr>
<td>Mamba</td>
<td>53.6 &#x00B1; 13.8</td>
<td>51.9 &#x00B1; 12.7</td>
<td>Projection layer from <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x2208; <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>24</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> to<break/> hidden size h</td>
</tr>
<tr>
<td>Mamba2</td>
<td>58.1 &#x00B1; 5.4</td>
<td>54.0 &#x00B1; 4.9</td>
<td>Averaged over three distinct weight initialisations</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Algorithmic efficacy, computational overhead, and hardware deployment.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Paradigm/ Algorithm</th>
<th>Target Application</th>
<th>Hardware Deployment</th>
<th>Computational Overhead Specifications</th>
<th>Baseline Comparison Metric</th>
</tr>
</thead>
<tbody>
<tr>
<td>eXteroceptive Behaviour Generation (XBG)</td>
<td>HRIC autonomous interaction</td>
<td>ergoCub Humanoid, Intel Realsense D450</td>
<td>Training: 12 epochs, batch size 64, 2&#x00D7; NVIDIA A100 GPUs, Learning rate: 5e&#x2212;4,<break/> 10 Hz inference.</td>
<td>XBG (RGB-D): Meta-F1: 71.3%, DTW: 2.18, SR: 70.0%, XBG-RGB: Meta-F1: 69.9%, DTW: 2.40, SR: 61.7% [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>Adaptive NN &#x002B; DMPs</td>
<td>Continuous complex manipulation</td>
<td>Baxter Robot</td>
<td>Calculates target trajectory velocity <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and position dynamically using high-priority spring-damper mechanisms.</td>
<td>Substantial reduction in velocity attenuation at junction points compared to standard end-to-end DMP sequences [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
<tr>
<td>DTW &#x002B; GMM-DMP</td>
<td>Minimally Invasive Surgery (MIS)</td>
<td>KUKA LWR4&#x002B;, da Vinci Research Kit</td>
<td>Translates kinesthetic teaching into dynamically warped geometric models.</td>
<td>Circumvents Remote Centre of Motion (RCM) constraint failures present in pure imitation models [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Advanced Algorithmic Interventions and Model Robustness</title>
<p>The clinical effectiveness of medical AI depends on high-quality data inputs and on the ability of models to perform across diverse real-world environments [<xref ref-type="bibr" rid="ref-45">45</xref>]. Variations in clinical hardware and procedures, such as disparate histological staining techniques across hospitals, often cause diagnostic models to underperform on out-of-distribution (OOD) data. Generative AI, specifically conditional diffusion models, directly addresses this by synthesising realistic medical imagery to compensate for underrepresented demographic subgroups and rare pathologies. Empirical evidence from the CAMELYON17 histopathology challenge demonstrated that training diffusion models on 455,954 labelled patches and 1.8 million unlabeled patches significantly minimised demographic fairness gaps and preserved high diagnostic accuracy under severe OOD conditions [<xref ref-type="bibr" rid="ref-42">42</xref>]. Deep Reinforcement Learning (DRL) in Dynamic Environments: DRL provides optimised, goal-oriented autonomy for managing unstructured clinical environments and complex biological data. In object manipulation, DRL-driven Viewpoint Adjusting and Grasping Synergy (VAGS) strategies have achieved an 83.50% grasp success rate and a 95% scene-clearing rate in highly cluttered simulations. In targeted biotechnology applications, Hierarchical Deep Reinforcement Learning (HDRL) models have successfully processed massive 3D time-lapse image sets to navigate <italic>C. elegans</italic> embryogenesis, map modular cellular movement pathways, and identify novel therapeutic targets [<xref ref-type="bibr" rid="ref-43">43</xref>].</p>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Regulatory Hurdles and the Imperative for Dependable Deployment</title>
<p>True clinical integration is hindered not merely by technological limitations but by the rigorous demands of ethical governance, data security, and legal accountability [<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>]. Stringent Regulatory Compliance: The reliable deployment of connected health robots is tightly governed by frameworks such as GDPR, which mandates explicit patient consent and comprehensive encryption for cross-border transmission of medical data [<xref ref-type="bibr" rid="ref-47">47</xref>]. In the United States, the FDA had cleared 222 AI-based medical devices by 2020; however, algorithms capable of continuous post-market learning pose an acute regulatory challenge, necessitating the development of novel oversight mechanisms to ensure ongoing safety [<xref ref-type="bibr" rid="ref-46">46</xref>]. Responsibility and Liability Attribution: The deployment of highly autonomous surgical and mobile robots significantly complicates legal accountability in the event of adverse events [<xref ref-type="bibr" rid="ref-46">46</xref>]. Retrospective analyses of FDA data over a 14-year period emphasise the genuine physical risks associated with robotic interventions [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>]. Algorithmic Opacity and Explainable AI (XAI): The inherent &#x201C;black box&#x201D; nature of deep neural networks obscures the logic driving clinical predictions, fundamentally undermining physician trust and patient safety. Interpretability is transitioning from an operational preference to a strict legal requirement under frameworks such as the EU Artificial Intelligence Act. For dependable deployment, developers must mandate XAI frameworks that transparently justify automated decisions to prevent automation bias and the entrenchment of existing health disparities [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. Because medical AI lacks independent moral status, human operators and institutional stakeholders remain the primary duty-bearers; however, automated systems that execute high-risk manoeuvres in sub-second timeframes effectively preclude human intervention, creating unresolved legal ambiguity regarding liability [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Emerging Trends</title>
<sec id="s6_1">
<label>6.1</label>
<title>Foundation Models for Robotics</title>
<p>The robotic field has definitively transitioned from fragmented, task-specific deep reinforcement learning (DRL) toward internet-scale Embodied Foundation Models. Driven by massive aggregation efforts such as the Open X-Embodiment (OXE) dataset, these architectures establish universal control priors that enable zero-shot transfer across diverse robot morphologies. State-of-the-art models like RDT-1B and Octo utilise diffusion-based policy modelling to handle highly multimodal visuomotor distributions. Rather than outputting a deterministic action, the diffusion foundation model learns to reverse a stochastic forward process. The training objective minimises a score-matching loss over action chunks: <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> where the generation is conditioned on language and visual observations. Quantitative Breakthroughs (2024&#x2013;2025): The shift to diffusion foundation models yields unprecedented generalisation. Octo, trained on over 4 million trajectories across 22 distinct robotic platforms, achieves highly robust cross-embodiment sim-to-real transfer. Similarly, RDT-1B (a 1.2B-parameter diffusion foundation model) demonstrates exceptional zero-shot generalisation in complex bimanual manipulation scenarios, resolving the sparse-reward bottleneck that traditionally paralysed DRL [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Large Language Models as Task Interfaces and Planners</title>
<p>Large Language Models (LLMs) are now utilised to bypass algorithmic abstraction barriers, acting as zero-shot semantic planners that translate human intent directly into logically structured sub-goals [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]. Seminal frameworks like SayCan and ProgPrompt formulate planning as a constrained probability maximisation problem. SayCan mathematically grounds LLM abstractions into physical affordances via: <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, multiplying the LLM&#x2019;s semantic prior by a learned visual affordance score [<xref ref-type="bibr" rid="ref-51">51</xref>]. ProgPrompt transforms situated environments into Pythonic APIs. The generation relies on a prompting function <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> that maps the state <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>s</mml:mi></mml:math></inline-formula> into an import-style header (e.g., from actions import grab, open). The LLM recursively predicts the next program string, integrating assert conditions to provide real-time state feedback and precondition checking. Quantitative Evaluation: LLM-generated Pythonic execution heavily mitigates the &#x201C;cascading error&#x201D; problem seen in sequential robotic tasks. By explicitly grounding preconditions, LLM-guided planners can drastically improve success rates in long-horizon task generation without requiring any domain-specific policy re-training [<xref ref-type="bibr" rid="ref-52">52</xref>].</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Language-Conditioned and Multimodal Policies</title>
<p>Vision-Language-Action (VLA) architectures fuse pre-trained vision encoders (e.g., SigLIP, DINOv2) and LLMs (e.g., LLaMA, Qwen) directly into action decoders. This end-to-end multimodal alignment achieves profound instruction-following capabilities [<xref ref-type="bibr" rid="ref-53">53</xref>]. Standard VLAs historically relied on left-to-right autoregressive decoding, treating continuous joint torques as discretised text tokens: <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x003A;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-12">12</xref>]. However, this severely bottlenecks inference speeds. The 2024 breakthrough OpenVLA-OFT replaces autoregression with Parallel Decoding and Action Chunking, optimising a continuous L1 regression loss directly on the output representations. Conversely, the Discrete Diffusion VLA applies flow-matching directly to tokenised action chunks via a transition matrix <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Quantitative Breakthroughs: Addressing the severe latency constraints of embodied control, the OpenVLA-OFT&#x002B; variant processes continuous actions via parallel decoding, achieving a real-time action-generation throughput of 77.9 Hz (31.1 ms latency) on an NVIDIA A100 GPU, outperforming the original OpenVLA&#x2019;s sluggish 1.8 Hz baseline [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Sim-to-Real Transfer at Scale</title>
<p>The Reality Gap (<inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) is being actively closed by highly parallelised, GPU-accelerated simulators (e.g., Isaac Sim, MuJoCo) executing massive Domain Randomisation (DR) and adversarial feature learning [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. Rather than attempting to model physical reality perfectly, DR perturbs the simulated transition dynamics <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0B5;</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> where physics parameters &#x00B5; (mass, friction, damping) are sampled from a broad distribution <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>&#x03BC;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. To ensure representation robustness, contrastive learning objectives enforce that the latent representation <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> remains invariant to rendering variations, mathematically binding the features to task-relevant physics rather than superficial textures [<xref ref-type="bibr" rid="ref-6">6</xref>]. Quantitative Evaluation: Modern massively parallelised RL algorithms can simulate thousands of environments simultaneously. In highly agile quadruped locomotion (e.g., ANYmal parkour), policies trained purely under heavy domain randomisation in simulation successfully execute highly dynamic blind obstacle traversal in the real world without any real-world fine-tuning [<xref ref-type="bibr" rid="ref-54">54</xref>].</p>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Safety, Verification, and Trustworthy Autonomy</title>
<p>Deploying neural policies in critical environments requires migrating from empirical &#x201C;success rates&#x201D; to formal control-theoretic safety certificates. Control Barrier Functions (CBFs) and Lyapunov Certification provide this mathematical rigour. For a control-affine system <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>u</mml:mi></mml:math></inline-formula>, a neural policy is strictly bounded by a continuously differentiable safe set defined by <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. The learning controller ensures safety by solving a quadratic program that strictly enforces the CBF derivative condition: <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>u</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Furthermore, deep networks parameterising the policy are regularised via spectral normalisation to strictly bound their Lipschitz constant, ensuring that bounded uncertainties in the physical dynamics translate directly to bounded, stable errors in the controller&#x2019;s time derivatives. Quantitative Evaluation: Architectures employing Lipschitz-bounded verification and CBFs mathematically eliminate hardware constraint violations during learning (achieving zero-collision exploration), which is critical when robotic hardware costs hundreds of thousands of dollars [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
</sec>
<sec id="s6_6">
<label>6.6</label>
<title>Hybrid Learning Stack for Robotics</title>
<p>Recognising the sample inefficiency of pure DRL and the latency of VLAs, hybrid learning stacks orchestrate a symbiotic pipeline: slow, high-level semantic foundation models orchestrating fast, low-level continuous dynamical systems (RL or classical MPC) [<xref ref-type="bibr" rid="ref-55">55</xref>]. Dual-system architectures decouple control frequency. A &#x201C;System 2&#x201D; LLM planner generates semantic spatial targets at <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mn>1</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mo>&#x223C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>3</mml:mn></mml:math></inline-formula> Hz, resolving the long-horizon sparse reward problem. A &#x201C;System 1&#x201D; low-level DDPG or PPO reinforcement learning agent operates at <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mn>50</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mo>&#x223C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>500</mml:mn></mml:math></inline-formula> Hz, processing real-time proprioceptive force feedback to minimise the residual error strictly within the localised semantic bounds established by System 2 [<xref ref-type="bibr" rid="ref-56">56</xref>]. Non-Obvious Gap: Current hybrid models lack bidirectional causal feedback. If the high-frequency RL agent encounters unmodeled object compliance (e.g., slipping), it cannot mathematically communicate the geometric reason for this failure back into the language embedding space of the LLM for re-routing. Future research must integrate visual-language failure evaluators (e.g., StepEval) to encode physical failures as text tokens, thereby creating a closed-loop, symbiotic pipeline [<xref ref-type="bibr" rid="ref-57">57</xref>]. <xref ref-type="table" rid="table-8">Table 8</xref> compares different model paradigms in terms of certain fields.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Comparative technical analysis of multimodal ML paradigms (2024&#x2013;2025).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model Paradigm</th>
<th>Vision/Language Backbone</th>
<th>Action Decoding Framework</th>
<th>Hardware Deployment</th>
<th>Quantitative Efficacy/Inference Speed</th>
</tr>
</thead>
<tbody>
<tr>
<td>OpenVLA-OFT&#x002B;</td>
<td>SigLIP &#x002B; DINOv2/ LLaMA-2 7B</td>
<td>Continuous L1 Regression, Parallel Decoding</td>
<td>NVIDIA A100 GPU (optimised via LoRA)</td>
<td>77.9 Hz throughput (31.1 ms latency); achieves 97.1% SR on complex LIBERO spatial targets [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>].</td>
</tr>
<tr>
<td>(Pi-0)</td>
<td>PaliGemma (Gemma-2B)/SigLIP</td>
<td>Cross-embodiment Diffusion Flow-Matching</td>
<td>Large-scale multi-GPU clusters</td>
<td>Zero-shot adaptable across 22&#x002B; platforms; utilises diffusion matching to achieve high dexterity [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>].</td>
</tr>
<tr>
<td>RDT-1B</td>
<td>OpenCLIP/Qwen</td>
<td>1.2B Parameter Diffusion Foundation Model</td>
<td>High-end GPU workstations</td>
<td>Excels at dynamic bimanual manipulation with near-perfect zero-shot transfer capabilities [<xref ref-type="bibr" rid="ref-2">2</xref>].</td>
</tr>
<tr>
<td>RT-2</td>
<td>PaLI-X/PaLM-E</td>
<td>Symbol-tuning Autoregressive Transformer</td>
<td>Cloud-tethered infrastructure</td>
<td>First true VLA to co-finetune internet VQA data with robotic data; 1.8 to<break/> 3 Hz (highly latency-bound) [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>].</td>
</tr>
<tr>
<td>ProgPrompt</td>
<td>External (e.g., YOLOX)/GPT-4</td>
<td>API-driven Python code generator with asserts</td>
<td>Edge CPU/GPU (Queries Cloud LLM API)</td>
<td>Outperforms open-loop semantic planners by verifying preconditions and mitigating cascading failures [<xref ref-type="bibr" rid="ref-52">52</xref>].</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Evaluation: From Benchmarks to Deployment Reality</title>
<p>The deployment of Deep Learning (DL) in robotic systems is fundamentally bottlenecked by the epistemological tension between the &#x201C;black-box&#x201D; nature of massive neural architectures and the strict deterministic requirements of physical robotics. DL models, especially foundation Vision-Language-Action (VLA) models, operate as highly non-linear, stochastic function approximators that map open-world, high-dimensional observations into latent manifolds. Because their exact decision boundaries are mathematically opaque, they are prone to unpredictable, potentially catastrophic extrapolation when applied to out-of-distribution (OOD) real-world physics. Conversely, robotics demands rigorous deterministic guarantees, such as formal Lyapunov stability and collision-free bounds, to prevent hardware destruction and ensure human safety [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. To bridge this divide without sacrificing the cognitive depth of modern foundation models, this review proposes a unique conceptual framework: Prediction-Bounded Verified Deployment (PBVD). Rather than relying on naive, uncertified deployment or exhaustive hardware testing, PBVD mathematically fuses Prediction-Powered Inference (PPI) with granular sequence meta-evaluation. Under PBVD, an uncertified neural policy is first heavily evaluated in a large-scale simulation. Instead of trusting the biased simulation output, PBVD utilises a minimal set of paired physical trials to compute a mathematically rigorous &#x201C;rectifier&#x201D;, This rectifier bounds the expected real-world safety and performance of the black-box policy via non-asymptotic confidence intervals. Consequently, the PBVD framework structurally guarantees that a policy&#x2019;s real-world failure probability is strictly quantified and constrained <italic>before</italic> it is granted continuous torque access to a physical machine [<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
<sec id="s7_1">
<label>7.1</label>
<title>Evaluation Blueprint: Formalising the Sim-to-Real Deployment Pipeline</title>
<p>Historically, robotic policies have been evaluated using a coarse, binary average Success Rate (SR) computed over a statistically insignificant number of physical trials (e.g., 20 to 30 rollouts) [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>]. This approach completely fails to capture the complexity of continuous movement and lacks statistical guarantees. To formalise deployment, we must transition to rigorous mathematical frameworks that leverage imperfect simulators to bound real-world performance.</p>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Prediction-Powered Inference (PPI) and the SureSim Benchmark</title>
<p>Because physical evaluation is prohibitively expensive, the SureSim framework formalises the use of Prediction-Powered Inference (PPI) to augment small-scale real tests with large-scale simulation. Given a real-to-sim mapping function <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>g</mml:mi><mml:mo>&#x003A;</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the framework collects paired real and simulated evaluations, and additional purely simulated evaluations. The uniform PPI estimator corrects the simulation bias to compute the true mean: <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-58">58</xref>]. The first term mathematically acts as a rectifier, compensating for the Reality Gap (the divergence between simulated dynamics and real dynamics). By applying the Waudby-Smith and Ramdas (WSR) algorithm to this estimator, SureSim establishes finite-sample valid confidence intervals. Quantitative Support: Empirical deployment of SureSim on physics-based manipulation benchmarks demonstrated that this framework reduces the required real-world hardware evaluation effort by 20%&#x2013;25% while achieving tighter confidence intervals than classical real-only statistical bounds, reducing the interval width by 14.4% when scaling to 700 simulations [<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
</sec>
<sec id="s7_3">
<label>7.3</label>
<title>Granular Subgoal Tracking and Neural Meta-Evaluation</title>
<p>Beyond sample efficiency, evaluating complex long-horizon behaviours requires abandoning the scalar pass/fail paradigm. The StepEval blueprint formalises task evaluation as a trajectory-level vector <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where each element corresponds to a specific sub-task. Rather than relying on human annotation, this framework utilises Vision-Language Models (VLMs) as automated judges, mapping visual trajectories to the predicted success vector <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-57">57</xref>]. However, VLMs cannot evaluate the continuous quality of joint kinematics. To address this, the Neural Meta Evaluator (NeME) frames trajectory assessment as an offline sequence classification problem. NeME processes a time window of joint trajectories using a neural sequence model (parameterised by) to predict behaviour <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mrow><mml:mover><mml:mi>b</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>&#x03A6;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. During inter-policy selection, standard validation loss fundamentally fails to identify optimal models. However, selecting policy weights using NeME&#x2019;s meta-F1 score (mF1) aligns perfectly with the epoch (Epoch 8) that achieves peak physical success rates on human-robot collaborative tasks. Furthermore, structural comparisons reveal that an LSTM-based NeME with a window length <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> achieves an mF1 score of 66.6 <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3%, vastly outperforming state-space models like Mamba, which suffered representational collapse at 51.9 <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 12.7% mF1 on robotic joint sequences [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
</sec>
<sec id="s7_4">
<label>7.4</label>
<title>Core Algorithmic Engines: Mathematical Nuances and Deployment Reality</title>
<p>To understand deployment reality, we must describe the core machine learning paradigms entirely in prose, embedding their foundational mathematical mechanics to reveal how their optimisation strategies dictate real-world latency, hardware constraints, and sim-to-real transferability. The Proximal Policy Optimisation (PPO) algorithm is an on-policy, actor-critic framework designed to guarantee monotonic policy improvement by mathematically preventing destructively large gradient updates that cause catastrophic failure in physical robots [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. It achieves this by updating the policy network using a clipped surrogate objective: <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mtext>clip</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability ratio between the new and old policies, and <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the advantage estimate. By clipping the ratio, PPO forces the policy to stay within a trusted region, yielding stable convergence [<xref ref-type="bibr" rid="ref-8">8</xref>]. While sample-inefficient compared to off-policy methods, PPO&#x2019;s gradient stability makes it highly preferred for transferring robust locomotion capabilities to real quadrupedal hardware without relying on massive replay buffers. Conversely, the Soft Actor-Critic (SAC) algorithm is an off-policy method formulated within a maximum-entropy framework, specifically designed to maximise sample efficiency in real-world sparse-reward environments. The SAC algorithm optimises a stochastic policy by maximising both the expected cumulative reward and the policy&#x2019;s entropy, as defined by the objective <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2211;</mml:mo><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>.</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The temperature parameter dictates the balance between exploiting the highest-value action and exploring diverse kinematic trajectories [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. While its inherent stochasticity provides robustness to external physical perturbations, SAC&#x2019;s off-policy updates and replay buffer management introduce severe computational bottlenecks during gradient updates, often making it less efficient in terms of pure wall-clock training time than PPO on physical hardware.</p>
<p>Moving beyond standard reinforcement learning, modern generative manipulation is dominated by Diffusion Policies. These architectures model the action generation process not as a direct prediction but as learning to reverse a stochastic forward noising process applied to highly complex, multimodal continuous action trajectories. The network starts with pure Gaussian noise and iteratively denoises it into an action chunk. The training objective minimises a score-matching loss: <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mn>0</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mfrac linethickness="0"><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> where the denoiser is conditioned on and fuses visual and language features. This framework resolves the sparse-reward bottleneck and naturally handles the multimodality of human demonstrations, making it highly effective for dexterous bimanual manipulation. Finally, Vision-Language-Action (VLA) architectures cast visual processing, language understanding, and physical action generation into a unified sequence modelling problem. In standard autoregressive VLAs, the policy is trained via behavioural cloning using a next-token prediction objective: <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x003A;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, treating continuous actions as discrete text tokens. However, this left-to-right generation poses a severe latency bottleneck. To address this, the OpenVLA-OFT formulation discards discrete tokens in favour of continuous L1 Regression applied across parallel-decoded action chunks [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Taking a different hybrid approach, the Discrete Diffusion VLA applies flow-matching directly to tokenised action chunks via a discrete transition matrix <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, allowing the transformer to adaptively unmask high-confidence discrete tokens in parallel [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
<sec id="s7_5">
<label>7.5</label>
<title>Comparative Technical Analysis of ML Paradigms</title>
<p>To objectively evaluate the deployment readiness of these paradigms, the following table synthesises their execution metrics, computational overhead, and hardware configurations. <xref ref-type="table" rid="table-9">Table 9</xref> shows the lgorithmic Efficacy and Hardware Deployment of ML Paradigms. While significant strides have been made in scaling VLA architectures and in deriving real-to-sim statistical bounds, a critical, non-obvious gap remains in dynamic vocabulary alignment during OOD recovery. Current Discrete Diffusion VLAs and autoregressive models rely on fixed patch embeddings and static BPE tokenisation for continuous states. If a deployed policy encounters anomalous tissue compliance or severe sensor noise that falls outside its tokenised distribution, it suffers an immediate latent misalignment between the visual encoder and the LLM backbone. Existing systems cannot dynamically request continuous recalibration of their action vocabularies without complete offline retraining. Future deployment blueprints must integrate verifiable fallback mechanisms (such as the PBVD framework&#x2019;s rectifier) that trigger automated execution halts and explicit semantic re-prompting when the confidence intervals of the diffusion generation collapse in real time [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>ML paradigms comparison in terms of 4 aspects.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>ML Paradigm</th>
<th>Action Decoding/Objective Strategy</th>
<th>Target Benchmark &#x0026; Evaluation Metric</th>
<th>Computational/<break/>Hardware Overhead</th>
<th>Quantitative Efficacy &#x0026; Inference Speed</th>
</tr>
</thead>
<tbody>
<tr>
<td>OpenVLA-OFT&#x002B;</td>
<td>Continuous L1 Regression, Parallel Decoding</td>
<td>ALOHA Bimanual (Multi-stage evaluation)</td>
<td>NVIDIA A100 GPU (LoRA optimised)</td>
<td>77.9 Hz throughput (0.321 s latency); solves complex long-horizon bimanual alignments at 89.6% avg SR [<xref ref-type="bibr" rid="ref-12">12</xref>].</td>
</tr>
<tr>
<td>Diffusion Policy (Image)</td>
<td>Denoising score-matching over continuous chunks</td>
<td>ALOHA Bimanual (Multi-stage evaluation)</td>
<td>NVIDIA A100 GPU</td>
<td>267.4 Hz throughput (0.090 s latency); excellent dexterity but lacks zero-shot semantic generalisation [<xref ref-type="bibr" rid="ref-12">12</xref>].</td>
</tr>
<tr>
<td>Discrete Diffusion VLA</td>
<td>Masked-token discrete flow-matching</td>
<td>LIBERO Benchmark (Spatial, Object, Goal, Long)</td>
<td>4&#x00D7; NVIDIA A800 GPUs (Training)</td>
<td>14.53 Hz inference; 96.3% avg. SR on LIBERO (outperforms autoregressive OpenVLA by &#x007E;20%) [<xref ref-type="bibr" rid="ref-21">21</xref>].</td>
</tr>
<tr>
<td>NeME (Meta-Evaluator)</td>
<td>Offline trajectory sequence classification (LSTM)</td>
<td>HRIC/ergoCub joint sequences</td>
<td>Edge-deployable CPU/GPU inference</td>
<td>Identifies optimal weights that perfectly match physical robot performance; 66.6% mF1 score [<xref ref-type="bibr" rid="ref-40">40</xref>].</td>
</tr>
<tr>
<td>SureSim (PPI Real2Sim)</td>
<td>Uniform PPI Estimator</td>
<td>Pick-and-place generalizability</td>
<td>Massively parallel physics simulator</td>
<td>Decreases required real-world trials by &#x003E;20%&#x2013;25%; shrinks confidence interval bounds by 14.4% [<xref ref-type="bibr" rid="ref-58">58</xref>].</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s8">
<label>8</label>
<title>System Design Blueprint for LLM-Enabled Robotic Assistance</title>
<sec id="s8_1">
<label>8.1</label>
<title>Bridging the Semantic-Kinematic Divide: The Interdisciplinary Rationale</title>
<p>The development of fully autonomous robotic assistants is fundamentally hindered by a dichotomous specialisation in artificial intelligence research. Specialists in Natural Language Processing (NLP) and Large Language Models (LLMs) focus on the semantic alignment of discrete tokens, yielding systems capable of open-world reasoning, deep common-sense logic, and hierarchical task decomposition [<xref ref-type="bibr" rid="ref-59">59</xref>]. Conversely, experts in Control Theory and Reinforcement Learning (RL) operate within the kinematic domain, focusing on continuous state spaces, high-frequency torque control, and formal stability guarantees necessary for safe physical interaction [<xref ref-type="bibr" rid="ref-20">20</xref>]. Addressing LLMs and RL within a unified manuscript is critical because neither paradigm can achieve reliable robotic autonomy in isolation. LLMs lack physical grounding and the ability to execute contact-rich, continuous control manoeuvres, while RL agents suffer from severe sample inefficiency and an inability to reason over long-horizon, abstract tasks without explicit, heavily engineered reward functions [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>]. By formulating a taxonomy that addresses both high-level semantic planning and low-level continuous control, this review bridges the gap between these distinct specialisations. We present a blueprint in which the LLM&#x2019;s abstract reasoning serves as a semantic manifold that strictly bounds the exploration space of the low-level RL controller, resulting in a cohesive framework for deployment in dynamic environments. To operationalise the integration of LLMs and RL, we propose a novel conceptual framework: Semantic-Kinematic Symbiosis (SKS). The SKS architecture structurally formalises robotic execution as a Hierarchical Partially Observable Markov Decision Process (H-POMDP). In this framework, the LLM acts as the high-level semantic planner, interpreting ambiguous human instructions into a sequence of intermediate goals. Rather than allowing the LLM to hallucinate physically impossible tasks, the SKS framework mathematically constrains the LLM using the RL policy&#x2019;s learned value functions, which act as representations of physical affordance. This integration is formally captured by the &#x201C;SayCan&#x201D; formulation. In prose, the probability of executing a successful robotic action is determined by the intersection of semantic utility and physical capability. The framework computes the optimal action by maximising the product of two probabilities: <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the language model predicting the semantic likelihood of a skill contributing to the high-level instruction, and <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the reinforcement learning value function providing the affordance, the probability that the robot can physically execute the skill from its current state [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
</sec>
<sec id="s8_2">
<label>8.2</label>
<title>High-Level Task Planning: LLMs as Semantic Oracles</title>
<p>The first stage of the SKS blueprint utilises autoregressive LLMs to translate natural language into discrete, executable task sequences. The fundamental operation of the LLM planner relies on maximising the conditional probability of a sequence of programming tokens <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>y</mml:mi></mml:math></inline-formula> given an input prompt <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>x</mml:mi></mml:math></inline-formula>, formalised as <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x003A;</mml:mo><mml:mi>x</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. To ensure the LLM does not generate actions beyond the robot&#x2019;s hardware limits, modern architectures rely on programmatic prompt structures. Frameworks such as ProgPrompt inject Pythonic API specifications directly into the LLM&#x2019;s context window. Instead of outputting free-form text, the LLM generates Python code that includes assert statements to continuously monitor the robot&#x2019;s state feedback during execution [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>]. Furthermore, the Interactive Predicate Learning (InterPreT) framework uses GPT-4 to generate complex semantic predicates as Python functions, which are iteratively refined with natural-language feedback from human operators, thereby seamlessly translating raw observation states into logical preconditions for task planners [<xref ref-type="bibr" rid="ref-50">50</xref>]. Utilising structured programming language prompts, ProgPrompt achieved an execution success rate of 0.87 on complex VirtualHome tasks, substantially reducing the computational overhead of plan generation while demonstrating near-linear scalability as environments expand, and drastically outperforming classical A&#x002A; planners [<xref ref-type="bibr" rid="ref-52">52</xref>].</p>
</sec>
<sec id="s8_3">
<label>8.3</label>
<title>Low-Level Continuous Control: Robust Reinforcement Learning</title>
<p>Once the semantic sub-goal is generated, the system delegates execution to a high-frequency RL policy capable of managing complex, contact-rich physical dynamics. To ensure that the robot learns safely without catastrophic hardware failure, the Proximal Policy Optimisation (PPO) algorithm is frequently deployed. In prose, PPO stabilises learning by preventing destructively large policy updates. It achieves this by updating the policy network using a clipped surrogate objective: <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mtext>clip</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mrow><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability ratio between the new and old policies, and is the advantage estimate. This mathematical clipping physically restricts the policy&#x2019;s step size, thereby guaranteeing stable convergence during complex manoeuvres [<xref ref-type="bibr" rid="ref-20">20</xref>]. For highly dexterous applications, the state-of-the-art RL-100 framework embeds RL directly into a continuous diffusion visuomotor policy. It utilises a consistency model distillation to compress multi-step diffusion into a single-step action generator, effectively bridging offline imitation learning with online, high-frequency RL fine-tuning. The RL-100 framework represents a definitive benchmark in deployment readiness. It achieved a 100% success rate across 1000 evaluated real-world episodes on diverse tasks (such as dynamic pushing and juicing) and demonstrated formidable robustness, maintaining a &#x007E;96% success rate even under aggressive human perturbations. Crucially, a juicing robot driven by this framework operated autonomously for seven continuous hours in a public shopping mall without a single failure [<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
</sec>
<sec id="s8_4">
<label>8.4</label>
<title>Integration Taxonomy: The Plan-Seq-Learn (PSL) Paradigm</title>
<p>To unify these domains operationally, the Plan-Seq-Learn (PSL) paradigm offers a highly scalable, modular pipeline. In PSL, the task is strictly decoupled: the LLM predicts a high-level sequence of target regions (Plan), an off-the-shelf vision-based motion planner moves the robotic arm into proximity of the target (Seq), and a localised RL policy is activated solely to manage the final centimetres of contact-rich manipulation (Learn). This approach explicitly isolates the sample inefficiency of RL to the final execution phase, relying on LLMs and classical motion planning to bypass the long-horizon sparse-reward problem. By confining RL exploration to localised regions specified by the LLM, the PSL architecture achieves a staggering 96.0% success rate on 10-stage, contact-rich operations such as &#x201C;NutAssembly&#x201D; directly from raw visual inputs, thereby comprehensively outperforming purely end-to-end systems that suffer from cascading sequence errors [<xref ref-type="bibr" rid="ref-20">20</xref>]. In related hybrid approaches that combine LLMs and RL for Franka Emika Panda manipulation, task completion times were reduced by 33.5% (from 18.5 to 12.3 s), while task adaptability increased by 36.4% compared to RL-only baselines [<xref ref-type="bibr" rid="ref-55">55</xref>].</p>
</sec>
<sec id="s8_5">
<label>8.5</label>
<title>Healthcare Focus Theme: Deployment in Clinical and Assistive Environments</title>
<p>The SKS blueprint is particularly transformative for the healthcare sector, where operations demand both empathetic human understanding and zero-tolerance kinematic precision. Socially Assistive Robots (SARs) in Eldercare: SARs deployed to support elderly cognitive function help mitigate global nursing shortages and loneliness. Integrating conversational agents powered by LLMs (e.g., GPT-3.5 on the Social Robot Mini) enables robots to contextualise nuanced, bidirectional emotional cues and patient histories into highly personalised care responses, rather than relying on rigid, pre-programmed dialogue trees [<xref ref-type="bibr" rid="ref-61">61</xref>]. Frameworks like RobotIQ integrate these LLM-driven interactions directly into Robot Operating System (ROS) libraries, translating an elderly patient&#x2019;s spoken request (e.g., &#x201C;I am thirsty&#x201D;) into precise localisation, navigation, and object-fetching APIs executed by low-level RL controllers [<xref ref-type="bibr" rid="ref-62">62</xref>]. Surgical Robotics: In precision environments, the RoboNurse-VLA framework functions as an automated scrub nurse. Processing voice prompts and dynamic visual scenes in real-time, the system executes surgical instrument handovers, adapting autonomously to unseen tools and rapidly changing operating room conditions [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
</sec>
<sec id="s8_6">
<label>8.6</label>
<title>Identification of Non-Obvious Gaps</title>
<p>While the SKS blueprint effectively integrates semantic intent with kinematic execution, a critical, non-obvious gap remains: the lack of bidirectional, continuous causal feedback. Current hybrid frameworks are predominantly top-down; the LLM instructs the RL policy. However, if the high-frequency RL agent encounters unmodeled physical compliance, such as a surgical tool slipping or unexpected patient resistance during rehabilitation, it currently cannot mathematically encode this continuous physical failure back into the LLM&#x2019;s discrete textual embedding space [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>]. Future architectures must develop dynamic, tactile-to-language vocabulary alignment algorithms. By utilising visual-language evaluators (e.g., StepEval) to classify sub-goal kinematic failures into descriptive language tokens [<xref ref-type="bibr" rid="ref-57">57</xref>], the system could establish a closed-loop symbiotic pipeline, allowing the LLM to dynamically re-route its semantic planning graph in response to localised physical constraints.</p>
</sec>
<sec id="s8_7">
<label>8.7</label>
<title>Integration of Frameworks</title>
<p>The individual frameworks are integrated into a cohesive functional stack as follows:<list list-type="bullet">
<list-item>
<p>Cognitive Layer (SKS/CSE): The Semantic-Kinematic Symbiosis (SKS) framework acts as the &#x2018;System 2&#x2019; reasoning engine, which is structurally governed by Certified-Semantic Embodiment (CSE) to filter semantic outputs through physical affordance bounds.</p></list-item>
<list-item>
<p>Stability Layer (LBSE): The Lyapunov-Bounded Semantic Execution (LBSE) framework provides the &#x2018;System 1&#x2019; fast-frequency control, translating the high-level semantic waypoints into deterministic, safety-guaranteed motor commands.</p></list-item>
<list-item>
<p>Feedback &#x0026; Verification Layer (PBVD/BCKG): The Prediction-Bounded Verified Deployment (PBVD) framework quantifies the reality gap during execution. Crucially, the BCKG framework provides the bidirectional causal link, allowing low-level kinematic failures to be re-encoded as text tokens to dynamically update the high-level SKS planning graph.</p></list-item>
</list></p>
<p>This bidirectional flow ensures that the system does not just operate top-down, but functions as a transparent, reproducible, and safety-critical closed loop. <xref ref-type="table" rid="table-10">Table 10</xref> shows our revised version, we demonstrate the &#x201C;Closed-Loop&#x201D; nature to the reviewer.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Framework taxonomy.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Architecture Layer</th>
<th>Framework</th>
<th>Primary Function</th>
<th>Loop Position</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td>Reasoning</td>
<td>SKS</td>
<td>Semantic goal decomposition</td>
</tr>
<tr>
<td>Certification</td>
<td>CSE</td>
<td>Safety-filtering of goals</td>
<td>Feed-forward Constraint</td>
</tr>
<tr>
<td>Execution</td>
<td>LBSE</td>
<td>Lyapunov-stable torque control</td>
<td>Low-level (Action)</td>
</tr>
<tr>
<td>Verification</td>
<td>PBVD</td>
<td>Real-time performance rectifier</td>
<td>Monitoring</td>
</tr>
<tr>
<td>Governance</td>
<td>BCKG</td>
<td>Causal failure feedback</td>
<td>Feedback (Loop Closure)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s9">
<label>9</label>
<title>Open Challenges and Research Directions</title>
<p>The transition of robotic policies from simulated environments or constrained datasets to dynamic, real-world deployment is bottlenecked by severe algorithmic and mathematical limitations. To address the complexities of these challenges, it is necessary to move beyond generalised descriptions and to critically dissect the mathematical formulations and architectural variations that define the field&#x2019;s current limitations.</p>
<sec id="s9_1">
<label>9.1</label>
<title>The Reality Gap and Sim-to-Real Transfer Discrepancies</title>
<p>The &#x201C;Reality Gap&#x201D; fundamentally stems from the inability of any simulator to perfectly model real-world physics, chaotic non-linearities, and sensor noise. Mathematically, this is framed as a discrepancy between a simulated Partially Observable Markov Decision Process (POMDP) <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the real-world POMDP <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The reality gap comprises distinct sub-gaps, notably the dynamics gap, defined as the expected divergence between transition models:</p>
<p><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. Because policies trained often exploit these inaccuracies to maximise rewards, evaluating the performance gap <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> requires specialised statistical frameworks [<xref ref-type="bibr" rid="ref-6">6</xref>]. To construct finite-sample valid confidence intervals without exhausting physical hardware, modern evaluation architectures employ Prediction-Powered Inference (PPI). Frameworks such as SureSim construct a paired dataset <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> mapping identical initial conditions across real and simulated trials. This allows the architecture to compute a rectifier variance, where high variance indicates low simulation-to-reality correlation, mathematically bounding the exact limits of the simulation&#x2019;s predictive utility [<xref ref-type="bibr" rid="ref-58">58</xref>]. To algorithmically mitigate <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, residual learning architectures diverge from standard Domain Randomisation by parameterising the dynamics as a composite function. Rather than assuming the simulator&#x2019;s analytical model is complete, residual networks learn a corrective function <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> over the simulator&#x2019;s predicted state transitions such that <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2248;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This structural separation allows the neural network to absorb unmodeled compliance and complex aerodynamic forces without requiring end-to-end retraining of the foundational physics engine [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
</sec>
<sec id="s9_2">
<label>9.2</label>
<title>Safety-Critical Machine Learning and State Constraints</title>
<p>In physical deployments, exploration and control policies cannot violate hard mechanical or environmental limits. Safe Reinforcement Learning reformulates the standard MDP into a Constrained Markov Decision Process (CMDP) defined by the tuple <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Here, policies must maximise the standard expected return while strictly bounding the expected discounted constraint cost <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msubsup><mml:mi>J</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2264;</mml:mo><mml:mi>d</mml:mi></mml:math></inline-formula>, where represents thresholded safety functions [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. Architecturally, enforcing these constraints relies on two distinct paradigms: Control Barrier Functions (CBFs), Lyapunov Certification, Variable Impedance and Safety Critics. In control-affine systems, safety is certified by maintaining the forward invariance of a robust safe set <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. If the true system dynamics <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are partially unknown, the time derivative of the CBF relies on an approximation <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mrow><mml:mover><mml:mi>f</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The architectural nuance lies in bounding the Lipschitz constant of the deep neural network parameterising the policy, often enforced via spectral normalisation. This ensures that the mapping of bounded uncertainty in the dynamics directly translates into bounded uncertainty in the CBF&#x2019;s time derivative, thereby mathematically guaranteeing safety [<xref ref-type="bibr" rid="ref-8">8</xref>]. To proactively avoid risky states prior to actuation, frameworks like SRL-VIC decouple task optimisation from safety through a dedicated Safety Critic network. The critic evaluates a recursive risk function <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, trained via Mean Squared Error, to estimate the probability of constraint violation. If <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msubsup><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> exceeds a strict threshold, the architecture bypasses the primary actor network and queries a pre-trained recovery policy (<inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) to sample a mathematically safe projection [<xref ref-type="bibr" rid="ref-63">63</xref>].</p>
</sec>
<sec id="s9_3">
<label>9.3</label>
<title>Sparse Reward Optimisation and Sample Inefficiency</title>
<p>For tasks demanding complex sequencing (e.g., long-horizon manipulation), defining a continuous reward gradient is highly susceptible to reward hacking. Consequently, the environment is often modelled with a sparse, binary reward structure (e.g., at the target, otherwise) [<xref ref-type="bibr" rid="ref-64">64</xref>]. This creates a severe sample inefficiency challenge, as standard Temporal Difference (TD) updates, such as <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, fail to propagate useful gradients when <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> remains uniformly <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-65">65</xref>]. To artificially synthesise dense gradients, architectures integrate Hindsight Experience Replay (HER) [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. HER alters the replay buffer dynamics by modifying the tuple <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> after a failed rollout. It replaces the unachieved original goal <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>g</mml:mi></mml:math></inline-formula> with an achieved posterior state <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo>&#x0060;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, forcing the Bellman equation to process a synthetic positive reward [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]. Advanced frameworks, such as Deep Value-and-Predictive-Model Control (DVPMC), integrate HER into a Model-Based RL (MBRL) pipeline [<xref ref-type="bibr" rid="ref-36">36</xref>]. DVPMC approximates the value function alongside a learned transition model and utilises sampling-based cross-entropy methods for action selection, thereby reducing the prediction horizon and circumventing the massive sample complexity inherent in model-free off-policy algorithms such as SAC or DDPG [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>].</p>
</sec>
<sec id="s9_4">
<label>9.4</label>
<title>Distributional Shift and Epistemic Uncertainty</title>
<p>Both offline RL and real-time perception models suffer catastrophic performance degradation when evaluating Out-of-Distribution (OOD) states, an issue rooted in epistemic uncertainty. When training purely from static datasets, high-capacity function approximators systematically overestimate the Q-values of Out-of-Distribution actions [<xref ref-type="bibr" rid="ref-30">30</xref>]. To mitigate this distributional shift, algorithms either directly constrain the learned policy distribution to match the behaviour policy&#x2019;s distribution or deploy Conservative Q-Learning (CQL) to strictly minimise a lower bound on the value function [<xref ref-type="bibr" rid="ref-66">66</xref>]. Furthermore, when goal-conditioned architectures encounter OOD goals, algorithms inject noise perturbation with probability over the goal states, explicitly penalising the corresponding values via negative Temporal Difference errors to suppress exploratory divergence [<xref ref-type="bibr" rid="ref-35">35</xref>]. In unstructured environments, perception architectures must mathematically distinguish between aleatoric uncertainty (inherent sensory noise) and epistemic uncertainty (OOD inputs lacking training support). Using evidential deep learning, the network parameterises a Dirichlet distribution rather than standard categorical logits. A normalising flow network tracks the latent density of the features <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The epistemic uncertainty is then thresholded using a confidence score formulation: <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>. If <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> falls below a predefined percentile threshold (e.g., the k-th percentile of the training distribution), the feature is mathematically flagged as OOD, and the downstream model-predictive controller applies auxiliary cost penalties to prevent the robot from navigating into untrusted regions [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
</sec>
</sec>
<sec id="s10">
<label>10</label>
<title>Conclusion</title>
<p>The integration of Artificial Intelligence into robotic systems has catalysed a profound paradigm shift, transitioning the field from rigid, pre-programmed automation to highly dynamic, open-world autonomy driven by Vision-Language-Action (VLA) models and Deep Reinforcement Learning (DRL). This critical review has systematically synthesised the algorithms, architectures, and evaluation frameworks required to deploy these autonomous systems. However, as robotics permeates high-stakes, unstructured environments, particularly within P5 (predictive, personalised, preventive, participatory, and precision) medicine, the epistemological tension between the stochastic, &#x201C;black-box&#x201D; nature of deep neural networks and the deterministic, safety-critical requirements of physical execution remains the primary bottleneck to real-world deployment. To resolve this tension and guide the next decade of robotic research, we outline a strategic research agenda. Furthermore, we propose a novel, unifying conceptual framework that integrates cognitive reasoning, kinematic execution, and stringent regulatory compliance. Throughout this review, a persistent, non-obvious gap has been identified across state-of-the-art hybrid models (such as Plan-Seq-Learn or OpenVLA): the distinct lack of bidirectional causal feedback. Currently, high-level Large Language Models (LLMs) issue top-down semantic commands, but if a low-level RL controller encounters an unmodeled physical constraint (e.g., anomalous tissue compliance during surgery or a slipping object), it cannot mathematically translate that continuous physical failure back into the discrete textual embedding space of the LLM to trigger dynamic re-routing. We propose the Bidirectional Causal-Kinematic Governance (BCKG) architecture as a definitive roadmap for the future. In the BCKG framework, the low-level Lyapunov-certified execution manifold is equipped with a Neuro-Symbolic Failure Encoder. When a physical constraint is violated, this encoder translates the kinematic discrepancy <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> (e.g., assert Object_Slip(Target)) into a formalised logic predicate (e.g., assert Object_Slip(Target)). This predicate is fed back into the LLM&#x2019;s context window as an active prompt, forcing the language model to reason causally about the failure and generate a revised topological plan. Crucially, the BCKG framework wraps this entire bidirectional loop in an immutable, cryptographic data logger, creating a transparent, reproducible audit trail of every semantic-to-kinematic decision, a strict necessity for adhering to the General Data Protection Regulation (GDPR) and the European Artificial Intelligence Act.</p>
<sec id="s10_1">
<label>10.1</label>
<title>Causal Reinforcement Learning and Neuro-Symbolic Integration</title>
<p>Future architectures must move beyond purely correlational deep learning by integrating Causal Reinforcement Learning and Neuro-Symbolic approaches. Relying strictly on model-free DRL is profoundly sample-inefficient and opaque. By incorporating bidirectional dynamics models, which simultaneously perform forward and inverse predictions in a latent space, algorithms can self-supervise the denoising of representations, substantially alleviating model bias and prediction errors. Furthermore, integrating structured symbolic logic with neural networks provides exact interpretability. Integrating advanced transfer learning with foundational models (such as the YOLO &#x002B; SAC pipeline) has already demonstrated the capacity to drastically reduce robotic training times, achieving convergence in 6443 s compared to 15.9 times longer without transfer techniques. Expanding these causal pipelines will be strictly necessary to achieve human-level, few-shot problem-solving capabilities.</p>
</sec>
<sec id="s10_2">
<label>10.2</label>
<title>Continual Lifelong Learning and Dynamic Modality Expansion</title>
<p>Robots deployed in dynamic clinical environments cannot remain frozen after offline training; they must continuously adapt without suffering from <italic>catastrophic forgetting</italic>. Future research must scale Hierarchical Lifelong Reinforcement Learning (HLifeRL) frameworks. Current implementations of HLifeRL utilising dynamic option libraries have demonstrated profound stability; when expanding a robotic skill library from four complex locomotion tasks to five, the master policy retained over 95% of its original performance metrics, effectively resisting the parameter interference that destroys standard monolithic networks. Concurrently, the input space of VLA models must expand beyond vision and text. To execute high-precision medical tasks, future foundation models must ingest multi-modal high-frequency streams, specifically integrating tactile force/torque sensors, auditory feedback, and depth data, which are currently severely underutilised in end-to-end architectures.</p>
</sec>
<sec id="s10_3">
<label>10.3</label>
<title>Dependable Clinical Deployment and Ethical Governance</title>
<p>The transformative impact of robotic automation in healthcare, ranging from Da Vinci surgical systems providing 3D-HD millimetre-level precision to Socially Assistive Robots (SARs) mitigating nursing workforce shortages, is fundamentally dictated by regulatory hurdles rather than pure algorithmic capability. Systems governed by the Food and Drug Administration (FDA) and the European Medicines Agency (EMA) demand comprehensive, verifiable safety and performance testing. Because learning algorithms inherently evolve over time, existing regulatory frameworks struggle to accommodate continuous online updates, severely decelerating clinical rollout. Future research must develop standardised, automated Verification and Validation (V&#x0026;V) benchmarks specifically designed for continuous-learning medical devices. From an ethical and legal standpoint, current medical AI and robotic systems lack independent moral status; humans and institutional stakeholders remain the absolute duty bearers. Future deployments must establish clear legal attribution frameworks, ensuring that when an autonomous robotic assistant executes a flawed manoeuvre, the liability can be transparently traced through the system&#x2019;s runtime monitors and algorithmic decision logs. Finally, the evaluation of healthcare robotics must move away from short-term laboratory simulations. Future studies demand extensive, longitudinal field trials within authentic clinical settings to rigorously evaluate the cost-effectiveness, long-term patient engagement, and human-robot trust transfer necessary to validate these intelligent physical robots as transformative, reliable medical tools.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no funding for this work.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualisation and supervision, Ahmed Ismail Ebada; methodology and data curation, Yasmeen Abu-Seif and Hrushikesh Pardeshi; investigation, Yasmeen Abu-Seif; writing original draft preparation, Ahmed Ismail Ebada, Yasmeen Abu-Seif and Hrushikesh Pardeshi; writing review and editing, Yasmeen Abu-Seif and Nesma El-Sayed. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>All data generated or analysed during this study are included in this published article. Any additional data supporting the findings of this study are available from the corresponding authors upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>ML</term>
<def>
<p>Machine Learning</p>
</def>
</def-item>
<def-item>
<term>RL</term>
<def>
<p>Reinforcement Learning</p>
</def>
</def-item>
<def-item>
<term>IL</term>
<def>
<p>Imitation Learning</p>
</def>
</def-item>
<def-item>
<term>SSL</term>
<def>
<p>Self-Supervised Learning</p>
</def>
</def-item>
<def-item>
<term>UL</term>
<def>
<p>Unsupervised Learning</p>
</def>
</def-item>
<def-item>
<term>DRL</term>
<def>
<p>Deep Reinforcement Learning</p>
</def>
</def-item>
<def-item>
<term>POMDP</term>
<def>
<p>Partially Observable Markov Decision Process</p>
</def>
</def-item>
<def-item>
<term>CVaR</term>
<def>
<p>Conditional Value at Risk</p>
</def>
</def-item>
<def-item>
<term>TDP</term>
<def>
<p>Thermal Design Heat</p>
</def>
</def-item>
<def-item>
<term>FLOPs</term>
<def>
<p>Floating Point Operations</p>
</def>
</def-item>
<def-item>
<term>RSSM</term>
<def>
<p>Recurrent State-Space Model</p>
</def>
</def-item>
<def-item>
<term>MSE</term>
<def>
<p>Mean Squared Error</p>
</def>
</def-item>
<def-item>
<term>MPC</term>
<def>
<p>Model Predictive Control</p>
</def>
</def-item>
<def-item>
<term>LLM</term>
<def>
<p>Large Language Model</p>
</def>
</def-item>
<def-item>
<term>VLA</term>
<def>
<p>Vision-Language-Action Model</p>
</def>
</def-item>
<def-item>
<term>VLM</term>
<def>
<p>Vision-Language Model</p>
</def>
</def-item>
<def-item>
<term>HRI</term>
<def>
<p>Human-Robot Interaction</p>
</def>
</def-item>
<def-item>
<term>MPC</term>
<def>
<p>Model Predictive Control</p>
</def>
</def-item>
<def-item>
<term>SAC</term>
<def>
<p>Soft Actor-Critic</p>
</def>
</def-item>
<def-item>
<term>PPO</term>
<def>
<p>Proximal Policy Optimisation</p>
</def>
</def-item>
<def-item>
<term>OOD</term>
<def>
<p>Out-of-Distribution</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Soori</surname> <given-names>M</given-names></string-name>, <string-name><surname>Arezoo</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dastres</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Artificial intelligence, machine learning and deep learning in advanced robotics, a review</article-title>. <source>Cogn Robot</source>. <year>2023</year>;<volume>3</volume>:<fpage>54</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cogr.2023.04.001</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Din</surname> <given-names>MU</given-names></string-name>, <string-name><surname>Akram</surname> <given-names>W</given-names></string-name>, <string-name><surname>Saoud</surname> <given-names>LS</given-names></string-name>, <string-name><surname>Rosell</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Vision language action models in robotic manipulation: a systematic review</article-title>. <comment>arXiv:2507.10672. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2507.10672</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alqobali</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alnasser</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rashidi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alshmrani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alhmiedat</surname> <given-names>T</given-names></string-name></person-group>. <article-title>A real-time semantic map production system for indoor robot navigation</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>20</issue>):<fpage>6691</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24206691</pub-id>; <pub-id pub-id-type="pmid">39460171</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Deep reinforcement learning-based safe interaction for industrial human-robot collaboration using intrinsic reward function</article-title>. <source>Adv Eng Inform</source>. <year>2021</year>;<volume>49</volume>(<issue>12</issue>):<fpage>101360</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2021.101360</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Salvato</surname> <given-names>E</given-names></string-name>, <string-name><surname>Fenu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Medvet</surname> <given-names>E</given-names></string-name>, <string-name><surname>Pellegrino</surname> <given-names>FA</given-names></string-name></person-group>. <article-title>Crossing the reality gap: a survey on sim-to-real transferability of robot controllers in reinforcement learning</article-title>. <source>IEEE Access</source>. <year>2021</year>;<volume>9</volume>:<fpage>153171</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2021.3126658</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aljalbout</surname> <given-names>E</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>J</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Akinola</surname> <given-names>I</given-names></string-name>, <string-name><surname>Garrett</surname> <given-names>CR</given-names></string-name>, <string-name><surname>Heiden</surname> <given-names>E</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The reality gap in robotics: challenges, solutions, and best practices</article-title>. <source>Annu Rev Control Robot Auton Syst</source>. <year>2026</year>;<volume>9</volume>(<issue>1</issue>):<fpage>403</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1146/annurev-control-031924-100130</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Matsuno</surname> <given-names>K</given-names></string-name>, <string-name><surname>Cheah</surname> <given-names>CC</given-names></string-name></person-group>. <article-title>Lyapunov-based deep learning control for robots with unknown Jacobian</article-title>. <comment>arXiv:2509.04984. 2025</comment>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brunke</surname> <given-names>L</given-names></string-name>, <string-name><surname>Greeff</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hall</surname> <given-names>AW</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Panerati</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Safe learning in robotics: from learning-based control to safe reinforcement learning</article-title>. <source>Annu Rev Control Robot Auton Syst</source>. <year>2022</year>;<volume>5</volume>(<issue>1</issue>):<fpage>411</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1146/annurev-control-042920-020211</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ancha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Osteen</surname> <given-names>PR</given-names></string-name>, <string-name><surname>Bucher</surname> <given-names>B</given-names></string-name>, <string-name><surname>Phillips</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>EVORA: deep evidential traversability learning for risk-aware off-road autonomy</article-title>. <source>IEEE Trans Robot</source>. <year>2024</year>;<volume>40</volume>:<fpage>3756</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tro.2024.3431828</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shojaeinasab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jalayer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baniasadi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Najjaran</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Unveiling the black box: a unified XAI framework for signal-based deep learning models</article-title>. <source>Machines</source>. <year>2024</year>;<volume>12</volume>(<issue>2</issue>):<fpage>121</fpage>. doi:<pub-id pub-id-type="doi">10.3390/machines12020121</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Oquab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Darcet</surname> <given-names>T</given-names></string-name>, <string-name><surname>Moutakanni</surname> <given-names>T</given-names></string-name>, <string-name><surname>Vo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Szafraniec</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khalidov</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Dinov2: learning robust visual features without supervision</article-title>. <comment>arXiv:2304.07193. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2304.07193</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Finn</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Fine-tuning vision-language-action models: optimizing speed and success</article-title>. <comment>arXiv:2502.19645. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2502.19645</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shanks</surname> <given-names>S</given-names></string-name>, <string-name><surname>Embley-Riches</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Delfaki</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Ciliberto</surname> <given-names>C</given-names></string-name>, <string-name><surname>DreamerNav</surname> <given-names>K D</given-names></string-name></person-group>. <article-title>Learning-based autonomous navigation in dynamic indoor environments using world models</article-title>. <source>Front Robot AI</source>. <year>2025</year>;<volume>12</volume>:<fpage>1655171</fpage>. doi:<pub-id pub-id-type="doi">10.3389/frobt.2025.1655171</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kontolati</surname> <given-names>K</given-names></string-name>, <string-name><surname>Goswami</surname> <given-names>S</given-names></string-name>, <string-name><surname>Em Karniadakis</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shields</surname> <given-names>MD</given-names></string-name></person-group>. <article-title>Learning nonlinear operators in latent spaces for real-time predictions of complex dynamics in physical systems</article-title>. <source>Nat Commun</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>5101</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41467-024-49411-w</pub-id>; <pub-id pub-id-type="pmid">38876997</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>El-Hussieny</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Real-time deep learning-based model predictive control of a 3-DOF biped robot leg</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>16243</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-66104-y</pub-id>; <pub-id pub-id-type="pmid">39004665</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Fr&#x00E4;nti</surname> <given-names>P</given-names></string-name></person-group>. <article-title>A review of motion planning algorithms for intelligent robots</article-title>. <source>J Intell Manuf</source>. <year>2022</year>;<volume>33</volume>(<issue>2</issue>):<fpage>387</fpage>&#x2013;<lpage>424</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10845-021-01867-z</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Waga</surname> <given-names>A</given-names></string-name>, <string-name><surname>Benhlima</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bekri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abdouni</surname> <given-names>J</given-names></string-name>, <string-name><surname>Saber</surname> <given-names>FZ</given-names></string-name></person-group>. <article-title>A survey on autonomous navigation for mobile robots: from traditional techniques to deep learning and large language models</article-title>. <source>J King Saud Univ Comput Inf Sci</source>. <year>2025</year>;<volume>37</volume>(<issue>7</issue>):<fpage>198</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s44443-025-00216-x</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Sanchez</surname> <given-names>FR</given-names></string-name>, <string-name><surname>McCarthy</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bulens</surname> <given-names>DC</given-names></string-name>, <string-name><surname>McGuinness</surname> <given-names>K</given-names></string-name>, <string-name><surname>O&#x2019;Connor</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Dexterous robotic manipulation using deep reinforcement learning and knowledge transfer for complex sparse reward-based tasks</article-title>. <source>Expert Syst</source>. <year>2023</year>;<volume>40</volume>(<issue>6</issue>):<fpage>e13205</fpage>. doi:<pub-id pub-id-type="doi">10.1111/exsy.13205</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Brohan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Brown</surname> <given-names>N</given-names></string-name>, <string-name><surname>Carbajal</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chebotar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dabis</surname> <given-names>J</given-names></string-name>, <string-name><surname>Finn</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RT-1: robotics transformer for real-world control at scale</article-title>. <comment>arXiv:2212.06817. 2022</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dalal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chiruvolu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chaplot</surname> <given-names>D</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Plan-Seq-Learn: language model guided RL for solving long horizon robotics tasks</article-title>. <comment>arXiv:2405.01534. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2405.01534</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nian</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Discrete diffusion VLA: bringing discrete diffusion to action decoding in vision-language-action policies</article-title>. <comment>arXiv:2508.20072. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2508.20072</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hong</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ping</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Education robot object detection with a brain-inspired approach integrating Faster R-CNN, YOLOv3, and semi-supervised learning</article-title>. <source>Front Neurorobot</source>. <year>2024</year>;<volume>17</volume>:<fpage>1338104</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fnbot.2023.1338104</pub-id>; <pub-id pub-id-type="pmid">38239759</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alharthi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Noreen</surname> <given-names>I</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aljrees</surname> <given-names>T</given-names></string-name>, <string-name><surname>Riaz</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Innab</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Novel deep reinforcement learning based collision avoidance approach for path planning of robots in unknown environment</article-title>. <source>PLoS One</source>. <year>2025</year>;<volume>20</volume>(<issue>1</issue>):<fpage>e0312559</fpage>. doi:<pub-id pub-id-type="doi">10.1371/journal.pone.0312559</pub-id>; <pub-id pub-id-type="pmid">39821118</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Das</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ryoo</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Does self-supervised learning really improve reinforcement learning from pixels?</article-title> In: <conf-name>36th Conference on Neural Information Processing Systems (NeurIPS 2022); 2022 Nov 28&#x2013;Dec 9</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>30865</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.52202/068431-2238</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>D</given-names></string-name>, <string-name><surname>Mulyana</surname> <given-names>B</given-names></string-name>, <string-name><surname>Stankovic</surname> <given-names>V</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A survey on deep reinforcement learning algorithms for robotic manipulation</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>7</issue>):<fpage>3762</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23073762</pub-id>; <pub-id pub-id-type="pmid">37050822</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali Shahid</surname> <given-names>A</given-names></string-name>, <string-name><surname>Piga</surname> <given-names>D</given-names></string-name>, <string-name><surname>Braghin</surname> <given-names>F</given-names></string-name>, <string-name><surname>Roveda</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning</article-title>. <source>Auton Rob</source>. <year>2022</year>;<volume>46</volume>(<issue>3</issue>):<fpage>483</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10514-022-10034-z</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Orr</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dutta</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Multi-agent deep reinforcement learning for multi-robot applications: a survey</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>7</issue>):<fpage>3625</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23073625</pub-id>; <pub-id pub-id-type="pmid">37050685</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kasaura</surname> <given-names>K</given-names></string-name>, <string-name><surname>Miura</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kozuno</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yonetani</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hoshino</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hosoe</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Benchmarking actor-critic deep reinforcement learning algorithms for robotics control with action constraints</article-title>. <source>IEEE Robot Autom Lett</source>. <year>2023</year>;<volume>8</volume>(<issue>8</issue>):<fpage>4449</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1109/LRA.2023.3284378</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RL-100: performant robotic manipulation with real-world reinforcement learning</article-title>. <comment>arXiv:2510.14830. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2510.14830</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Figueiredo Prudencio</surname> <given-names>R</given-names></string-name>, <string-name><surname>Maximo</surname> <given-names>MROA</given-names></string-name>, <string-name><surname>Colombini</surname> <given-names>EL</given-names></string-name></person-group>. <article-title>A survey on offline reinforcement learning: taxonomy, review, and open problems</article-title>. <source>IEEE Trans Neural Netw Learning Syst</source>. <year>2024</year>;<volume>35</volume>(<issue>8</issue>):<fpage>10237</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2023.3250269</pub-id>; <pub-id pub-id-type="pmid">37030754</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>G&#x00FC;rtler</surname> <given-names>N</given-names></string-name>, <string-name><surname>Blaes</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kolev</surname> <given-names>P</given-names></string-name>, <string-name><surname>Widmaier</surname> <given-names>F</given-names></string-name>, <string-name><surname>W&#x00FC;thrich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bauer</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Benchmarking offline reinforcement learning on real-robot hardware</article-title>. <comment>arXiv:2307.15690. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2307.15690</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Janner</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Levine</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Offline reinforcement learning as one big sequence modeling problem</article-title>. <comment>arXiv:2106.02039. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2106.02039</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>F</given-names></string-name></person-group>. <article-title>HLifeRL: a hierarchical lifelong reinforcement learning framework</article-title>. <source>J King Saud Univ Comput Inf Sci</source>. <year>2022</year>;<volume>34</volume>(<issue>7</issue>):<fpage>4312</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jksuci.2022.05.001</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Morales</surname> <given-names>EF</given-names></string-name>, <string-name><surname>Murrieta-Cid</surname> <given-names>R</given-names></string-name>, <string-name><surname>Becerra</surname> <given-names>I</given-names></string-name>, <string-name><surname>Esquivel-Basaldua</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>A survey on deep learning and deep reinforcement learning in robotics with a tutorial on deep reinforcement learning</article-title>. <source>Intell Serv Robot</source>. <year>2021</year>;<volume>14</volume>(<issue>5</issue>):<fpage>773</fpage>&#x2013;<lpage>805</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11370-021-00398-z</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tomizuka</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Hierarchical planning through goal-conditioned offline reinforcement learning</article-title>. <source>IEEE Robot Autom Lett</source>. <year>2022</year>;<volume>7</volume>(<issue>4</issue>):<fpage>10216</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/lra.2022.3190100</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Antonyshyn</surname> <given-names>L</given-names></string-name>, <string-name><surname>Givigi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Deep model-based reinforcement learning for predictive control of robotic systems with dense and sparse rewards</article-title>. <source>J Intell Rob Syst</source>. <year>2024</year>;<volume>110</volume>(<issue>3</issue>):<fpage>100</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10846-024-02118-y</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Taniguchi</surname> <given-names>T</given-names></string-name>, <string-name><surname>Murata</surname> <given-names>S</given-names></string-name>, <string-name><surname>Suzuki</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ognibene</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lanillos</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ugur</surname> <given-names>E</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>World models and predictive coding for cognitive and developmental robotics: frontiers and challenges</article-title>. <source>Adv Robot</source>. <year>2023</year>;<volume>37</volume>(<issue>13</issue>):<fpage>780</fpage>&#x2013;<lpage>806</lpage>. doi:<pub-id pub-id-type="doi">10.1080/01691864.2023.2225232</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kong</surname> <given-names>LH</given-names></string-name>, <string-name><surname>He</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>WS</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>YN</given-names></string-name></person-group>. <article-title>Dynamic movement primitives based robot skills learning</article-title>. <source>Mach Intell Res</source>. <year>2023</year>;<volume>20</volume>(<issue>3</issue>):<fpage>396</fpage>&#x2013;<lpage>407</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11633-022-1346-z</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Navarro-Alarcon</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Human-in-the-loop robot learning for smart manufacturing: a human-centric perspective</article-title>. <source>IEEE Trans Automat Sci Eng</source>. <year>2025</year>;<volume>22</volume>:<fpage>11062</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tase.2025.3528051</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tiezzi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Apicella</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cardenas-Perez</surname> <given-names>C</given-names></string-name>, <string-name><surname>Fregonese</surname> <given-names>G</given-names></string-name>, <string-name><surname>Dafarra</surname> <given-names>S</given-names></string-name>, <string-name><surname>Morerio</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning to evaluate autonomous behaviour in human-robot interaction</article-title>. <comment>arXiv:2507.06404. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2507.06404</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Denecke</surname> <given-names>K</given-names></string-name>, <string-name><surname>Baudoin</surname> <given-names>CR</given-names></string-name></person-group>. <article-title>A review of artificial intelligence and robotics in transformed health ecosystems</article-title>. <source>Front Med</source>. <year>2022</year>;<volume>9</volume>:<fpage>795957</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fmed.2022.795957</pub-id>; <pub-id pub-id-type="pmid">35872767</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ktena</surname> <given-names>I</given-names></string-name>, <string-name><surname>Wiles</surname> <given-names>O</given-names></string-name>, <string-name><surname>Albuquerque</surname> <given-names>I</given-names></string-name>, <string-name><surname>Rebuffi</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Tanno</surname> <given-names>R</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>AG</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generative models improve fairness of medical classifiers under distribution shifts</article-title>. <source>Nat Med</source>. <year>2024</year>;<volume>30</volume>(<issue>4</issue>):<fpage>1166</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41591-024-02838-6</pub-id>; <pub-id pub-id-type="pmid">38600282</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Hamadani</surname> <given-names>MNA</given-names></string-name>, <string-name><surname>Fadhel</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Alzubaidi</surname> <given-names>L</given-names></string-name>, <string-name><surname>Balazs</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Reinforcement learning algorithms and applications in healthcare and robotics: a comprehensive and systematic review</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>8</issue>):<fpage>2461</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24082461</pub-id>; <pub-id pub-id-type="pmid">38676080</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Unleashing the potential of AI in modern healthcare: machine learning algorithms and intelligent medical robots</article-title>. <source>Res Intell Manuf Assem</source>. <year>2024</year>;<volume>3</volume>(<issue>1</issue>):<fpage>100</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.25082/rima.2024.01.002</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Masala</surname> <given-names>GL</given-names></string-name>, <string-name><surname>Giorgi</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Artificial intelligence and assistive robotics in healthcare services: applications in silver care</article-title>. <source>Int J Environ Res Public Health</source>. <year>2025</year>;<volume>22</volume>(<issue>5</issue>):<fpage>781</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ijerph22050781</pub-id>; <pub-id pub-id-type="pmid">40427894</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>ZM</given-names></string-name></person-group>. <article-title>Ethics and governance of trustworthy medical artificial intelligence</article-title>. <source>BMC Med Inform Decis Mak</source>. <year>2023</year>;<volume>23</volume>(<issue>1</issue>):<fpage>7</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s12911-023-02103-9</pub-id>; <pub-id pub-id-type="pmid">36639799</pub-id></mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Hsu</surname> <given-names>FY</given-names></string-name>, <string-name><surname>Lien</surname> <given-names>AS</given-names></string-name></person-group>. <article-title>Health care professionals&#x2019; perspectives of socially assistive robots in health care settings: systematic review</article-title>. <source>J Med Internet Res</source>. <year>2025</year>;<volume>27</volume>:<fpage>e79634</fpage>. doi:<pub-id pub-id-type="doi">10.2196/79634</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Silvera-Tawil</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Robotics in healthcare: a survey</article-title>. <source>SN Comput Sci</source>. <year>2024</year>;<volume>5</volume>(<issue>1</issue>):<fpage>189</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s42979-023-02551-0</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nicora</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Santangelo</surname> <given-names>G</given-names></string-name>, <string-name><surname>Billeci</surname> <given-names>L</given-names></string-name>, <string-name><surname>Aprile</surname> <given-names>IG</given-names></string-name>, <string-name><surname>Germanotta</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Systematic review of AI/ML applications in multi-domain robotic rehabilitation: trends, gaps, and future directions</article-title>. <source>J NeuroEng Rehabil</source>. <year>2025</year>;<volume>22</volume>(<issue>1</issue>):<fpage>79</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s12984-025-01605-z</pub-id>; <pub-id pub-id-type="pmid">40205472</pub-id></mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>InterPreT: interactive predicate learning from language feedback for generalizable task planning</article-title>. <comment>arXiv:2405.19758. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2405.19758</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kawaharazuka</surname> <given-names>K</given-names></string-name>, <string-name><surname>Matsushima</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gambardella</surname> <given-names>A</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Paxton</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Real-world robot applications of foundation models: a review</article-title>. <source>Adv Robot</source>. <year>2024</year>;<volume>38</volume>(<issue>18</issue>):<fpage>1232</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1080/01691864.2024.2408593</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>I</given-names></string-name>, <string-name><surname>Blukis</surname> <given-names>V</given-names></string-name>, <string-name><surname>Mousavian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tremblay</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ProgPrompt: program generation for situated robot task planning using large language models</article-title>. <source>Auton Robot</source>. <year>2023</year>;<volume>47</volume>(<issue>8</issue>):<fpage>999</fpage>&#x2013;<lpage>1012</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10514-023-10135-3</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kawaharazuka</surname> <given-names>K</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yamada</surname> <given-names>J</given-names></string-name>, <string-name><surname>Posner</surname> <given-names>I</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Vision-language-action models for robotics: a review towards real-world applications</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>162467</fpage>&#x2013;<lpage>504</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2025.3609980</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hoeller</surname> <given-names>D</given-names></string-name>, <string-name><surname>Rudin</surname> <given-names>N</given-names></string-name>, <string-name><surname>Sako</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hutter</surname> <given-names>M</given-names></string-name></person-group>. <article-title>ANYmal parkour: learning agile navigation for quadrupedal robots</article-title>. <source>Sci Robot</source>. <year>2024</year>;<volume>9</volume>(<issue>88</issue>):<fpage>eadi7566</fpage>. doi:<pub-id pub-id-type="doi">10.1126/scirobotics.adi7566</pub-id>; <pub-id pub-id-type="pmid">38478592</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Saad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>S</given-names></string-name>, <string-name><surname>Suhaib</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Hybrid framework for robotic manipulation: integrating reinforcement learning and large language models</article-title>. <comment>arXiv:2603.30022. 2026</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2603.30022</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Shao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Large VLM-based vision-language-action models for robotic manipulation: a survey</article-title>. <comment>arXiv:2508.13073. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2508.13073</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>ElMallah</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chhajer</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>CG</given-names></string-name></person-group>. <article-title>Score the steps, not just the goal: VLM-based subgoal evaluation for robotic manipulation</article-title>. <comment>arXiv:2509.19524. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2509.19524</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Badithela</surname> <given-names>A</given-names></string-name>, <string-name><surname>Snyder</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zha</surname> <given-names>L</given-names></string-name>, <string-name><surname>Mikhail</surname> <given-names>J</given-names></string-name>, <string-name><surname>O&#x2019;Kelly</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dixit</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Reliable and scalable robot policy evaluation with imperfect simulators</article-title>. <comment>arXiv:2510.04354. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2510.04354</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>C</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on large language model based autonomous agents</article-title>. <source>Front Comput Sci</source>. <year>2024</year>;<volume>18</volume>(<issue>6</issue>):<fpage>186345</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s11704-024-40231-1</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ahn</surname> <given-names>M</given-names></string-name>, <string-name><surname>Brohan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Brown</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chebotar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cortes</surname> <given-names>O</given-names></string-name>, <string-name><surname>David</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Do as I can, not as I say: grounding language in robotic affordances</article-title>. <comment>arXiv:2204.01691. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2204.01691</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rincon Arango</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Marco-Detchart</surname> <given-names>C</given-names></string-name>, <string-name><surname>Julian Inglada</surname> <given-names>VJ</given-names></string-name></person-group>. <article-title>Personalized cognitive support via social robots</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>3</issue>):<fpage>888</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25030888</pub-id>; <pub-id pub-id-type="pmid">39943528</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Raptis</surname> <given-names>EK</given-names></string-name>, <string-name><surname>Kapoutsis</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Kosmatopoulos</surname> <given-names>EB</given-names></string-name></person-group>. <article-title>RobotIQ: empowering mobile robots with human-level planning for real-world execution</article-title>. <source>Int J Adv Rob Syst</source>. <year>2026</year>;<volume>23</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1177/17298806261423235</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Solak</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lahr</surname> <given-names>GJG</given-names></string-name>, <string-name><surname>Ajoudani</surname> <given-names>A</given-names></string-name></person-group>. <article-title>SRL-VIC: a variable stiffness-based safe reinforcement learning for contact-rich robotic tasks</article-title>. <source>IEEE Robot Autom Lett</source>. <year>2024</year>;<volume>9</volume>(<issue>6</issue>):<fpage>5631</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/lra.2024.3396368</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Robotic control in adversarial and sparse reward environments: a robust goal-conditioned reinforcement learning approach</article-title>. <source>IEEE Trans Artif Intell</source>. <year>2024</year>;<volume>5</volume>(<issue>1</issue>):<fpage>244</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAI.2023.3237665</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Knoll</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep reinforcement learning based trajectory planning under uncertain constraints</article-title>. <source>Front Neurorobot</source>. <year>2022</year>;<volume>16</volume>:<fpage>883562</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fnbot.2022.883562</pub-id>; <pub-id pub-id-type="pmid">35586262</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>RM-RL: role-model reinforcement learning for precise robot manipulation</article-title>. <comment>arXiv:2510.15189. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2510.15189</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>