<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">83294</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.083294</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Graph-Based Constrained PPO for Low-Latency and Energy-Aware AI Agent Migration in Internet of Vehicular Agents</article-title>
<alt-title alt-title-type="left-running-head">Graph-Based Constrained PPO for Low-Latency and Energy-Aware AI Agent Migration in Internet of Vehicular Agents</alt-title>
<alt-title alt-title-type="right-running-head">Graph-Based Constrained PPO for Low-Latency and Energy-Aware AI Agent Migration in Internet of Vehicular Agents</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jiang</surname><given-names>Kanyang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kang</surname><given-names>Yingkai</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Ming</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>mingli@gdut.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Automation, Guangdong University of Technology and Key Laboratory of Intelligent Detection and the Internet of Things in Manufacturing, Ministry of Education</institution>, <addr-line>Guangzhou</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Automation, Guangdong University of Technology</institution>, <addr-line>Guangzhou</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Ming Li. Email: <email>mingli@gdut.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>61</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>04</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_83294.pdf"></self-uri>
<abstract>
<p>The Internet of Vehicular Agents (IoVA) interconnects distributed AI agents across vehicular networks to deliver real-time intelligent services for vehicular users. Due to the limited computing capacity of vehicles, AI agents are deployed on nearby RoadSide Units (RSUs) to perform computation-intensive inference. As vehicles traverse RSU coverage boundaries, AI agents must migrate to target RSUs to maintain service continuity. However, the communication and computing resources at each RSU are shared among multiple co-served vehicles, creating coupled allocation decisions that jointly determine system latency and energy consumption. To address this challenge, we propose a low-latency and energy-aware AI agent migration framework that models the end-to-end system latency and vehicle energy consumption in the IoVA. Since the cumulative nature of energy consumption introduces long-term constraints that cannot be handled by instantaneous optimization, we formulate the resource allocation problem as a constrained Markov decision process and develop a Graph-based Constrained Proximal Policy Optimization (GCPPO) algorithm to solve it. GCPPO employs a bidirectional graph attention network to extract the relational features between heterogeneous vehicles and RSUs, thereby enabling topology-aware resource allocation, and adopts a Lagrangian dual mechanism to adaptively enforce the long-term energy constraints. Simulation results demonstrate the effectiveness and scalability of the proposed algorithm, which achieves a <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>31.3</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> reduction in average system latency over baselines while attaining a <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mn>96.4</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> constraint satisfaction rate.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Internet of vehicular agents</kwd>
<kwd>AI agent migration</kwd>
<kwd>constrained deep reinforcement learning</kwd>
<kwd>graph attention network</kwd>
<kwd>resource allocation</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>2024 Guangdong Province Education Science Planning Project</funding-source>
<award-id>2024GXJK621</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding, complex reasoning, and content generation [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. AI agents leverage LLMs as their cognitive core, evolving from task-specific tools into autonomous entities capable of perceiving, reasoning, and acting across diverse domains [<xref ref-type="bibr" rid="ref-3">3</xref>]. In vehicular networks, AI agents are increasingly deployed to deliver real-time intelligent services [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. However, the rapidly changing network topology and fluctuating wireless channel conditions impose stringent requirements on service continuity [<xref ref-type="bibr" rid="ref-6">6</xref>]. The Internet of Vehicular Agents (IoVA) has emerged as a promising paradigm in which distributed AI agents are seamlessly interconnected and dynamically coordinated across the vehicular environment [<xref ref-type="bibr" rid="ref-7">7</xref>]. In the IoVA, AI agents continuously perceive the surrounding environment and user requests to construct real-time situational awareness. Leveraging this awareness, the AI agents formulate context-aware decisions and convert them into executable actions, thereby delivering intelligent services to vehicular users [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>However, AI agent decision-making relies on computation-intensive inference that far exceeds the limited computing capacity of vehicles [<xref ref-type="bibr" rid="ref-8">8</xref>]. To sustain the inference process, AI agents are deployed on RoadSide Units (RSUs) with sufficient computing resources [<xref ref-type="bibr" rid="ref-9">9</xref>]. As vehicles continuously traverse RSU coverage boundaries, AI agents must migrate to target RSUs to maintain service continuity [<xref ref-type="bibr" rid="ref-10">10</xref>]. The resulting service latency and energy consumption are jointly determined by the allocation of communication and computing resources at the RSUs. Although increasing resource allocation can effectively reduce service latency, it also raises energy consumption, which is bounded by strict vehicle energy budgets [<xref ref-type="bibr" rid="ref-11">11</xref>]. Therefore, it remains a significant challenge to optimize resource allocation for AI agent migration in the IoVA while jointly reducing service latency and satisfying vehicle energy constraints.</p>
<p>Traditional optimization methods for resource allocation in vehicular networks typically rely on accurate instantaneous channel state information and quasi-static network assumptions [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>]. Nevertheless, the high mobility of vehicles introduces significant channel estimation errors [<xref ref-type="bibr" rid="ref-14">14</xref>], and the dynamically changing network topology further limits the applicability of these methods [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. Deep Reinforcement Learning (DRL)-based methods offer a promising alternative by learning effective policies without prior knowledge of system dynamics [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Despite this advantage, most existing DRL approaches encode the environment state as a concatenated observation vector [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. This flat representation fails to capture the topological relationships among vehicles and RSUs, and generalizes poorly as the network scales. Furthermore, conventional DRL methods lack systematic constraint-handling mechanisms [<xref ref-type="bibr" rid="ref-18">18</xref>], making the learned policies prone to violating vehicle energy constraints in the highly dynamic IoVA.</p>
<p>To address the above challenges, at the system level, we develop a low-latency and energy-aware AI agent migration framework in the IoVA. The framework jointly models end-to-end service latency and vehicle energy consumption under dynamic channel conditions. We further formulate the multi-vehicle resource allocation problem as a Constrained Markov Decision Process (CMDP) with long-term cumulative energy budgets. At the algorithmic level, we design the Graph-based Constrained Proximal Policy Optimization (GCPPO) algorithm to learn an effective policy for the formulated CMDP. The main contributions of this paper are summarized as follows:<list list-type="bullet">
<list-item>
<p>We develop a low-latency and energy-aware AI agent migration framework in the IoVA, which jointly characterizes the end-to-end service latency and vehicle energy consumption in dynamic vehicular environments. Specifically, the latency model captures the latency across the communication, inference, and migration phases, while the energy model captures both the transmission and circuit power consumption of each vehicle over uplink and downlink channels.</p></list-item>
<list-item>
<p>We formulate the multi-vehicle resource allocation problem for AI agent migration as a Constrained Markov Decision Process (CMDP), which aims to minimize the long-term average system latency subject to cumulative energy constraints.</p></list-item>
<list-item>
<p>We design a novel GCPPO algorithm to solve the formulated CMDP. GCPPO leverages a bidirectional Graph Attention Network (GAT) to capture the relational features between heterogeneous vehicles and RSUs, and incorporates a Lagrangian dual method to adaptively enforce the long-term energy constraints. Simulation results demonstrate that GCPPO reduces average system latency by <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mn>31.3</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> compared with baselines while achieving a <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>96.4</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> constraint satisfaction rate.</p></list-item>
</list></p>
<p>The rest of the paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews the related work. <xref ref-type="sec" rid="s3">Section 3</xref> introduces the proposed low-latency and energy-aware AI agent migration framework in the IoVA. In <xref ref-type="sec" rid="s4">Section 4</xref>, we present the architecture of the GCPPO algorithm. <xref ref-type="sec" rid="s5">Section 5</xref> provides simulation results to demonstrate the performance of the GCPPO algorithm. In <xref ref-type="sec" rid="s6">Section 6</xref>, we conclude the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>AI Agents in Vehicular Networks</title>
<p>Driven by the rapid advancement of LLMs, researchers have increasingly explored integrating AI agents into vehicular networks [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. To address the substantial computing demand of AI agents, the authors in [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a cloud-edge collaborative architecture that distributed multimodal LLM inference between edge servers and the cloud for intelligent driver assistance. Furthermore, the authors in [<xref ref-type="bibr" rid="ref-22">22</xref>] developed an Agent-as-a-Service paradigm, where AI agents autonomously performed computing and communication tasks, effectively reducing service latency for edge-assisted autonomous driving.</p>
<p>Given the high mobility of vehicles, AI agent migration across edge servers has emerged as a promising approach to maintain service continuity [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. Specifically, the authors in [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed a generative diffusion-based contract design to incentivize RSU participation in AI agent migration. To address security threats during AI agent migration, the authors in [<xref ref-type="bibr" rid="ref-5">5</xref>] developed a secure online migration framework with trust assessment, effectively mitigating network attacks while maintaining low migration latency. However, the above works focus on incentive mechanisms and migration security, while overlooking vehicle energy consumption, which critically constrains vehicles with limited resources across successive migrations in the IoVA.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Constrained Deep Reinforcement Learning for Resource Optimization</title>
<p>Resource optimization in wireless networks typically involves long-term constraints (e.g., energy budgets and service quality guarantees) that conventional DRL methods fail to satisfy [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. Constrained DRL (CDRL) methods address this limitation by formulating the problem as a CMDP, which extends the MDP with cumulative cost constraints [<xref ref-type="bibr" rid="ref-25">25</xref>]. Among CDRL methods, Lagrangian dual optimization is the most widely adopted approach, which solves the CMDP by converting constraints into penalty terms in the objective function [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. For instance, the authors in [<xref ref-type="bibr" rid="ref-24">24</xref>] employed Lagrangian dual optimization to achieve near-optimal network capacity through joint UAV altitude control and channel access under energy harvesting constraints. Since Lagrangian dual methods may still produce infeasible actions during execution [<xref ref-type="bibr" rid="ref-18">18</xref>], recent studies have incorporated safety mechanisms to provide stronger constraint guarantees [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]. For instance, the authors in [<xref ref-type="bibr" rid="ref-28">28</xref>] embedded a safety layer that projects each action onto the feasible set to satisfy latency constraints for edge offloading. However, the above CDRL methods rely on flat state representations, which cannot capture the spatial relationships among interacting entities. When applied to IoVA scenarios with heterogeneous vehicles and RSUs under dynamic topologies, this limitation degrades the expressiveness and scalability of the learned policies.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Low-Latency and Energy-Aware AI Agent Migration Framework</title>
<p>In this section, we present the proposed low-latency and energy-aware AI agent migration framework in the IoVA, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The framework considers a system where AI agents are deployed on RSUs to perform computation-intensive inference, and migrate to target RSUs as vehicles traverse coverage boundaries. We first model the service latency and vehicle energy consumption, and then formulate the resource allocation problem.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Illustration of the proposed low-latency and energy-aware AI agent migration framework in the IoVA.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Service Latency Model</title>
<p>We consider an IoVA system that consists of a set <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> of <italic>V</italic> vehicles and a set <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>E</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> of <italic>E</italic> RSUs. Due to limited on-board computing capacity, AI agents are deployed on nearby RSUs for computation-intensive inference [<xref ref-type="bibr" rid="ref-9">9</xref>]. The service duration is discretized into <italic>T</italic> time slots. Each time slot is sufficiently short that the vehicle-RSU associations and channel conditions remain approximately constant [<xref ref-type="bibr" rid="ref-29">29</xref>]. At time slot <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>t</mml:mi></mml:math></inline-formula>, vehicle <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>v</mml:mi></mml:math></inline-formula> is associated with its serving RSU <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:math></inline-formula>, and the set of vehicles served by RSU <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>e</mml:mi></mml:math></inline-formula> is denoted by <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. The service process comprises uplink transmission, RSU-side inference, AI agent migration upon RSU handover, and downlink delivery.</p>
<p>In the communication phases, each vehicle exchanges data with its serving RSU through shared wireless bandwidth. Let <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msubsup><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>u</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:msubsup></mml:math></inline-formula> denote the total uplink and downlink bandwidth of RSU <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>e</mml:mi></mml:math></inline-formula>, and let <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the allocation ratios for vehicle <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>v</mml:mi></mml:math></inline-formula>, subject to <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Based on the Shannon-Hartley theorem, the uplink and downlink transmission rates can be expressed as
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:msub><mml:mi>log</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>e</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:msub><mml:mi>log</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the uplink transmit power of vehicle <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>v</mml:mi></mml:math></inline-formula> at time slot <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:msubsup></mml:math></inline-formula> is the fixed downlink transmit power of RSU <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>e</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the channel gain between vehicle <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>v</mml:mi></mml:math></inline-formula> and RSU <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>e</mml:mi></mml:math></inline-formula> at time slot <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>t</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>e</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> represent the noise power at the RSU and the vehicle, respectively. In the uplink phase, the vehicle transmits environmental perception data of size <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and user instructions of size <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. In the downlink phase, the RSU returns structured actions of size <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and textual outputs of size <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The uplink and downlink communication latencies can be obtained as
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the computation phase, the RSU allocates its computing resources to the AI agents for inference [<xref ref-type="bibr" rid="ref-8">8</xref>]. Specifically, the AI agent first encodes the multimodal perception data into a joint representation of the driving environment and task intent. This representation is then processed in parallel during prefill to construct the Key-Value (KV) cache for the session. Based on this cache, the AI agent performs autoregressive decoding to generate the response. The encoding, prefill, and decode stages process data volumes of <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively. Let <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the fraction of computing resources of RSU <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>e</mml:mi></mml:math></inline-formula> allocated to vehicle <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>v</mml:mi></mml:math></inline-formula>, subject to <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, and let <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>C</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula> denote the computational intensity in GPU cycles per unit of data and the processing speed of RSU <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>e</mml:mi></mml:math></inline-formula>, respectively. The resulting inference latency is expressed as
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Since the AI agent is deployed on the RSU, it must be migrated upon RSU handover to preserve service continuity [<xref ref-type="bibr" rid="ref-30">30</xref>]. Let <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the serving RSU at the previous time slot and define the migration indicator <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msubsup><mml:mi>&#x03B4;</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2260;</mml:mo><mml:mi>e</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, which equals <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mn>1</mml:mn></mml:math></inline-formula> if a handover occurs and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mn>0</mml:mn></mml:math></inline-formula> otherwise. As the vehicle moves, the AI agent processes newly received perception data and user instructions at each time slot, appending new entries to the KV cache at a rate of <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msubsup><mml:mi>S</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> per context token. The accumulated cache of <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msubsup><mml:mi>N</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> context tokens is compressed via quantization with coefficient <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>q</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> to reduce the transfer volume. Together with a fixed overhead <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> covering the model parameters, runtime environment, and protocol data, the migration latency over the inter-RSU link with bandwidth <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is given by
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>q</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Thus, the end-to-end service latency for vehicle <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>v</mml:mi></mml:math></inline-formula> at time slot <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>t</mml:mi></mml:math></inline-formula> is given by
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>D</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03B4;</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Energy Consumption Model</title>
<p>In the IoVA, vehicles operate on limited onboard batteries while RSUs are grid-powered, we model only the vehicle-side energy consumption during uplink and downlink communication [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
<p>In the uplink phase, the vehicle transmits data to the serving RSU through its radio frequency chain [<xref ref-type="bibr" rid="ref-29">29</xref>]. Due to the limited amplifier efficiency, the power amplifier draws input power that exceeds the intended transmit power. The active transmit chain further consumes static circuit power from the supporting circuitry and a bandwidth-dependent dynamic component that arises from baseband signal processing across the allocated bandwidth. The resulting uplink energy consumption can be expressed as
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:mfrac><mml:mo>+</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>u</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03BA;</mml:mi><mml:mi>u</mml:mi></mml:msup><mml:msubsup><mml:mi>B</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msubsup><mml:mi>&#x03B7;</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the power amplifier efficiency, <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>u</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the uplink static circuit power, and <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msup><mml:mi>&#x03BA;</mml:mi><mml:mi>u</mml:mi></mml:msup></mml:math></inline-formula> is the bandwidth-dependent circuit power coefficient for the uplink chain.</p>
<p>In the downlink phase, the vehicle receives the service response without power amplification. The receive chain consumes a demodulation power <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>d</mml:mi></mml:msubsup></mml:math></inline-formula>, a static circuit power <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and a bandwidth-dependent component <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msup><mml:mi>&#x03BA;</mml:mi><mml:mi>d</mml:mi></mml:msup><mml:msubsup><mml:mi>B</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, yielding the downlink energy consumption
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03BA;</mml:mi><mml:mi>d</mml:mi></mml:msup><mml:msubsup><mml:mi>B</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Thus, the total energy consumption of vehicle <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>v</mml:mi></mml:math></inline-formula> at time slot <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>t</mml:mi></mml:math></inline-formula> is given by
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>u</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Problem Formulation</title>
<p>We aim to optimize the resource allocation in the IoVA to minimize the cumulative service latency across all vehicles while satisfying per-vehicle energy budget constraints. However, increasing the transmit power of a vehicle reduces its latency but raises its energy consumption, while allocating more bandwidth to one vehicle limits the resources available to others. To balance these competing objectives, we jointly determine the computing resource allocation ratio <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the uplink and downlink bandwidth allocation ratios <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and the vehicle transmit power <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, collectively denoted as <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The optimization problem can be formulated as
<disp-formula id="eqn-9a"><label>(9a)</label><mml:math id="mml-eqn-9a" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mspace width="1em" /></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>D</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd><mml:mtd /></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9b"><label>(9b)</label><mml:math id="mml-eqn-9b" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mspace width="1em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9c"><label>(9c)</label><mml:math id="mml-eqn-9c" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mspace width="1em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9d"><label>(9d)</label><mml:math id="mml-eqn-9d" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mspace width="1em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9e"><label>(9e)</label><mml:math id="mml-eqn-9e" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mspace width="1em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9f"><label>(9f)</label><mml:math id="mml-eqn-9f" display="block"><mml:mtable columnalign="right center left" rowspacing="3pt" columnspacing="0 thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>u</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mspace width="1em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Constraints <xref ref-type="disp-formula" rid="eqn-9a">(9b</xref>,<xref ref-type="disp-formula" rid="eqn-9c">c)</xref> ensure that the computing and bandwidth allocation ratios assigned by each RSU do not exceed unity. Constraint <xref ref-type="disp-formula" rid="eqn-9d">(9d)</xref> bounds the uplink transmit power of each vehicle in its feasible range, where <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msubsup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup></mml:math></inline-formula> denotes the maximum transmit power. Constraint <xref ref-type="disp-formula" rid="eqn-9e">(9e)</xref> restricts the serving RSU of each vehicle to its communicable RSU set <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msubsup><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Constraint <xref ref-type="disp-formula" rid="eqn-9f">(9f)</xref> imposes that the cumulative energy consumption of each vehicle over the entire service period does not exceed its energy budget <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>u</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Graph-Based Constrained Proximal Policy Optimization Algorithm</title>
<sec id="s4_1">
<label>4.1</label>
<title>CMDP Formulation</title>
<p>The resource allocation problem in the IoVA involves a non-convex objective, temporally coupled energy constraints, and non-stationary system dynamics, which render conventional approaches intractable. We therefore reformulate it as a CMDP [<xref ref-type="bibr" rid="ref-25">25</xref>], characterized by the tuple <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow></mml:math></inline-formula> is the state space, <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow></mml:math></inline-formula> is the action space, <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>r</mml:mi></mml:math></inline-formula> is the reward function, <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>c</mml:mi></mml:math></inline-formula> is the cost function, and <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is the discount factor. The detailed definitions are as follows:
<list list-type="simple">
<list-item><label>(1)</label><p><italic>State space:</italic> At each time slot <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>t</mml:mi></mml:math></inline-formula>, the CDRL agent observes the system state <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mrow><mml:mi mathvariant="bold">s</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula> <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mrow><mml:mi mathvariant="bold">m</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>E</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which encompasses channel gains, RSU associations, service data volumes, KV cache sizes, and cumulative energy consumption. Here, <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mrow><mml:mi mathvariant="bold">m</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> collects all service data volumes of vehicle <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>v</mml:mi></mml:math></inline-formula> defined in <xref ref-type="sec" rid="s3_1">Section 3.1</xref>, and <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mrow><mml:mover><mml:mi>E</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> tracks the energy consumed up to time slot <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>t</mml:mi></mml:math></inline-formula>.</p></list-item>
<list-item><label>(2)</label><p><italic>Action space:</italic> At each time slot <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>t</mml:mi></mml:math></inline-formula>, the CDRL agent determines the action <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula> <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which jointly specifies the computing, bandwidth, and power allocation for each vehicle.</p></list-item>
<list-item><label>(3)</label><p><italic>Reward function:</italic> Since the goal is to minimize the cumulative service latency, the immediate reward is defined as <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>D</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>.</p></list-item>
<list-item><label>(4)</label><p><italic>Cost function:</italic> Among the constraints in <xref ref-type="disp-formula" rid="eqn-9a">(9)</xref>, constraints <xref ref-type="disp-formula" rid="eqn-9b">(9b</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-9d">d)</xref> are enforced via per-slot action projection, and constraint <xref ref-type="disp-formula" rid="eqn-9e">(9e)</xref> is determined by the physical distance between each vehicle and its nearest RSU. However, constraint <xref ref-type="disp-formula" rid="eqn-9f">(9f)</xref> couples decisions across all time slots. To encode this long-term constraint, we define the immediate cost as <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Accordingly, the CMDP optimization objective is to maximize the expected cumulative reward subject to the energy budget constraint <xref ref-type="disp-formula" rid="eqn-9f">(9f)</xref>, which can be formulated as</p></list-item></list></p>
<p><disp-formula id="eqn-10a"><label>(10a)</label><mml:math id="mml-eqn-10a" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:munder><mml:mspace width="1em" /></mml:mtd><mml:mtd><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi>r</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-10b"><label>(10b)</label><mml:math id="mml-eqn-10b" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /></mml:mtd><mml:mtd><mml:msubsup><mml:mi>J</mml:mi><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msub><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>u</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Architecture of the GCPPO Algorithm</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Bidirectional Graph Attention Network for State Representation</title>
<p>To capture the relational structure among vehicles and RSUs in the IoVA, we represent the system state as a bipartite graph <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Each vehicle node <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:math></inline-formula> carries a feature vector <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> that encodes its channel condition, service demands, and energy status, and each RSU node <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:math></inline-formula> carries <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula> that encodes its computing and communication capacity. The edge set <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> connects each vehicle to its serving RSU, which encodes the resource-sharing structure [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>The bipartite graph <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow></mml:math></inline-formula> is then processed by a bidirectional GAT [<xref ref-type="bibr" rid="ref-31">31</xref>]. The node features serve as the initial embeddings, i.e., <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula>. At layer <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>l</mml:mi></mml:math></inline-formula>, the bidirectional aggregations are given by
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> is the set of RSU nodes adjacent to vehicle <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>v</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the set of vehicles served by RSU <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>e</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> are direction-specific learnable weight matrices, and <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the activation function. The attention coefficient <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is computed as
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the target node, <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> denotes its neighboring nodes, <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> is a learnable attention vector, and the parameters <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> take direction-specific values <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for RSU-to-vehicle edges and <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for vehicle-to-RSU edges. The resulting vehicle embeddings thus reflect the resource availability of nearby RSUs, while the RSU embeddings capture the aggregate demand of served vehicles.</p>
<p>Since the aggregations in <xref ref-type="disp-formula" rid="eqn-11">(11)</xref> only access immediate neighbors, we stack <italic>L</italic> layers to expand the receptive field to <italic>L</italic>-hop neighbors, which enables each vehicle node to incorporate information from co-served vehicles that compete for the same RSU resources. After the final layer, we concatenate all vehicle node embeddings to obtain the state vector <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">h</mml:mi></mml:mrow><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which serves as the input to the actor and critic networks of the CDRL agent.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Graph-Based PPO for Energy-Constrained AI Agent Migration</title>
<p>The actor network <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo>&#x2223;</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> parameterizes a Gaussian policy over the continuous action space, where separate fully connected layers output the mean <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and standard deviation <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi mathvariant="bold-italic">&#x03C3;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. At time slot <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>t</mml:mi></mml:math></inline-formula>, the actor outputs a raw action <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which is then projected onto the feasible set defined by <xref ref-type="disp-formula" rid="eqn-9b">(9b</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-9d">d)</xref> prior to execution. Specifically, since constraints <xref ref-type="disp-formula" rid="eqn-9b">(9b</xref>,<xref ref-type="disp-formula" rid="eqn-9c">c)</xref> couple the per-RSU shares across co-served vehicles, we apply a per-RSU softmax over <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> that maps <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> onto the simplex via <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The bandwidth shares <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are obtained in the same way, while the per-vehicle transmit power is mapped via <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:msubsup><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to satisfy <xref ref-type="disp-formula" rid="eqn-9d">(9d)</xref>. The resulting projection is a deterministic mapping applied to the sampled action, thereby preserving the validity of the PPO policy gradient computed on <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and ensuring strict feasibility. To support the policy update, the cumulative reward and energy cost under the current policy are evaluated by a reward critic <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msubsup><mml:mi>V</mml:mi><mml:mi>&#x03D5;</mml:mi><mml:mi>r</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and a cost critic <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msubsup><mml:mi>V</mml:mi><mml:mi>&#x03C8;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively. Based on the cost critic estimates, we adopt Lagrangian relaxation [<xref ref-type="bibr" rid="ref-25">25</xref>] to enforce the long-term energy budget constraint in <xref ref-type="disp-formula" rid="eqn-10b">(10b)</xref>, introducing a non-negative Lagrange multiplier <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> for each vehicle <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mi>v</mml:mi></mml:math></inline-formula> to penalize energy budget violations. The Lagrangian objective and the corresponding per-step advantage can be expressed as
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>J</mml:mi><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>u</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>L</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>r</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:math></inline-formula> denotes the vector of Lagrange multipliers, and <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>r</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the reward and cost advantages estimated via Generalized Advantage Estimation (GAE) [<xref ref-type="bibr" rid="ref-32">32</xref>]. As <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> increases, the policy is steered away from energy-intensive actions toward more conservative resource allocation. To update the policy, we adopt the PPO clipped surrogate objective [<xref ref-type="bibr" rid="ref-33">33</xref>] with clipping parameter <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> and importance sampling ratio <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mi>&#x03C1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">a</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The GCPPO policy objective can be expressed as
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>C</mml:mi><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>O</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>L</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>clip</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mo>&#x03F5;</mml:mo><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mo>&#x03F5;</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>L</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Since both the policy objective and the constraint enforcement rely on accurate value estimates, the reward critic and cost critic are trained to minimize the mean squared error between their predictions and the empirical returns, with the critic loss functions given by
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>&#x03D5;</mml:mi><mml:mi>r</mml:mi></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>&#x03C8;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:msub><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the discounted cumulative reward and cost returns computed from the collected trajectories, respectively. After each policy update, the Lagrange multipliers are updated via dual gradient ascent based on the cost critic estimates as
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>&#x03BB;</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>J</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>u</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>&#x03BB;</mml:mi></mml:msub></mml:math></inline-formula> is the dual learning rate and <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msubsup><mml:mrow><mml:mover><mml:mi>J</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>c</mml:mi><mml:mi>v</mml:mi></mml:msubsup></mml:math></inline-formula> is the estimated cumulative energy cost of vehicle <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi>v</mml:mi></mml:math></inline-formula> under the current policy. The overall training procedure is presented in Algorithm 1, with per-iteration training complexity <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo>+</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mi>K</mml:mi><mml:mi>T</mml:mi><mml:mi>V</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-31">31</xref>] and per-step inference complexity <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo>+</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mi>V</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mi>d</mml:mi></mml:math></inline-formula> is the GAT embedding dimension, <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:msub><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> are the hidden dimensions of the actor and critic networks, and <italic>K</italic> is the number of policy update epochs.</p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-8.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Simulation Results</title>
<sec id="s5_1">
<label>5.1</label>
<title>Experimental Setup</title>
<p>We consider an IoVA system in which vehicles travel along a road segment at speeds of 30 to 120 km/h and are served by uniformly deployed RSUs. The initial positions of vehicles are randomly distributed along the road, and each vehicle maintains a constant speed throughout each episode. Channel gains between each vehicle and RSU combine a distance-dependent path loss with exponent <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>3.76</mml:mn></mml:math></inline-formula> and Rayleigh small-scale fading [<xref ref-type="bibr" rid="ref-14">14</xref>]. The uplink sensing and instruction data volumes range over <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>15</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.02</mml:mn><mml:mo>,</mml:mo><mml:mn>0.625</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> MB, respectively, while the downlink actuation and text response volumes range over <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.001</mml:mn><mml:mo>,</mml:mo><mml:mn>0.05</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.004</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> MB, respectively. The inference data volumes range over <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>6.25</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> MB across the encoding, prefill, and decoding stages, and the context token counts range over <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mo stretchy="false">[</mml:mo><mml:mn>1024</mml:mn><mml:mo>,</mml:mo><mml:mn>8192</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The detailed simulation parameters are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Simulation parameters.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="2"><italic>Environment settings</italic></td>
<td align="center" colspan="2"><italic>GCPPO hyperparameters</italic></td>
</tr>
<tr>
<td>Number of vehicles</td>
<td><inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mo stretchy="false">[</mml:mo><mml:mn>20</mml:mn><mml:mo>,</mml:mo><mml:mn>100</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula></td>
<td>Number of GAT layers</td>
<td>2</td>
</tr>
<tr>
<td>Number of RSUs</td>
<td>6</td>
<td>GAT embedding dimension</td>
<td>128</td>
</tr>
<tr>
<td>Uplink and downlink bandwidth</td>
<td><inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mo stretchy="false">[</mml:mo><mml:mn>20</mml:mn><mml:mo>,</mml:mo><mml:mn>120</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> MHz</td>
<td>Hidden dimension of actor and critic</td>
<td>128</td>
</tr>
<tr>
<td>Max vehicle transmit power</td>
<td><inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> W</td>
<td>Actor learning rate</td>
<td><inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mn>2.4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>RSU downlink transmit power</td>
<td>1.8 W</td>
<td>Critic learning rate</td>
<td><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Receiver noise power</td>
<td><inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mo>&#x2212;</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> dBm</td>
<td>Discount factor</td>
<td>0.99</td>
</tr>
<tr>
<td>RSU processing speed</td>
<td><inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mn>7.8</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> cycles/s</td>
<td>PPO clipping parameter</td>
<td>0.2</td>
</tr>
<tr>
<td>Inter-RSU link bandwidth</td>
<td>500 MHz</td>
<td>Policy update epochs per episode</td>
<td>3</td>
</tr>
<tr>
<td>Energy constraint threshold</td>
<td><inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>160</mml:mn><mml:mo>,</mml:mo><mml:mn>200</mml:mn><mml:mo>,</mml:mo><mml:mn>240</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> J</td>
<td>Lagrange multiplier learning rate</td>
<td>0.035</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For the GCPPO algorithm, the actor and critic networks share a two-layer bidirectional GAT encoder. All algorithms are implemented in PyTorch and trained for <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>5</mml:mn></mml:msup></mml:math></inline-formula> environment steps on an NVIDIA RTX 4090 GPU with <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mn>5</mml:mn></mml:math></inline-formula> random seeds. Each algorithm&#x2019;s final policy is evaluated on <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mn>20</mml:mn></mml:math></inline-formula> test episodes per seed, and the mean and standard deviation of each performance metric are computed over the resulting <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mn>100</mml:mn></mml:math></inline-formula> episodes. The GCPPO algorithm completes training in <inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mn>17.87</mml:mn></mml:math></inline-formula> min. During inference, it takes <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mn>1.44</mml:mn></mml:math></inline-formula> ms on average to generate a joint resource allocation decision.</p>
<p>To validate the effectiveness of the proposed GCPPO algorithm in the IoVA, we compare it with the following baseline algorithms:<list list-type="bullet">
<list-item>
<p><bold>CPO [<xref ref-type="bibr" rid="ref-34">34</xref>]:</bold> Constrained Policy Optimization (CPO) solves the CMDP via trust region updates with linearized cost constraints, providing first-order feasibility guarantees.</p></list-item>
<list-item>
<p><bold>PPO-Lag [<xref ref-type="bibr" rid="ref-35">35</xref>]:</bold> PPO-Lagrangian (PPO-Lag) augments the standard PPO objective with a Lagrangian penalty for energy constraint enforcement. It serves as an ablation variant of GCPPO without the bidirectional GAT encoder, isolating the contribution of the graph-based state representation.</p></list-item>
<list-item>
<p><bold>SAC-Lag [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]:</bold> Soft Actor-Critic with Lagrangian (SAC-Lag) extends the entropy-regularized off-policy framework with dual variables for cost constraint satisfaction.</p></list-item>
<list-item>
<p><bold>DAPA:</bold> Demand-Aware Proportional Allocation (DAPA) is a heuristic baseline that allocates computing and bandwidth resources at each RSU in proportion to the service demand of served vehicles, and adjusts each vehicle&#x2019;s transmit power according to its demand priority and residual energy.</p></list-item>
<list-item>
<p><bold>Random:</bold> Random uniformly samples resource allocation decisions from the feasible action space.</p></list-item>
</list></p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Performance Evaluation</title>
<p>In <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, we present the training performance of GCPPO and the baseline algorithms. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, GCPPO achieves the highest converged reward and converges in <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>5</mml:mn></mml:msup></mml:math></inline-formula> training steps, substantially faster than all baselines. Among the constrained baselines, CPO and PPO-Lag plateau at lower rewards, while SAC-Lag exhibits high variance due to the instability of off-policy Lagrangian updates. <xref ref-type="fig" rid="fig-2">Fig. 2b</xref> illustrates the cost curves under an energy budget of <inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mn>200</mml:mn></mml:math></inline-formula> J. GCPPO maintains its cost consistently below the energy constraint after initial training, whereas CPO converges below the threshold but adopts a more conservative policy that limits reward improvement. In contrast, SAC-Lag frequently exceeds the budget. <xref ref-type="fig" rid="fig-2">Fig. 2c</xref> depicts the convergence trajectories in the reward-cost plane, where GCPPO converges to the upper-left corner of the feasible region, achieving the highest reward while satisfying the energy budget. The superior performance of GCPPO is attributed to the combined effect of the bidirectional GAT, which improves reward through topology-aware allocation, and the Lagrangian dual mechanism, which constrains the trajectory in the feasible region.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Comparison of training performance among GCPPO and baseline algorithms. (<bold>a</bold>) Reward curves for different algorithms. (<bold>b</bold>) Cost curves for different algorithms under the energy constraint threshold. (<bold>c</bold>) Convergence trajectories of different algorithms in the reward-cost plane.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-2.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, we present the evaluation results on independent test episodes. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>, GCPPO achieves the lowest average system latency, reducing it by <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mn>36.0</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> over SAC-Lag and by approximately <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mn>4</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> over CPO and PPO-Lag. <xref ref-type="fig" rid="fig-3">Fig. 3b</xref> shows that GCPPO matches CPO in energy consumption while consuming <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mn>20.9</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> less than SAC-Lag. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3c</xref>, GCPPO achieves a constraint satisfaction rate of <inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mn>96.4</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, substantially outperforming PPO-Lag and SAC-Lag. Although DAPA attains the highest satisfaction rate of <inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mn>99.2</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, this comes at a <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mn>76.6</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> increase in system latency. These results indicate that the bidirectional GAT encoder enables GCPPO to exploit the vehicle-RSU topology for effective resource allocation, while the Lagrangian dual mechanism maintains energy feasibility without sacrificing latency performance.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Performance evaluation of GCPPO and baseline algorithms for low-latency and energy-aware AI agent migration optimization. (<bold>a</bold>) Average system latency. (<bold>b</bold>) Average energy consumption. (<bold>c</bold>) Constraint satisfaction rate.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-3.tif"/>
</fig>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Sensitivity Analysis</title>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the impact of the energy budget on GCPPO in the IoVA. As the budget increases from <inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mn>160</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mn>240</mml:mn></mml:math></inline-formula> J, the converged reward improves from approximately <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mo>&#x2212;</mml:mo><mml:mn>1500</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:mo>&#x2212;</mml:mo><mml:mn>1200</mml:mn></mml:math></inline-formula>, and the constraint satisfaction rate rises from <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:mn>90.8</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mn>98.1</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. This is because a more relaxed budget allows the CDRL agent to allocate higher transmit power for latency reduction without exceeding the energy limit. Under the tightest budget of <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:mn>160</mml:mn></mml:math></inline-formula> J, the Lagrangian multipliers automatically increase, yielding a more energy-conservative policy that trades higher latency for constraint satisfaction.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Reward curves and constraint satisfaction rates of the GCPPO algorithm under different energy budgets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-4.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-5">Figs. 5</xref> and <xref ref-type="fig" rid="fig-6">6</xref> illustrate the impact of wireless bandwidth and vehicle transmit power in the IoVA, respectively. As bandwidth increases from <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:mn>20</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:mn>120</mml:mn></mml:math></inline-formula> MHz, both system latency and energy consumption decrease across all algorithms, since higher bandwidth improves the transmission rate, thereby reducing communication time and the associated energy expenditure. Conversely, increasing the transmit power from <inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mn>0.1</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:mn>0.5</mml:mn></mml:math></inline-formula> W reduces system latency but simultaneously raises energy consumption. Under both parameter sweeps, GCPPO consistently achieves the lowest normalized average latency, outperforming the strongest baseline by approximately <inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mn>12</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> in the bandwidth sweep. This advantage grows under resource-scarce conditions, where topology-aware allocation via the bidirectional GAT yields the largest gains. In contrast, SAC-Lag consumes approximately <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:mn>80</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> more normalized energy than GCPPO in the power sweep, indicating that its off-policy Lagrangian enforcement cannot effectively regulate the growing energy cost.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Impact of wireless bandwidth on the performance of GCPPO and baseline algorithms. (<bold>a</bold>) System latency curves and normalized average latency. (<bold>b</bold>) Energy consumption curves and normalized average energy consumption.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Impact of vehicle transmit power on the performance of GCPPO and baseline algorithms. (<bold>a</bold>) System latency curves and normalized average latency. (<bold>b</bold>) Energy consumption curves and normalized average energy consumption.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-6.tif"/>
</fig>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Scalability Analysis</title>
<p>In <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, we evaluate the scalability of all algorithms as the number of vehicles increases from <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mn>20</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:mn>100</mml:mn></mml:math></inline-formula>. We quantify it using the scalability exponent <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, estimated via least-squares linear regression in log-log space, where per-vehicle metrics scale as <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mi>&#x03B1;</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> with the number of vehicles <italic>V</italic>. An exponent below <inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mn>1.0</mml:mn></mml:math></inline-formula> indicates sub-linear growth, meaning that per-vehicle metrics grow more slowly than the network size.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Impact of the number of vehicles on the performance of GCPPO and baseline algorithms, where the scalability exponent is estimated via least-squares linear regression in log-log space. (<bold>a</bold>) Per-vehicle latency curves and scalability exponent. (<bold>b</bold>) Per-vehicle energy consumption curves and scalability exponent.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83294-fig-7.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="fig-7">Fig. 7a</xref>, GCPPO achieves the lowest latency scalability exponent of <inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mn>0.91</mml:mn></mml:math></inline-formula>, indicating sub-linear growth in per-vehicle latency as the network expands. CPO, PPO-Lag, and SAC-Lag also achieve sub-linear exponents, while DAPA and Random exhibit super-linear growth. For energy consumption, <xref ref-type="fig" rid="fig-7">Fig. 7b</xref> shows that GCPPO attains the lowest exponent of <inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mn>0.60</mml:mn></mml:math></inline-formula>, significantly outperforming SAC-Lag and Random, both of which approach or exceed linear growth. The sub-linear scaling of GCPPO is attributed to the bidirectional GAT encoder, which generalizes allocation decisions to larger vehicle populations, and the Lagrangian multiplier updates, which adaptively tighten the energy penalty as network density increases.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this paper, we have proposed a low-latency and energy-aware AI agent migration framework in the IoVA, jointly characterizing the end-to-end service latency and vehicle energy consumption across communication, inference, and migration phases. To solve the formulated CMDP, we have designed the GCPPO algorithm, which leverages a bidirectional GAT encoder to capture the relational structure among vehicles and RSUs, thereby enabling topology-aware resource allocation. It further incorporates a Lagrangian dual mechanism to adaptively enforce the long-term energy constraints without requiring predefined penalty weights. Simulation results demonstrate the effectiveness and scalability of GCPPO, which achieves a <inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mn>31.3</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> reduction in average system latency over baselines while attaining a <inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mn>96.4</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> constraint satisfaction rate. In future work, we plan to extend the proposed framework to multi-agent scenarios where distributed CDRL agents collaboratively optimize AI agent migration across heterogeneous edge networks.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the 2024 Guangdong Province Education Science Planning Project (Higher Education Special Project) under Grant 2024GXJK621.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Kanyang Jiang, Yingkai Kang and Ming Li; methodology, Kanyang Jiang and Yingkai Kang; software, Kanyang Jiang and Yingkai Kang; validation, Kanyang Jiang, Yingkai Kang and Ming Li; formal analysis, Kanyang Jiang and Yingkai Kang; investigation, Kanyang Jiang and Yingkai Kang; resources, Ming Li; data curation, Kanyang Jiang and Yingkai Kang; writing&#x2014;original draft preparation, Kanyang Jiang and Yingkai Kang; writing&#x2014;review and editing, Ming Li; visualization, Kanyang Jiang and Yingkai Kang; supervision, Ming Li; project administration, Ming Li. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on evaluation of large language models</article-title>. <source>ACM Trans Intell Syst Technol</source>. <year>2024</year>;<volume>15</volume>(<issue>3</issue>):<fpage>39</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3641289</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Plaat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wong</surname> <given-names>A</given-names></string-name>, <string-name><surname>Verberne</surname> <given-names>S</given-names></string-name>, <string-name><surname>Broekens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Van Stein</surname> <given-names>N</given-names></string-name>, <string-name><surname>B&#x00E4;ck</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Multi-step reasoning with large language models, a survey</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>58</volume>(<issue>6</issue>):<fpage>160</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3774896</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Pei</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chawla</surname> <given-names>NV</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Large language model based multi-agents: a survey of progress and challenges</article-title>. In: <conf-name>Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24</conf-name>. <publisher-loc>CA, USA</publisher-loc>: <publisher-name>International Joint Conferences on Artificial Intelligence Organization</publisher-name>; <year>2024</year>. p. <fpage>8048</fpage>&#x2013;<lpage>57</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mahmud</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hajmohamed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Almentheri</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alqaydi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aldhaheri</surname> <given-names>L</given-names></string-name>, <string-name><surname>Khalil</surname> <given-names>RA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Integrating LLMs with ITS: recent advances, potentials, challenges, and future directions</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2025</year>;<volume>26</volume>(<issue>5</issue>):<fpage>5674</fpage>&#x2013;<lpage>709</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2025.3528116</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Defending against network attacks for secure AI agent migration in vehicular metaverses</article-title>. <source>IEEE Internet Things J</source>. <year>2026</year>;<volume>13</volume>(<issue>3</issue>):<fpage>4153</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2025.3633501</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Clancy</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mullins</surname> <given-names>D</given-names></string-name>, <string-name><surname>Deegan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Horgan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ward</surname> <given-names>E</given-names></string-name>, <string-name><surname>Eising</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Wireless access for V2X communications: research, challenges and opportunities</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2024</year>;<volume>26</volume>(<issue>3</issue>):<fpage>2082</fpage>&#x2013;<lpage>119</lpage>. doi:<pub-id pub-id-type="doi">10.1109/COMST.2024.3384132</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Luan</surname> <given-names>TH</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Internet of agents: fundamentals, applications, and challenges</article-title>. <source>IEEE Trans Cognit Commun Netw</source>. <year>2026</year>;<volume>12</volume>:<fpage>4476</fpage>&#x2013;<lpage>501</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCCN.2025.3623369</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A review on edge large language models: design, execution, and applications</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>8</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3719664</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Mobile edge intelligence for large language models: a contemporary survey</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2025</year>;<volume>27</volume>(<issue>6</issue>):<fpage>3820</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.36227/techrxiv.172115025.57884352/v1</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Min</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ning</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Mobility-aware seamless service migration and resource allocation in multi-edge IoV systems</article-title>. <source>IEEE Trans Mob Comput</source>. <year>2025</year>;<volume>24</volume>(<issue>7</issue>):<fpage>6315</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TMC.2025.3540407</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qiu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Deep reinforcement learning-based adaptive computation offloading and power allocation in vehicular edge computing networks</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2024</year>;<volume>25</volume>(<issue>10</issue>):<fpage>13339</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2024.3391831</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shui</surname> <given-names>T</given-names></string-name>, <string-name><surname>Saad</surname> <given-names>W</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Resilient vehicular communications under imperfect channel state information</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2026</year>;<volume>25</volume>:<fpage>6442</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1109/twc.2025.3625199</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep reinforcement learning for multi-objective resource allocation in multi-platoon cooperative vehicular networks</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2023</year>;<volume>22</volume>(<issue>9</issue>):<fpage>6185</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/twc.2023.3240425</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>G</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Joint spectrum and power allocation for V2X communications with imperfect CSI</article-title>. <source>IEEE Trans Vehic Technol</source>. <year>2023</year>;<volume>72</volume>(<issue>12</issue>):<fpage>16338</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tvt.2023.3299691</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ju</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Pei</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Joint secure offloading and resource allocation for vehicular edge computing network: a multi-agent deep reinforcement learning approach</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2023</year>;<volume>24</volume>(<issue>5</issue>):<fpage>5555</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2023.3242997</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Iqbal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Al-Dulaimi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chih-Lin</surname> <given-names>I</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep reinforcement learning-based task scheduling and resource allocation for vehicular edge computing: a survey</article-title>. <source>IEEE Tran Intell Transp Syst</source>. <year>2025</year>;<volume>26</volume>(<issue>12</issue>):<fpage>21472</fpage>&#x2013;<lpage>501</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2025.3607910</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Graph neural networks and deep reinforcement learning-based resource allocation for V2X communications</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>4</issue>):<fpage>3613</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3469547</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wachi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sui</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A survey of constraint formulations in safe reinforcement learning</article-title>. In: <conf-name>Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24</conf-name>. <publisher-loc>CA, USA</publisher-loc>: <publisher-name>International Joint Conferences on Artificial Intelligence Organization</publisher-name>; <year>2024</year>. p. <fpage>8262</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>K</given-names></string-name>, <string-name><surname>Du</surname> <given-names>H</given-names></string-name>, <string-name><surname>Niyato</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generative AI-enabled vehicular networks: fundamentals, framework, and case study</article-title>. <source>IEEE Netw</source>. <year>2024</year>;<volume>38</volume>(<issue>4</issue>):<fpage>259</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MNET.2024.3391767</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>R</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guizani</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>GAI-IoV: bridging generative AI and vehicular networks for ubiquitous edge intelligence</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2024</year>;<volume>23</volume>(<issue>10</issue>):<fpage>12799</fpage>&#x2013;<lpage>814</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TWC.2024.3396276</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A cloud-edge collaborative architecture for multimodal LLM-based advanced driver assistance systems in IoT networks</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>10</issue>):<fpage>13208</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2024.3509628</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Agent-as-a-service: an AI-native edge computing framework for 6G networks</article-title>. <source>IEEE Netw</source>. <year>2025</year>;<volume>39</volume>(<issue>2</issue>):<fpage>44</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mnet.2024.3520987</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>D</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Niyato</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generative diffusion-based contract design for efficient AI twin migration in vehicular embodied AI networks</article-title>. <source>IEEE Trans Mobile Comput</source>. <year>2025</year>;<volume>24</volume>(<issue>5</issue>):<fpage>4573</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmc.2025.3526230/mm1</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khairy</surname> <given-names>S</given-names></string-name>, <string-name><surname>Balaprakash</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>LX</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Constrained deep reinforcement learning for energy sustainable multi-UAV based random access IoT networks with NOMA</article-title>. <source>IEEE J Selected Areas Commun</source>. <year>2021</year>;<volume>39</volume>(<issue>4</issue>):<fpage>1101</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSAC.2020.3018804</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Altman</surname> <given-names>E</given-names></string-name></person-group>. <source>Constrained markov decision processes</source>. <publisher-loc>Abingdon, UK</publisher-loc>: <publisher-name>Routledge</publisher-name>; <year>1999</year>. doi:<pub-id pub-id-type="doi">10.1201/9781315140223</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koursioumpas</surname> <given-names>N</given-names></string-name>, <string-name><surname>Magoula</surname> <given-names>L</given-names></string-name>, <string-name><surname>Petropouleas</surname> <given-names>N</given-names></string-name>, <string-name><surname>Thanopoulos</surname> <given-names>AI</given-names></string-name>, <string-name><surname>Panagea</surname> <given-names>T</given-names></string-name>, <string-name><surname>Alonistioti</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A safe deep reinforcement learning approach for energy efficient federated learning in wireless communication networks</article-title>. <source>IEEE Trans Green Commun Netw</source>. <year>2024</year>;<volume>8</volume>(<issue>4</issue>):<fpage>1862</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TGCN.2024.3372695</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Constrained reinforcement-learning-enabled policies with augmented lagrangian for cooperative intersection management</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>5</issue>):<fpage>5396</fpage>&#x2013;<lpage>411</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2024.3487854</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Safety-critical offloading with constrained reinforcement learning for multi-access edge computing</article-title>. <source>ACM Trans Sens Netw</source>. <year>2025</year>;<volume>21</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3715695</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jeong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Energy-efficient vehicular edge computing with one-by-one access scheme</article-title>. <source>IEEE Wirel Commun Lett</source>. <year>2024</year>;<volume>13</volume>(<issue>1</issue>):<fpage>39</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/LWC.2023.3318632</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Du</surname> <given-names>H</given-names></string-name>, <string-name><surname>Niyato</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hybrid-generative diffusion models for attack-oriented twin migration in vehicular metaverses</article-title>. <source>IEEE Trans Vehic Technol</source>. <year>2025</year>;<volume>74</volume>(<issue>9</issue>):<fpage>14720</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tvt.2025.3566034</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Velickovic</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cucurull</surname> <given-names>G</given-names></string-name>, <string-name><surname>Casanova</surname> <given-names>A</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li&#x00F2;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph attention networks</article-title>. <comment>arXiv:1710.10903. 2018</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Schulman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Moritz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Levine</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jordan</surname> <given-names>MI</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>High-dimensional continuous control using generalized advantage estimation</article-title>. <comment>arXiv:1506.02438. 2016</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Schulman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wolski</surname> <given-names>F</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Klimov</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Proximal policy optimization algorithms</article-title>. <comment>arXiv:1707.06347. 2017</comment>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Achiam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Held</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tamar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Constrained policy optimization</article-title>. In: <conf-name>Proceedings of the 34th International Conference on Machine Learning</conf-name>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2017</year>. p. <fpage>22</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ray</surname> <given-names>A</given-names></string-name>, <string-name><surname>Achiam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Amodei</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Benchmarking safe exploration in deep reinforcement learning</article-title>. <comment>arXiv:1910.01708. 2019</comment>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Haarnoja</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name>, <string-name><surname>Levine</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor</article-title>. In: <conf-name>Proceedings of the 35th International Conference on Machine Learning</conf-name>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2018</year>. p. <fpage>1861</fpage>&#x2013;<lpage>70</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>