<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">51408</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.051408</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Distributed Resource Allocation in Dispersed Computing Environment Based on UAV Track Inspection in Urban Rail Transit</article-title>
<alt-title alt-title-type="left-running-head">Distributed Resource Allocation in Dispersed Computing Environment Based on UAV Track Inspection in Urban Rail Transit</alt-title>
<alt-title alt-title-type="right-running-head">Distributed Resource Allocation in Dispersed Computing Environment Based on UAV Track Inspection in Urban Rail Transit</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Gan</surname><given-names>Tong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Dong</surname><given-names>Shuo</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Shiyou</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Jiaxin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>ll_jiaxin@163.com</email></contrib>
<aff id="aff-1"><label>1</label><institution>Division of Consulting</institution>, <institution>Beijing Metro Consultancy Corporation Ltd</institution>., <addr-line>Beijing,</addr-line> <addr-line>100037</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer and Communication Engineering, University of Science and Technology Beijing</institution>, <addr-line>Beijing, 100083</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jiaxin Li. Email: <email>ll_jiaxin@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>18</day><month>7</month><year>2024</year></pub-date>
<volume>80</volume>
<issue>1</issue>
<fpage>643</fpage>
<lpage>660</lpage>
<history>
<date date-type="received">
<day>05</day>
<month>3</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>5</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Gan et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Gan et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_51408.pdf"></self-uri>
<abstract>
<p>With the rapid development of urban rail transit, the existing track detection has some problems such as low efficiency and insufficient detection coverage, so an intelligent and automatic track detection method based on UAV is urgently needed to avoid major safety accidents. At the same time, the geographical distribution of IoT devices results in the inefficient use of the significant computing potential held by a large number of devices. As a result, the Dispersed Computing (DCOMP) architecture enables collaborative computing between devices in the Internet of Everything (IoE), promotes low-latency and efficient cross-wide applications, and meets users&#x2019; growing needs for computing performance and service quality. This paper focuses on examining the resource allocation challenge within a dispersed computing environment that utilizes UAV inspection tracks. Furthermore, the system takes into account both resource constraints and computational constraints and transforms the optimization problem into an energy minimization problem with computational constraints. The Markov Decision Process (MDP) model is employed to capture the connection between the dispersed computing resource allocation strategy and the system environment. Subsequently, a method based on Double Deep Q-Network (DDQN) is introduced to derive the optimal policy. Simultaneously, an experience replay mechanism is implemented to tackle the issue of increasing dimensionality. The experimental simulations validate the efficacy of the method across various scenarios.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>UAV track inspection</kwd>
<kwd>dispersed computing</kwd>
<kwd>resource allocation</kwd>
<kwd>deep reinforcement learning</kwd>
<kwd>Markov decision process</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Beijing Metro Consultancy Corporation Ltd</funding-source>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Urban rail transit, serving as a crucial national infrastructure, holds significant significance in the economic development of a nation. As urban rail transit expands swiftly, the density of road networks has witnessed a continuous rise, safety problems along urban rail transit lines are particularly prominent. Therefore, timely troubleshooting of equipment failures and potential safety hazards along the line is becoming more and more important [<xref ref-type="bibr" rid="ref-1">1</xref>]. To guarantee the safety of urban rail transit operations, concerned urban rail transit entities perform regular inspections of various professions within the urban rail transit system. Many recent urban rail transit safety accidents can be attributed to disasters occurring within the operating environment. Therefore, it is of utmost importance to perform efficient inspections of the urban rail transit operating environment and accurately identify potential hazards and risks. This is critical for ensuring the safety of urban rail transit operations.</p>
<p>The current inspection techniques for urban rail transit primarily involve manual inspections and inspections conducted using rail inspection vehicles. Manual inspection refers to the routine inspection of equipment along the line by using visual inspection or simple inspection tools after professional training. The main feature of rail inspection vehicle inspection is to load the inspection equipment on the rail inspection vehicle and use image processing technology, ultrasonic detection, and laser detection technology to detect faults in key equipment and facilities such as urban rail transit works and power supply. Combining the working principles of the two inspection methods, it can be seen that the two generally have problems such as low efficiency, poor night inspection conditions, low inspection frequency, constraints of comprehensive maintenance skylights, narrow inspection areas, and poor automated analysis capabilities. Moreover, even if various types of inspections are used repeatedly, there are still many dead ends of inspection, and it is difficult to achieve full coverage of inspections. Unmanned Aerial Vehicles (UAVs) offer benefits such as excellent flight maneuverability, low individual flight expenses, extensive surveillance capabilities, and the ability to operate independently of train operations. Using UAVs for urban rail transit inspections can effectively solve the drawbacks of existing methods. It is especially suitable as a supplement and replacement for the current inspection methods of urban rail transit. At present, the UAV inspection method has been widely used in power equipment inspection, engineering survey, geological exploration, fire early warning, highway inspection, and bridge inspection, and has achieved ideal results [<xref ref-type="bibr" rid="ref-2">2</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>], gradually becoming an irreplaceable inspection method. In particular, the power industry has begun to use UAV inspection methods and systems on a large scale. Inspection schemes, route planning, and inspection safety guarantees are becoming more and more mature. It is also necessary to use UAVs to inspect urban rail transit lines and operating environments, it will become a trend [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Simultaneously, as industrial technology advances rapidly, smart devices are steadily gaining increased computing and communication capabilities while also becoming smaller in size. IoT devices have seamlessly integrated into our daily lives, forming an essential component of our infrastructure [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. However, due to the geographical distribution of equipment (i.e., UAV) is too scattered, computing resources have not been fully utilized. The researchers propose to combine the scattered resources through the network to build a scalable, secure, and robust new system driven by global computing tasks [<xref ref-type="bibr" rid="ref-12">12</xref>]. Nevertheless, the efficient management of extensive computing tasks and real-time monitoring of smart devices entail significant demands on computing paradigms, leading to the emergence of the Dispersed Computing (DCOMP) paradigm [<xref ref-type="bibr" rid="ref-13">13</xref>]. The DCOMP paradigm, as a novel architectural approach, was initially devised to address challenges pertaining to the heterogeneous geographical distribution, collaborative computation, and communication between diverse devices. This is achieved by introducing a dispersed computation layer that serves as a task execution agent, strategically embedded between each network node&#x2019;s application and resource layers. The primary objectives are to optimize overall performance and network efficiency, as well as enhance the reliability of computing and communication processes [<xref ref-type="bibr" rid="ref-14">14</xref>]. With further research, dispersed computing is defined as connecting devices with general computing and communication properties into a network organism in a wide area, so as to achieve the purpose of device interconnection and capability interoperability. According to the specific requirements of users and applications, computing tasks can automatically access multiple devices with the environment and idle resources from the network, complete the computing tasks in a cooperative way among nodes, and return the computing results [<xref ref-type="bibr" rid="ref-15">15</xref>]. The dispersed computing paradigm, as a new form of dispersed computing, does not aim to replace the existing traditional computing architectures such as cloud computing, mobile cloud computing, fog computing, and mobile edge computing. Instead, it serves as an augmentation and natural progression of these traditional computing paradigms. Its primary objective is to address the evolving requirements of users in the Internet of Everything (IoE) domain, ensuring the fulfillment of increasingly demanding computing quality standards.</p>
<p>The primary contributions of this paper can be summarized as follows:</p>
<p>1) This paper explores a distributed computing environment that utilizes UAV track inspection, comprising multiple mobile users (MUs) and networked computation points (NCPs). One dispersed computing platform exists between the Mus and the NCPs, to achieve task allocation and resource management.</p>
<p>2) The NCPs offer computational resources to mobile users (MUs) based on task requirements, ensuring the timely completion of tasks on the dispersed computing platform. The allocation of computing resources is presented as a binary optimization problem, aiming to minimize energy consumption during task execution.</p>
<p>3) To minimize energy consumption across all the networked computation points (NCPs) in the investigated dispersed computing environment, a Markov decision process (MDP) model is employed. Additionally, a resource allocation approach based on Double Q Network (DDQN) is introduced to determine the optimal allocation strategies.</p>
<p>The paper is presented in the form of the following: <xref ref-type="sec" rid="s2">Section 2</xref> provides the related works. <xref ref-type="sec" rid="s3">Section 3</xref> gives the system model. The solutions for the proposed resource allocation problem are provided in <xref ref-type="sec" rid="s4">Section 4</xref>. Numerical simulations are provided in <xref ref-type="sec" rid="s5">Section 5</xref>. Finally, the work in <xref ref-type="sec" rid="s6">Section 6</xref> is concluded.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Resource Allocation for Edge Computing</title>
<p>As commonly understood, the cloud computing paradigm [<xref ref-type="bibr" rid="ref-16">16</xref>] often faces the challenge of geographical distance between a multitude of terminal devices and the centralized cloud center, and cloud computing also consumes substantial bandwidth from the backbone network, leading to network congestion and elevated bandwidth expenses. Mobile cloud computing, as an inherent progression and enhancement of cloud computing, addresses these challenges, but it still encounters challenges related to network congestion and significant computational expenses [<xref ref-type="bibr" rid="ref-17">17</xref>]. Subsequently, fog computing [<xref ref-type="bibr" rid="ref-18">18</xref>&#x2013;<xref ref-type="bibr" rid="ref-20">20</xref>] was introduced. Compared with mobile cloud computing, fog computing is closer to the user than mobile cloud computing. It can not only provide the resources needed for computing but also provide some other services, such as image recognition. It also faces a lot of security and privacy issues [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. Mobile edge computing (MEC) [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>] bears a resemblance to fog computing, wherein a versatile computing server is deployed at the network edge, and tasks are likewise transmitted to nearby high-performance servers via wireless links. The difference is that MEC supports user location awareness, and MEC servers belong to mobile operators and are generally deployed on the base station side [<xref ref-type="bibr" rid="ref-25">25</xref>], while fog computing nodes are generally managed by individuals. From the perspective of privacy protection, MEC is relatively more reliable. From the design of MEC paradigm, if the terminal nodes and edge nodes in MEC are not under load for a long time, the computing and communication resources will be idle. Currently, resource allocation has grown into a crucial branch of edge computing.</p>
<p>Liu et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] explored the application of non-orthogonal multiple access (NOMA) technologies to enhance connectivity in IoT networks, enabling energy-efficient mobile edge computing (MEC). Additionally, Yang et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] presented a method of resource allocation in multi-access MEC systems that are designed to minimize mobile device energy consumption while meeting competing deadlines. The scenarios consider a scenario in which several mobile devices with different computing capabilities and energy restrictions offload their MEC tasks to a nearby server. Moreover, Chen et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] analyzed the coordination optimization of computation allocation and resource allocation in edge computing, and proposed a new method: a time attention-deterministic strategy gradient (TADPG) for solving the problem. Zhu et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] presented a synthesis of hybrid NOMA and MEC techniques in their work. They solved the optimization problem with respect to power and time allocation for each group and alleviated latency and energy consumption. The method utilizes a matching algorithm for optimizing power and time allocation to cut down the complexity of the method further. Otherwise, Ali et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] posed the Pow Mig Expand algorithm, based on a comprehensive utility function, which allocates the requests received to the optimal server. In addition, they also introduced the Energy Efficient Intelligent Splitter (EESA) algorithm, which utilizes deep learning to allocate the required energy savings to the most suitable servers. Proposed methods are all using fixed infrastructure and lack scalability, making them unable to flexibly respond to sudden traffic surges or damage caused by disasters or other factors. In order to improve the flexibility of computing networks, recent attention has been focused on mobile device-assisted work. Sun et al. presented a novel system architecture in their work [<xref ref-type="bibr" rid="ref-31">31</xref>] that utilizes unmanned aerial vehicles (UAVs) and multiple Industrial Internet of Things (IIoT) nodes. The proposed architecture enables the direct transmission of resources collected by IoT nodes to the UAV for processing. The main goal of this system is to optimize response time for forest fire monitoring considering some constraints between nodes, i.e., minimize the maximum time. The system proposes a Markov random field-based decomposition (MRF) method and a learning-based collaborative particle swarm optimization (LCPSO) scheme to achieve the optimal allocation. These techniques work together to allocate the resources optimally in the UAV system, facilitating efficient forest fire monitoring. In their work, Chai et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] considered users with task computing needs who may choose to perform edge computing offloading, local computing, or device-to-device offloading. The authors proposed a joint optimization method which wants to minimize costs by dividing the task into small data segments and enabling various computational modes to execute them simultaneously. On a similar note, Wang et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] utilized MEC and D2D communication to facilitate offloading from the core network. They studied the optimal resource allocation strategy in the ICWN (Cognitive Wireless Network Internet) scenario. However, it is worth noting that the assistive technology may involve additional costs. In addition, the above research overlooked the potential untapped resources within the network and the achievement of utilizing neighboring computing nodes. If these factors are considered, it may enhance the efficiency and performance of the system.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Resource Allocation for Dispersed Computing</title>
<p>Following the emergence of MEC, DCOMP was introduced as a complementary solution, aiming to address this gap and harness the untapped computing resources effectively. The objective is to establish a dispersed and distributed network ecosystem with heterogeneous resources, maximizing the utilization of all available resources within the network and delivering users with enhanced efficiency and reliability in computing services. Dispersed computing possesses distinct advantages compared to other paradigms. The objectives of conventional computing paradigms involve task offloading to either the cloud, which possesses greater computing power, or edge nodes with shorter communication distances. This is done to minimize transmission delays and energy consumption. However, central nodes like fog computing nodes, MEC nodes, or service gateways in cloud computing exhibit limitations [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>With a massive increase in users and tasks, the central node fails, services will be interrupted and tasks will be lost, resulting in a poor user experience. Dispersed computing is an architecture for dispersed resources. In this architecture, nodes with computing power are abstracted as general computing points, and the DCOMP intermediate layer is responsible for security authentication, privacy protection, task decision-making and data routing transmission between nodes [<xref ref-type="bibr" rid="ref-36">36</xref>], and there is no centralized scheduling center. In addition, in a dispersed computing environment, the joining and leaving of nodes are very efficient and transparent. Simultaneously, it also addresses the drawbacks of the cloud computing paradigm, including high delay and high cost, while resolving the limitations of edge computing, such as the limited coverage area of edge nodes and the inability of terminal nodes to collaborate directly.</p>
<p>Indeed, in contrast to mobile edge computing, there has been growing attention to resource allocation in dispersed computing. Several studies [<xref ref-type="bibr" rid="ref-37">37</xref>&#x2013;<xref ref-type="bibr" rid="ref-39">39</xref>] have emerged in this area, taking into account the presence of idle devices and limited resources within the network. A noteworthy example is the work of Rahimzadeh et al. [<xref ref-type="bibr" rid="ref-37">37</xref>], who introduced SPARCLE. SPARCLE is a scheduling scheme for dispersed computing systems and is designed especially for dispersed computing environments which are provided with a network-aware algorithm for allocating tasks and resources for polynomial time flow processes. The proposed schemes are very useful to optimize resources and improve overall system performance. The collaborative management model for the task resource in a dispersed computing system is proposed by Zhou et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] and focuses on the quantification of dynamic dependencies between the resources occupied in a dispersed computing system and the consumed resources of task requests. This can better understand the resource utilization modes and achieve effective resource allocation. A kind of intelligent controlling mechanism to solve the problem of the overburdened load is brought into the collaborative management model of task resources. These intelligent control technologies aim to dynamically allocate and manage resources to avoid overload and ensure system stability. Through intelligent control, this model can optimize resource allocation and improve the whole performance and reliability of dispersed computing systems. The load-balancing scheduling algorithms and the prioritization-aware resource allocation strategies of the open-source dispersed computing system are studied by Paulos et al. [<xref ref-type="bibr" rid="ref-39">39</xref>]. These strategies are very important and have been integrated into a core feature in middleware. However, the research on resource allocation in dispersed environments is relatively limited to the field of dispersed computing. It is necessary to conduct further exploration and investigation to deepen our understanding of resource allocation in this situation. In addition, considering the limited energy resources of devices and the need to optimize energy consumption, energy cost has become a crucial indicator, especially for mobile devices. Therefore, when designing resource allocation strategies in a dispersed computing environment, special attention should be paid to addressing energy efficiency.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>System Model and Problem Formulation</title>
<p>In this section, we describe the application scenarios and system models in dispersed computing environments with the goal of minimizing the energy consumption of networked computing nodes.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Scenario Description</title>
<p>Considering a dispersed computing environment including one dispersed computing platform, <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>N</mml:mi></mml:math></inline-formula> tasks published on the platform by <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>L</mml:mi></mml:math></inline-formula> mobile users, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>M</mml:mi></mml:math></inline-formula> networked computation points (NCPs) that can complete the tasks published on the platform using their computing abilities. The platform can publish and assign tasks for the mobile users, and can schedule computing resources of NCPs for various tasks. In the dispersed computing environment, the set of mobile users is denoted by <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. Each mobile user publishes the tasks in the platform and requires the computing resources from the NCPs. Assuming the tasks published by user should be completed within a time duration <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>T</mml:mi></mml:math></inline-formula>. In the researched dispersed computing environment, the position of mobile users is assumed to be unchanged with in the time duration for tasks to be completed.</p>
<p>The set of NCPs is given by <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, which can provide computing abilities to the tasks in the platform. For each NCP, the available computing resource of NCP <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>j</mml:mi></mml:math></inline-formula> for <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:math></inline-formula> that can be allocated for the tasks is denoted by <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. By using these computing resources, the NCPs can complete the tasks published on the platform by the mobile users.</p>
<p>The set of tasks published on the platform is denoted by <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, which is running on the NCPs. For each computing task <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>, the resources scheduled by the platform is denoted by <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, for <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>. Generally, the tasks <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> published by user <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>k</mml:mi></mml:math></inline-formula> will be uniformly treated as the related tasks in the set <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. The maximum number of tasks that each mobile user can publish on the platform is given by <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>Assuming <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to denote the computing resources of NCP <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> allocated for the task <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. We design a binary variable <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to indicated whether the task <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is running on the NCP <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. If <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, the task <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is executed on NCP <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. If <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the task <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is not executed on NCP <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Then we have <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In order to achieve load balancing among different NCPs, we assume that the maximum number of tasks that is computed on NCP <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Meanwhile, the maximum number of NCPs working on the same task <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is assumed to be <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and we have <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>System Model</title>
<p>In the researched dispersed computing environment, we assume all the tasks published by the mobile users should be completed only by the NCPs. The mobile users should publish the computation requirements on the dispersed computing platform. Once the requirements are answered, they should upload the related tasks to the NCPs and wait for the NCPs to complete the corresponding computing tasks. Therefore, for each mobile user, the whole-time duration can be divided into two parts using TDMA. The first part is to upload the computation tasks to the NCPs, where the second part is to complete the computation tasks by the NCPs. The diagram is given in the following figure.</p>
<p>Assuming the transmit power of mobile user is <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>p</mml:mi></mml:math></inline-formula>, and the upload transmission is undertaken on the frequency band with channel bandwidth <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>. To avoid co-channel interference among different users, we assumed that the task could be uploaded onto the NCP using orthogonal frequency patterns by different users. Then the transmission time for the tasks upload can be written by [<xref ref-type="bibr" rid="ref-32">32</xref>]
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>p</mml:mi><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Once the tasks are upload to the NCPs, it will be computed by the NCPs. The NCPs should use the limited time to complete the tasks to ensure the real-time satisfaction. The limited time available for computation of task <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the NCPs is <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Assuming the workload of task <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is given by <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and the related time for computing the tasks is <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In order to complete the tasks, the NCPs should contribute their computing abilities, which would bring energy consumption for task due to the control variable <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Therefore, the total cost for computing the task <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the NCPs can be given by
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the unit energy consumption of computation. The gross energy cost for the entire dispersed computing environment can be given:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>J</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Problem Formulation</title>
<p>It follows that we aim at minimizing the consumption of a distributed computing environment and that the minimization problem can be drawn out mathematically as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:munder></mml:mtd><mml:mtd><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mo>.</mml:mo></mml:mtd><mml:mtd><mml:mi>C</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x003A;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>C</mml:mi><mml:mn>2</mml:mn><mml:mo>&#x003A;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>C</mml:mi><mml:mn>3</mml:mn><mml:mo>&#x003A;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>To eliminate this problem <xref ref-type="disp-formula" rid="eqn-4">(4)</xref> and solve it, the optimal solution to <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the discrete offloading decision is found. In this way, the task offloading in inter-slot time is coordinated under a limited time constraint, and the total computational energy consumed is minimized for a network computing node. It is worth noting that <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a binary variable which is the offloading decision variable and is why the network computing nodes can choose user tasks intelligently. With the increase of mobile users and computing tasks, traditional methods may not adapt to deal with complex scenarios with dynamic characteristics and make intelligent decisions. A method of solution for a certain problem based on reinforcement learning is explored.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Proposed Solutions</title>
<p>This paper uses the Markov Decision Process (MDP) model to calculate the optimization problem. The goal is to minimize the network computing nodes&#x2019; energy consumption in the dispersed computing environment. Firstly, reward functions, actions and the state within the MDP model are defined. Secondly, a method based on Double Deep Q Network (DDQN) is proposed. This approach effectively addresses the problem of dimensionality by learning the Q function.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Markov Decision Process Model</title>
<p>State space: The state in the MDP model represents the space of the dispersed computing paradigm. For available resources, on time slot t, the state is decided by <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in our proposed system, where the former is the residual available computing resource of network computing nodes and the latter is the maximum number of offloaded nodes. These two parameters are observed to maintain the constraints of computing and resource capacity of network computing nodes. Additionally, if it reaches the suitable state, it is important to consider the energy consumption state <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>E</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> at each time slot t to determine. Therefore, the state vector is defined as <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> at the time node <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>T</mml:mi></mml:math></inline-formula>.</p>
<p>Action: The objective is to select the best action by mapping the state space to the action space. In our proposed dispersed computing environment, agents are responsible for determining the resource allocation of users&#x2019; computing tasks in each time period.</p>
<p>Reward: The actions of each agent are motivated by rewards. In the dispersed computing system, the reward function is given as follows to minimize energy consumption. We believe that the best action is produced with the lowest task reward value.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Markov Decision Process</title>
<p>The MDP problem is effectively solved through dynamic programming or linear programming. However, when dealing with unknown and time-varying state transition probabilities and immediate rewards in discrete time steps, reinforcement learning methods become more suitable for tackling such problems. MDP is commonly used as the theoretical framework for reinforcement learning [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
<p>We approach the optimization problem by formulating it as an MDP. The agent interacts with the environment through iterative processes and adaptive learning to determine the optimal actions. The agent finds the current state at each time step <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>t</mml:mi></mml:math></inline-formula> which is written as <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and chooses an appropriate action <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on the policy <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula>. The action space <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>a</mml:mi></mml:math></inline-formula> is finite, it is the same for state space <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>s</mml:mi></mml:math></inline-formula>. The policy <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula> maps the current state to the corresponding action, providing guidance for decision-making under different states. Subsequently, the agent switches to its next condition <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> based on a switch probability <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>p</mml:mi></mml:math></inline-formula> which receives a reward <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>R</mml:mi></mml:math></inline-formula>. Over the long term, the state value function <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>V</mml:mi></mml:math></inline-formula> represents the cumulative discounted reward <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>R</mml:mi></mml:math></inline-formula> of the agent in state <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. It quantifies the long-term impact of the policy <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula> in the current state and serves as a measure of the state-value pair.</p>
<p>Now, let us consider the state after executing action <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> as <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where the transition probability from <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is denoted as <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>p</mml:mi></mml:math></inline-formula>. By employing the MDP model, the state value function is expressed using temporal differences through the Bellman equation [<xref ref-type="bibr" rid="ref-41">41</xref>]:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>V</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:munderover><mml:mi>&#x03C6;</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="negativethinmathspace" /><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:munderover><mml:mi>&#x03C6;</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="negativethinmathspace" /><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03C6;</mml:mi><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Building upon the aforementioned process, the agent strives to develop an optimal control strategy <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>&#x03C0;</mml:mi></mml:math></inline-formula> that maximizes the expected cumulative discount reward at the current state. This essentially transforms the optimization problem into an optimal state function problem <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>V</mml:mi></mml:math></inline-formula>, as depicted below [<xref ref-type="bibr" rid="ref-41">41</xref>]:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>V</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03C6;</mml:mi><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mspace width="1em" /><mml:mi>C</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>C</mml:mi><mml:mn>3.</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Therefore, the optimal strategy is obtained by solving the above problems. The optimal action in state <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> can be given as [<xref ref-type="bibr" rid="ref-41">41</xref>]
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="normal">a</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msup><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">s</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">a</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Solution Based on DDQN</title>
<p>In practical scenarios, the curse of dimensionality often arises when dealing with high-dimensional states and action spaces. To mitigate this issue, this paper proposes a novel DRL algorithm that deviates from traditional Q-learning methods.</p>
<p>In the conventional implementation of a DQN, a single Q network is employed to evaluate actions and is similar to multiple action-value functions, which can be numerous. However, this approach may lead to significant errors and an overestimation of the action-value function, ultimately affecting the choice of the optimal strategy. To address this concern, we adopt the DDQN algorithm. The fundamental principle behind DDQN is to decouple action selection from evaluation by constructing distinct action-value functions.</p>
<p>By updating DNN&#x2019;s parameters, the DDQN-based approach enhances the approximation of the optimal Q-values. In DDQN, the Q-value expression can be formulated by
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In the DDQN framework, the primary neural network&#x2019;s weight parameters are denoted as <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. Additionally, it will be discussed for a target network later.</p>
<p>Unlike the general DQN algorithm, the DDQN algorithm incorporates an experience replay buffer. The agent converts relevant experiences into memory tuples <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and stores them in the replay buffer at each time step <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>t</mml:mi></mml:math></inline-formula>. The agent utilizes the stored experiences from the buffer which is to train the neural network. It randomly selects a small batch of experiences to fine-tune the neural network and adjust its parameters. This process allows DDQN to leverage randomly chosen previous experiences during each update. Several studies, such as those cited in [<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>], have demonstrated that experience replay improves training efficiency, reduces correlation between training data, and expedites the convergence of DDQN.</p>
<p>In the DDQN algorithm, a target Q-network is introduced. The target Q-network and the Q-network employ different neural network parameters although they have the same neural network structure. The temporal difference signals from the Q-network are received by the target Q-network, while the latter serves as an evaluator for the Q-function.</p>
<p>However, the target Q-network updates its weights periodically instead of continuously. This approach enhances stability and accelerates the learning process.</p>
<p>The agent randomly selects a small batch of sample <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>D</mml:mi></mml:math></inline-formula> with size <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>R</mml:mi></mml:math></inline-formula> from the experience replay buffer to remove temporal correlations in the training data. The loss function <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>L</mml:mi></mml:math></inline-formula> is continuously minimized to approach the target value and improve the accuracy of Q-function approximation during each iteration, as shown below [<xref ref-type="bibr" rid="ref-34">34</xref>]:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03C6;</mml:mi><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The structure process and training process of DDQN are illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The initialization phase involves setting the weights of the neural network and the experience replay buffer&#x2019;s size. In DDQN, the Q-network and the target Q-network initially share the same weights, denoted as <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Next, a neural network is employed to establish the relationship between its corresponding value function Q and each state-action pair. For further details, the decision maker needs to preprocess the allocation strategy of the distributed computing platform adequately.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The training process of DDQN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-1.tif"/>
</fig>
<p>During the training process of each episode, the agent first views the distributed computing platform&#x2019;s state. Then, to choose the action <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>a</mml:mi></mml:math></inline-formula> for execution, the <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>-greedy algorithm is applied. With a probability of <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>, the agent randomly selects an action. Otherwise, the state-action pair&#x2019;s Q-value is estimated using the Q-network to select action <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>a</mml:mi></mml:math></inline-formula>.</p>
<p>Next, the agent accepts the reward value <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>R</mml:mi></mml:math></inline-formula> and the next status by executing action <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>a</mml:mi></mml:math></inline-formula>. It constructs an experience tuple <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and stores it. The Q-network randomly selects a small batch of samples from the replay buffer using an experience replay mechanism in the training process. Through minimizing the loss function <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, DDQN updates the network parameters. This involves updating the Q-network parameters using Stochastic Gradient Descent (SGD) and calculating the target Q-value. In Algorithm 1, the specific process is provided.</p>
<p>Then, the agent gets the reward value <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>R</mml:mi></mml:math></inline-formula> and the next status by executing action <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>a</mml:mi></mml:math></inline-formula> builds the experience tuple <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and saves the experience tuple. The Q-network randomly selects small batch samples through an experience replay mechanism in view of the training process. Through minimizing the loss function <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, DDQN renews the network parameters, that is, the algorithm updates the Q-network parameters by taking the Stochastic Gradient Descent (SGD) method for the loss function <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> after calculating the target Q-value. Finally, the specific implementation process is provided in Algorithm 1.</p>
<fig id="fig-10">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-10.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Numerical Simulations</title>
<sec id="s5_1">
<label>5.1</label>
<title>Simulation Set</title>
<p>We consider the dispersed computing environment as a smaller area with multiple mobile users and three network computing nodes. We consider user devices that are not computationally capable or cannot meet the demand of computational tasks. Each user has a different size task. Here, we consider two cases: (i) each mobile user has only a single task. The size of tasks per user is randomly distributed in the range of 1&#x2013;2.5 MB; (ii) each mobile user has multiple tasks. Each mobile user has multiple computing tasks of different types and sizes. The number of tasks per mobile user is randomly distributed in the span of 2&#x2013;4, and the task&#x2019;s size is distributed in 0.1&#x2013;1 MB randomly. The user&#x2019;s transfer rate is equal to 1 MB/s. In the approach of this paper, the experience replay buffer&#x2019;s capacity is <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>C</mml:mi></mml:math></inline-formula> &#x003D; 200 and the selected small batch of samples&#x2019; capacity is <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>U</mml:mi></mml:math></inline-formula> &#x003D; 32. Set The learning rate parameter to <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mi>l</mml:mi><mml:mi>r</mml:mi></mml:math></inline-formula> &#x003D; 0.01 and set the reward decay parameter to <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>&#x03C6;</mml:mi></mml:math></inline-formula> &#x003D; 0.9. The energy consumption of users is decided by the task size of users, i.e., larger tasks consume more energy.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Performance Analysis</title>
<p>First, the convergence property of the resource allocation method based on DDQN is verified. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the convergence property of the DDQN-based method which is proposed with a total of 24 user tasks. For the DDQN-based resource allocation method the energy consumption of the system gradually drops while the number of training sets rises. In general, the DDQN-based method with multiple iterations of training is able to converge quickly to a relatively stable value.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Convergence of the proposed DDQN-based method</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-2.tif"/>
</fig>
<sec id="s5_2_1">
<label>5.2.1</label>
<title>One Task per User</title>
<p><xref ref-type="fig" rid="fig-3">Figs. 3</xref>&#x2013;<xref ref-type="fig" rid="fig-5">5</xref> show the energy consumption, number of tasks computed, and resource consumption of each dispersed computing platform under different numbers of users, respectively. To be closer to the realistic scenario, each task is selected with a random size. Therefore, the task size is different for each user with the different number of users. It is evident from <xref ref-type="fig" rid="fig-4">Fig. 4</xref> that the method based on DDQN is capable of accomplishing tasks across various task quantities. The energy consumption of each network computing node can be observed in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The network computing node 2's energy consumption is low compared to other network computing nodes because of the constraint on the number of offloads of network computing node 2. It is evident from <xref ref-type="fig" rid="fig-5">Fig. 5</xref> that with the tasks&#x2019; number drop, the computational tasks will be gradually allocated to the nodes with more resources. Network computing node 3 has the most resources compared to the other two nodes and can accommodate the largest number of computational tasks.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Task energy consumption under different task scales</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Task allocation under different task scales</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Computing resource consumption under different task scales</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-5.tif"/>
</fig>
</sec>
<sec id="s5_2_2">
<label>5.2.2</label>
<title>Multiple Tasks per User</title>
<p><xref ref-type="fig" rid="fig-6">Figs. 6</xref>&#x2013;<xref ref-type="fig" rid="fig-8">8</xref> show that each user has a different number and size of tasks. Since different users have different numbers of tasks, the network nodes need to consider the service count constraint on the user tasks. The number of tasks varies at different numbers of users, i.e., the corresponding number of tasks is [23,29,32,42,46,55]. In <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, while the number of tasks offloaded rises, multiple network computing nodes are required to collaborate with each other to complete the user&#x2019;s computing tasks. In <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, while the number of user tasks rises, task allocation is assigned to the node with more resources remaining to meet the user&#x2019;s computational tasks. In <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, when there are fewer tasks, each node satisfies the task demand well. With the increasing number of tasks, the agent gives preference to the network computing nodes with abundant remaining resources.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Task energy consumption under different task scales (multiple tasks)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Task allocation under different task scales (multiple tasks)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Consumption resource allocation under different task scales (multiple tasks)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-8.tif"/>
</fig>
<p>In order to further analyze the advantages and disadvantages of the algorithms. We conducted comparison experiments with other algorithms, mainly including the offloading scheme based on random offloading strategy, and the offloading scheme based on dueling DQN (dueling DQN). In the random scheme, the user randomly selects NCPs that satisfy the offloading requirements as offloading nodes. In greedy scheme, the user selects the NCP with the richest computational resources as the offloading node. In dueling scheme, the user offloading policy is obtained based on dueling DQN algorithm. The results of the comparison experiments are shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. We consider different transmission rates for each node, which ranges from [0.8, 1.2] Mb/s, and the CPU energy consumption per cycle is <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>10</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. As shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, we can observe that the algorithm proposed in this paper is able to complete the computation of the task with lower total energy consumption.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The total energy consumption for the three algorithms to complete the task computing</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_51408-fig-9.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this paper, we tackle the challenge of allocating resources for multiple users in a dispersed computing environment, where each user has distinct resource requirements and the system conditions vary over time. Our objective is to minimize energy consumption while considering both resource and computational constraints. To tackle this problem, we formulate it as an energy minimization problem with computational constraints and adopt an MDP model to capture the interplay between the dispersed computational resource allocation policy and the system environment. To obtain the optimal policy, we propose an approach based on DDQN. This approach utilizes the capabilities of deep neural networks to approximate optimal Q-values and inform resource allocation decisions. Additionally, we incorporate the experience replay mechanism to mitigate the dimensionality issues caused by action spaces and high-dimensional states. To determine whether our approach is effective, we conduct experimental simulations in various scenarios. The results demonstrate the efficacy of our approach in achieving efficient resource allocation while considering the dynamic nature of the dispersed computing environment.</p>
</sec>
</body>
<back>
<glossary content-type="abbreviations" id="glossary-1">
<title>Nomenclature</title>
<def-list>
<def-item>
<term>UAV</term>
<def>
<p>Unmanned aerial vehicle</p>
</def>
</def-item>
<def-item>
<term>DCOMP</term>
<def>
<p>Dispersed Computing</p>
</def>
</def-item>
<def-item>
<term>IoE</term>
<def>
<p>Internet of Everything</p>
</def>
</def-item>
<def-item>
<term>MDP</term>
<def>
<p>Markov Decision Process</p>
</def>
</def-item>
<def-item>
<term>MEC</term>
<def>
<p>Mobile edge computing</p>
</def>
</def-item>
<def-item>
<term>MUs</term>
<def>
<p>Multiple mobile users</p>
</def>
</def-item>
<def-item>
<term>NCPs</term>
<def>
<p>Networked computation points</p>
</def>
</def-item>
<def-item>
<term>DQN</term>
<def>
<p>Double Q Network</p>
</def>
</def-item>
<def-item>
<term>DDQN</term>
<def>
<p>Double Deep Q Network</p>
</def>
</def-item>
<def-item>
<term>NOMA</term>
<def>
<p>Nonorthogonal multiple access</p>
</def>
</def-item>
<def-item>
<term>JCORA</term>
<def>
<p>Joint optimization problem of Computation Offloading and Resource Allocation</p>
</def>
</def-item>
<def-item>
<term>TADPG</term>
<def>
<p>Temporal Attention Deterministic Policy Gradient</p>
</def>
</def-item>
<def-item>
<term>PowMigExpand</term>
<def>
<p>Power Migration Expand</p>
</def>
</def-item>
<def-item>
<term>EESA</term>
<def>
<p>Energy Efficient Smart Allocator</p>
</def>
</def-item>
<def-item>
<term>UE</term>
<def>
<p>User Equipment</p>
</def>
</def-item>
<def-item>
<term>IIoT</term>
<def>
<p>Industrial Internet of Things</p>
</def>
</def-item>
<def-item>
<term>LCPSO</term>
<def>
<p>Learning-based Collaborative Particle Swarm Optimization</p>
</def>
</def-item>
<def-item>
<term>MRF</term>
<def>
<p>Markov random field</p>
</def>
</def-item>
<def-item>
<term>D2D</term>
<def>
<p>Device-to-Device</p>
</def>
</def-item>
<def-item>
<term>NCPs</term>
<def>
<p>Networked Computation Points</p>
</def>
</def-item>
<def-item>
<term>SGD</term>
<def>
<p>Stochastic Gradient Descent</p>
</def>
</def-item>
</def-list>
</glossary>
<ack>
<p>None.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by Beijing Metro Consultancy Corporation Ltd. (the Research on 5G Key Technologies of Urban Rail Transit I).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, T.G. and J.L.; methodology, T.G., J.L. and S.D.; software, T.G., S.W. and J.L.; validation, T.G., S.D. and J.L.; formal analysis, S.D. and S.W.; investigation, T.G., S.D., S.W. and J.L.; resources, S.D. and S.W.; writing&#x2014;original draft preparation, T.G. and J.L.; writing&#x2014;review and editing, T.G., S.D., S.W. and J.L.; visualization, T.G., S.D., S.W. and J.L. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>N. R.</given-names> <surname>AlNaimi</surname></string-name></person-group>, &#x201C;<article-title>Rail robot for rail track inspection</article-title>,&#x201D; <comment>MS thesis</comment>, <publisher-name>Qatar University</publisher-name>, <publisher-loc>Qatar</publisher-loc>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Siebert</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Teizer</surname></string-name></person-group>, &#x201C;<article-title>Mobile 3D mapping for surveying earthwork projects using an Unmanned Aerial Vehicle (UAV) system</article-title>,&#x201D; <source>Autom. Constr.</source>, vol. <volume>14</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1016/j.autcon.2014.01.004</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Automatic power line inspection using UAV images</article-title>,&#x201D; <source>Remote Sens.</source>, vol. <volume>9</volume>, no. <issue>8</issue>, pp. <fpage>824</fpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.3390/rs9080824</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Ai</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>I.</given-names> <surname>You</surname></string-name>, <string-name><given-names>K. K. R.</given-names> <surname>Choo</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Smart collaborative automation for receive buffer control in multipath industrial networks</article-title>,&#x201D; <source>IEEE Trans. Ind. Inform.</source>, vol. <volume>16</volume>, no. <issue>2</issue>, pp. <fpage>1385</fpage>&#x2013;<lpage>1394</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2019.2950109</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. F.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhai</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>L. M.</given-names> <surname>Guan</surname></string-name>, <string-name><given-names>K. N.</given-names> <surname>Mu</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Measurement for cracks at the bottom of bridges based on tethered creeping unmanned aerial vehicle</article-title>,&#x201D; <source>Autom. Constr.</source>, vol. <volume>119</volume>, no. <issue>11</issue>, pp. <fpage>103330</fpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.autcon.2020.103330</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Gao</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Conditional probability based multi-objective cooperative task assignment for heterogeneous UAVs</article-title>,&#x201D; <source>Eng. Appl. Artif. Intell.</source>, vol. <volume>123</volume>, no. <issue>1</issue>, pp. <fpage>106404</fpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.engappai.2023.106404</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Inzerillo</surname></string-name>, <string-name><given-names>G. Di</given-names> <surname>Mino</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Roberts</surname></string-name></person-group>, &#x201C;<article-title>Image-based 3D reconstruction using traditional and UAV datasets for analysis of road pavement distress</article-title>,&#x201D; <source>Autom. Constr.</source>, vol. <volume>96</volume>, pp. <fpage>457</fpage>&#x2013;<lpage>469</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.autcon.2018.10.010</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>P&#x00E1;li</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Mathe</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Tamas</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Bu&#x015F;oniu</surname></string-name></person-group>, &#x201C;<article-title>Railway track following with the AR.Drone using vanishing point detection</article-title>,&#x201D; in <conf-name>2014 IEEE Int. Conf. Autom. Quality Testing, Robot.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2014</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Flammini</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Naddei</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Pragliola</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Smarra</surname></string-name></person-group>, &#x201C;<article-title>Towards automated drone surveillance in railways: State-of-the-art and future directions</article-title>,&#x201D; in <conf-name>Adv. Concepts Intell. Vis. Syst.: 17th Int. Conf., Lecce, Italy</conf-name>, <publisher-name>Springer International Publishing</publisher-name>, Oct. 24&#x2013;27, <year>2016</year>, pp. <fpage>336</fpage>&#x2013;<lpage>348</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>I.</given-names> <surname>You</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Smart collaborative evolvement for virtual group creation in customized industrial IoT</article-title>,&#x201D; <source>IEEE Trans. Netw. Sci. Eng.</source>, vol. <volume>10</volume>, no. <issue>5</issue>, pp. <fpage>2514</fpage>&#x2013;<lpage>2524</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSE.2022.3203790</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Manyika</surname></string-name> <etal>et al.</etal></person-group>, <source>The Internet of Things: Mapping the Value Beyond the Hype</source>. vol. <volume>3</volume>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>McKinsey Global Institute</publisher-name>, <year>2015</year>, pp. <fpage>1</fpage>&#x2013;<lpage>24</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Calhoun</surname></string-name></person-group>, &#x201C;<article-title>DARPA emerging technologies</article-title>,&#x201D; <source>Strateg. Stud. Q.</source>, vol. <volume>10</volume>, no. <issue>3</issue>, pp. <fpage>91</fpage>&#x2013;<lpage>113</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. R.</given-names> <surname>Schurgot</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>A. E.</given-names> <surname>Conway</surname></string-name>, <string-name><given-names>L. G.</given-names> <surname>Greenwald</surname></string-name>, and <string-name><given-names>P. D.</given-names> <surname>Lebling</surname></string-name></person-group>, &#x201C;<article-title>A dispersed computing architecture for resource-centric computation and communication</article-title>,&#x201D; <source>IEEE Commun. Mag.</source>, vol. <volume>57</volume>, no. <issue>7</issue>, pp. <fpage>13</fpage>&#x2013;<lpage>19</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/MCOM.2019.1800776</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. S.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Pedarsani</surname></string-name>, and <string-name><given-names>A. S.</given-names> <surname>Avestimehr</surname></string-name></person-group>, &#x201C;<article-title>Communication-aware scheduling of serial tasks for dispersed computing</article-title>,&#x201D; <source>IEEE/ACM Trans. Netw.</source>, vol. <volume>27</volume>, no. <issue>4</issue>, pp. <fpage>1330</fpage>&#x2013;<lpage>1343</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TNET.2019.2919553</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Garc&#x00ED;a-Valls</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dubey</surname></string-name>, and <string-name><given-names>V.</given-names> <surname>Botti</surname></string-name></person-group>, &#x201C;<article-title>Introducing the new paradigm of social dispersed computing: Applications, technologies and challenges</article-title>,&#x201D; <source>J. Syst. Archit.</source>, vol. <volume>91</volume>, no. <issue>7</issue>, pp. <fpage>83</fpage>&#x2013;<lpage>102</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.sysarc.2018.05.007</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Husain</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Kunz</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Prasad</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Samdanis</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Mobile edge computing with network resource slicing for Internet-of-Things</article-title>,&#x201D; in <conf-name>2018 IEEE 4th World Forum Internet Things (WF-IoT)</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2018</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Alzahrani</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Alalwan</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Sarrab</surname></string-name></person-group>, &#x201C;<article-title>Mobile cloud computing: Advantage, disadvantage and open challenge</article-title>,&#x201D; in <conf-name>Proc. 7th Euro Am. Conf. Telemat. Inf. Syst.</conf-name>, <year>2014</year>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. V.</given-names> <surname>Dastjerdi</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Buyya</surname></string-name></person-group>, &#x201C;<article-title>Fog computing: Helping the Internet of Things realize its potential</article-title>,&#x201D; <source>Computer</source>, vol. <volume>49</volume>, no. <issue>8</issue>, pp. <fpage>112</fpage>&#x2013;<lpage>116</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1109/MC.2016.245</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Ai</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>I.</given-names> <surname>You</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Smart collaborative balancing for dependable network components in cyber-physical systems</article-title>,&#x201D; <source>IEEE Trans. Ind. Inform.</source>, vol. <volume>17</volume>, no. <issue>10</issue>, pp. <fpage>6916</fpage>&#x2013;<lpage>6924</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2020.3029766</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ge</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Fan</surname></string-name> and <string-name><given-names>K. B.</given-names> <surname>Letaief</surname></string-name></person-group>, &#x201C;<article-title>Delay-sensitive task offloading in vehicular fog computing-assisted platoons</article-title>,&#x201D; <source>IEEE Trans. Netw. Serv. Manag.</source>, vol. <volume>21</volume>, no. <issue>2</issue>, pp. <fpage>2012</fpage>&#x2013;<lpage>2026</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSM.2023.3322881</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Alrawais</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alhothaily</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Hu</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Cheng</surname></string-name></person-group>, &#x201C;<article-title>Fog computing for the internet of things: Security and privacy issues</article-title>,&#x201D; <source>IEEE Internet Comput.</source>, vol. <volume>21</volume>, no. <issue>2</issue>, pp. <fpage>34</fpage>&#x2013;<lpage>42</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1109/MIC.2017.37</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>T. M.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>I.</given-names> <surname>You</surname></string-name> and <string-name><given-names>H. K.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Smart collaborative distribution for privacy enhancement in moving target defense</article-title>,&#x201D; <source>Inf. Sci.</source>, vol. <volume>479</volume>, pp. <fpage>593</fpage>&#x2013;<lpage>606</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ins.2018.06.002</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Mach</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Becvar</surname></string-name></person-group>, &#x201C;<article-title>Mobile edge computing: A survey on architecture and computation offloading</article-title>,&#x201D; <source>IEEE Commun. Surv. Tutorials</source>, vol. <volume>19</volume>, no. <issue>3</issue>, pp. <fpage>1628</fpage>&#x2013;<lpage>1656</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1109/COMST.2017.2682318</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>K. B.</given-names> <surname>Letaief</surname></string-name></person-group>, &#x201C;<article-title>URLLC-awared resource allocation for heterogeneous vehicular edge computing</article-title>,&#x201D; <source>IEEE Trans. Vehicular Technol.</source>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1109/TVT.2024.3370196</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>C.</given-names> <surname>You</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name>, and <string-name><given-names>K. B.</given-names> <surname>Letaief</surname></string-name></person-group>, &#x201C;<article-title>A survey on mobile edge computing: The communication perspective</article-title>,&#x201D; <source>IEEE Commun. Surv. Tut.</source>, vol. <volume>19</volume>, no. <issue>4</issue>, pp. <fpage>2322</fpage>&#x2013;<lpage>2358</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1109/COMST.2017.2745201</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Peng</surname></string-name></person-group>, &#x201C;<article-title>Resource allocation for energy-efficient MEC in NOMA-enabled massive IoT networks</article-title>,&#x201D; <source>IEEE J. Sel. Areas Commun.</source>, vol. <volume>39</volume>, no. <issue>4</issue>, pp. <fpage>1015</fpage>&#x2013;<lpage>1027</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/JSAC.2020.3018809</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Huang</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>Energy efficiency based joint computation offloading and resource allocation in multi-access MEC systems</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>117054</fpage>&#x2013;<lpage>117062</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2936435</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Xing</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xu</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Tao</surname></string-name></person-group>, &#x201C;<article-title>A DRL agent for jointly optimizing computation offloading and resource allocation in MEC</article-title>,&#x201D; <source>IEEE Internet Things J.</source>, vol. <volume>8</volume>, no. <issue>24</issue>, pp. <fpage>17508</fpage>&#x2013;<lpage>17524</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/JIOT.2021.3081694</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Navaie</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Ding</surname></string-name></person-group>, &#x201C;<article-title>Resource allocation for hybrid NOMA MEC offloading</article-title>,&#x201D; <source>IEEE Trans. Wirel. Commun.</source>, vol. <volume>l9</volume>, no. <issue>7</issue>, pp. <fpage>4964</fpage>&#x2013;<lpage>4977</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TWC.2020.2988532</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Ali</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Khaf</surname></string-name>, <string-name><given-names>Z. H.</given-names> <surname>Abbas</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Abbas</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Muhammad</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>A deep learning approach for mobility-aware and energy-efficient resource allocation in MEC</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, no. <issue>7</issue>, pp. <fpage>179530</fpage>&#x2013;<lpage>179546</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3028240</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wan</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Learning-based resource allocation strategy for Industrial IoT in UAV-enabled MEC systems</article-title>,&#x201D; <source>IEEE Trans. Industr. Inform.</source>, vol. <volume>17</volume>, no. <issue>7</issue>, pp. <fpage>5031</fpage>&#x2013;<lpage>5040</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2020.3024170</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Chai</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>Q.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Task execution cost minimization-based joint computation offloading and resource allocation for cellular D2D MEC systems</article-title>,&#x201D; <source>IEEE Syst. J.</source>, vol. <volume>13</volume>, no. <issue>4</issue>, pp. <fpage>4110</fpage>&#x2013;<lpage>4121</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/JSYST.2019.2921115</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Du</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Guizani</surname></string-name></person-group>, &#x201C;<article-title>Resource allocation in information-centric wireless networking with D2D-enabled MEC: A deep reinforcement learning approach</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>114935</fpage>&#x2013;<lpage>114944</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2935545</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>I.</given-names> <surname>You</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Zhan</surname></string-name></person-group>, &#x201C;<article-title>Smart collaborative tracking for ubiquitous power IoT in edge-cloud interplay domain</article-title>,&#x201D; <source>IEEE Internet Things J.</source>, vol. <volume>7</volume>, no. <issue>7</issue>, pp. <fpage>6046</fpage>&#x2013;<lpage>6055</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/JIOT.2019.2958097</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Bellendorf</surname></string-name> and <string-name><given-names>Z. &#x00C1;.</given-names> <surname>Mann</surname></string-name></person-group>, &#x201C;<article-title>Classification of optimization problems in fog computing</article-title>,&#x201D; <source>Future Gener. Comput. Syst.</source>, vol. <volume>107</volume>, no. <issue>5</issue>, pp. <fpage>158</fpage>&#x2013;<lpage>176</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.future.2020.01.036</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Yang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Dispersed computing for tactical edge in future wars: Vision, architecture, and challenges</article-title>,&#x201D; <source>Wirel. Commun. Mob. Comput.</source>, pp. <fpage>1</fpage>&#x2013;<lpage>31</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1155/2021/8899186</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Rahimzadeh</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>SPARCLE: Stream processing applications over dispersed computing networks</article-title>,&#x201D; in <conf-name>2020 IEEE 40th Int. Conf. Distrib. Comput. Syst. (ICDCS)</conf-name>, <year>2020</year>, pp. <fpage>1067</fpage>&#x2013;<lpage>1078</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gong</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Hui</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Lin</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Zeng</surname></string-name></person-group>, &#x201C;<article-title>A task-resource joint management model with intelligent control for mission-aware dispersed computing</article-title>,&#x201D; <source>China Commun.</source>, vol. <volume>18</volume>, no. <issue>10</issue>, pp. <fpage>214</fpage>&#x2013;<lpage>232</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.23919/JCC.2021.10.016</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Paulos</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Priority-enabled load balancing for dispersed computing</article-title>,&#x201D; in <conf-name>2021 IEEE 5th Int. Conf. Fog Edge Comput. (ICFEC)</conf-name>, <year>2021</year>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Deep reinforcement learning-based air combat maneuver decision-making: Literature review, implementation tutorial and future direction</article-title>,&#x201D; <source>Artif. Intell. Rev.</source>, vol. <volume>57</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1007/s10462-023-10620-2</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. C.</given-names> <surname>Luong</surname></string-name>, <string-name><given-names>D. T.</given-names> <surname>Hoang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Gong</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Niyato</surname></string-name>, and <string-name><given-names>I. K.</given-names> <surname>Dong</surname></string-name></person-group>, &#x201C;<article-title>Applications of deep reinforcement learning in communications and networking: A Survey</article-title>,&#x201D; <source>IEEE Commun. Surv. Tut.</source>, vol. <volume>21</volume>, no. <issue>4</issue>, pp. <fpage>3133</fpage>&#x2013;<lpage>3174</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/COMST.2019.2916583</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>He</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Zhao</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Yin</surname></string-name></person-group>, &#x201C;<article-title>Integrated networking, caching, and computing for connected vehicles: A deep reinforcement learning approach</article-title>,&#x201D; <source>IEEE Trans. Vehicular Technol.</source>, vol. <volume>67</volume>, no. <issue>1</issue>, pp. <fpage>44</fpage>&#x2013;<lpage>55</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1109/TVT.2017.2760281</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>