<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80115</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.080115</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Monitoring and Observability in Edge Computing Systems: Taxonomy, Comparative Analysis, and Research Directions</article-title>
<alt-title alt-title-type="left-running-head">Monitoring and Observability in Edge Computing Systems: Taxonomy, Comparative Analysis, and Research Directions</alt-title>
<alt-title alt-title-type="right-running-head">Monitoring and Observability in Edge Computing Systems: Taxonomy, Comparative Analysis, and Research Directions</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Ahmed</surname><given-names>Hamza</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Syed</surname><given-names>Hassan Jamil</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>hassan.jamil@apu.edu.my</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Aslam</surname><given-names>Aqsa</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Zehra</surname><given-names>Sehar</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Faseeha</surname><given-names>Ummay</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Othman</surname><given-names>Nurzati Iwani</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Computing, National University of Computer and Emerging Sciences</institution>, <addr-line>Karachi</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Technology, Asia Pacific University of Technology &#x0026; Innovation</institution>, <addr-line>Kuala Lumpur</addr-line>, <country>Malaysia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Hassan Jamil Syed. Email: <email>hassan.jamil@apu.edu.my</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>6</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_80115.pdf"></self-uri>
<abstract>
<p>Edge computing is an emerging model for latency-sensitive and distributed applications. However, the observability of edge computing systems in heterogeneous environments remains a challenge, as most existing approaches are limited to only the system, service, application, and network layers. This paper surveys state-of-the-art solutions for edge observability and monitoring. The paper further introduces a thematic taxonomy that groups the state-of-the-art edge observability and monitoring literature based on monitoring intent, telemetry indicators, observability scope, architectural layers, deployment environments, and observability toolchains. Finally, we compare representative solutions in terms of latency, system overhead, bandwidth consumption, and detection accuracy. Our analysis reveals the trade-offs present in existing solutions regarding the overhead and performance metrics, as well as the limitations they impose when scaling, addressing mobility, or augmenting telemetry. Lastly, we conclude with a discussion on open challenges and future directions for designing next-generation, edge observability frameworks.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Edge computing</kwd>
<kwd>edge monitoring</kwd>
<kwd>monitoring solutions</kwd>
<kwd>observability</kwd>
<kwd>telemetry</kwd>
<kwd>microservices</kwd>
<kwd>performance analysis</kwd>
<kwd>distributed systems</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, there has been an increasing shift in computing towards edge computing. This is achieved by shifting information and data processing closer to the point of origin, which was previously done at centralized servers. This computing shift offers several advantages in terms of lower latency, bandwidth optimization, and enhanced data security [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. It can be done by the reduction in physical distance between devices and servers. Edge computing has become essential for applications where the latency needs to be minimal, even a millisecond matters, such as autonomous vehicles and industrial robotics [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>]. The transmission delay is reduced by processing the data locally, which reduces the latency and requires less bandwidth. It preserves bandwidth usage and ensures more reliable data transfers. However, with these benefits, there needs to be a look at security and management issues that may arise with these solutions [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>For different cases, there are different edge architectures available. The hierarchical models process data layer by layer across the architecture, moving from local devices up to regional servers or the cloud for heavy lifting. This structure provides a high degree of control and central oversight [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. In contrast, flat architectures favor direct peer-to-peer communication. For a low number of devices, this simplifies the setup, whereas the management becomes increasingly difficult as the network grows. Fog computing has emerged as a middle layer to balance the load between the edge and the cloud as a bridge, to maintain low latency while providing extra storage and compute power [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>In the field of transportation, healthcare, and smart cities [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>], edge computing has a direct influence on manufacturing. In healthcare, real-time patient monitoring is done by edge computing using data analysis through wearable devices and remote sensors, which helps in improving patient care and outcomes [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. The use of connected vehicles, as well as intelligent traffic management systems in the transportation industry, requires low-latency processing of data to ensure the safety and efficiency of these systems. Smart cites use edge computing to allocate and monitor resources, improve safety, and control pollution [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates the conceptual role of eBPF (Extended Berkeley Packet Filter) in enhancing edge monitoring across multiple operational domains. It enables efficient kernel-level observability, and eBPF strengthens the monitoring and performance capabilities that are essential in modern edge computing applications. Edge computing itself plays an essential role in the sectors of manufacturing, healthcare, transportation, and smart cities, where the decision-making needs to be done quickly and critically. In healthcare, real-time monitoring is done through wearable sensors and medical IoT devices that depend on low-latency processing at the edge to have continuous monitoring of patient assessment and early anomaly detection. Similarly, transportation systems that include connected vehicles and intelligent traffic management depend on fast processing to maintain safety and operating efficiency. Smart city infrastructures benefit from edge-enabled monitoring for resource optimization, environmental sensing, and public safety management.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>eBPF-driven monitoring architecture for edge environments.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80115-fig-1.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the use cases of eBPF highlight how it improves observability and performance. This includes network tracing, runtime security enforcement, storage I/O optimization, and container-level isolation. In the distributed edge environment, the eBPF provides the mechanisms that are required for reliable real-time monitoring. Therefore, edge computing is adapted very highly due to its observability capabilities enabled by eBPF, which allows these systems to address complex, real-world challenges across diverse industrial and social domains.</p>
<p>In the future, it is clearly seen that edge computing is closely connected with emerging technologies, which include 5G networks [<xref ref-type="bibr" rid="ref-16">16</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>], the Internet of Things (IoT) [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>], and artificial intelligence (AI) [<xref ref-type="bibr" rid="ref-3">3</xref>]. The integration of 5G networks is ideal for real-time edge networks and applications that promise faster data speeds and lower latency. The IoT devices&#x2019; data can be processed locally at edge servers because it generates large amounts of data, which reduces the need for constant cloud communication. Furthermore, deploying the machine learning algorithms and AI applications on edge devices allows real-time data analysis, which helps in decision-making, significantly enhancing the capabilities of edge computing systems [<xref ref-type="bibr" rid="ref-22">22</xref>]. However, security concerns required a widespread adoption will require addressing security concerns, standardizing protocols, and exploring novel applications across various industries.</p>
<p>In this paper, we present a detailed survey of the state-of-the-art edge monitoring solutions. To the best of our knowledge, only a limited number of existing studies have comprehensively reviewed cloud-edge monitoring solutions, which highlights the uniqueness and significance of our work. For instance, previous studies [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>] have explored various features of edge-cloud monitoring, but have been deficient in considering all relevant dimensions comprehensively. In this survey, we cover a wide range of critical aspects of edge monitoring solutions, including architecture, monitoring perspectives, communication models, scalability, and the overhead of these solutions. These features are essential in defining the attributes and effectiveness of monitoring solutions.</p>
<p>The rest of the paper is as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews related work on edge monitoring solutions. <xref ref-type="sec" rid="s3">Section 3</xref> proposed the thematic taxonomy for edge monitoring solutions. <xref ref-type="sec" rid="s4">Section 4</xref> provides an overview of the state-of-the-art edge monitoring solutions. <xref ref-type="sec" rid="s5">Section 5</xref> discusses open research issues and challenges in edge computing. <xref ref-type="sec" rid="s6">Section 6</xref> presents a conclusion of the whole paper to present the findings.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>This section reviews and analyses existing studies that are related to edge computing-based monitoring and observability frameworks. The study examines their proposed architectures, evaluation methodologies, and application domains. The discussion highlights the key contributions and limitations of these works, with particular emphasis on performance monitoring, anomaly detection, and reliability analysis in distributed edge environments. Furthermore, the motivations and research gaps identified in the literature are critically examined to establish the need for more lightweight, accurate, and scalable monitoring solutions suitable for resource-constrained edge systems.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Research Motivation</title>
<p>The fast-evolving mobile ecosystem and the intersections of new service paradigms, comprised of Internet-of-Things (IoT) ecosystems and sensor-based services, are driving the need for more real-time computational capabilities at the edge of the network. With new workloads and the adaptive allocation of resources with very low latency, and the availability of cloud resources in the center of the network, the classical cloud-based architectures are not effective. The growing number of mobile and IoT devices generates constant streams of data, and sending all the data to cloud data centers adds communication delays, soaks the bandwidth of the network, and increases energy consumption. With the described challenges, Mobile Edge Computing (MEC) appears to be one of the most important paradigm shifts for the design of the network, allowing for real-time computation to be done closer to the sources of data.</p>
<p>Task scheduling and offloading, in terms of the advantages, still remain difficult challenges in MEC. These challenges are the result of the irregular nature of the wireless channel&#x2019;s quality, the mobility of users, the various levels of capability of devices, the workload, the utilization of edge devices, and the position of the edge servers [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>]. Transfer events, where devices transition between network coverage zones, can interrupt or degrade the performance of a task being executed without the scheduling controlling for the mobility pattern. All of the above-described complexity and interrelated factors contribute to the determination of a suitable task placement. The number of factors describing the system can be used to create a multi-dimensional model, which can be used to describe task placement and latency, the cost of the data to be sent, the energy consumption of the devices, and the computation of the devices [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<p><xref ref-type="table" rid="table-1">Table 1</xref> provides a consolidated perspective on existing edge computing and monitoring approaches, a structured comparison of representative studies. In this work, we have explored diverse architectural models, monitoring strategies, and edge-enabled applications, all of which have their contributions that vary significantly in terms of empirical validation, scalability, and practical deployment. A comparative assessment of these models will help in identifying both their strengths and limitations, thereby illustrating unresolved challenges in real-time edge monitoring and observability. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, the primary edge computing studies provide a conceptual and architectural basis focusing on the reduction of the latency and support of the mobility, but the empirical evaluation is missing. Then, studies started emerging that provided domain-specific monitoring solutions, especially in the fields of healthcare, industrial systems, and smart infrastructure, which showcased optimized responsiveness and on-site analytics. Nevertheless, they were still constrained with regard to the application scope, with limited privacy, and minimal scalability. Contemporary kernel-level and smart monitoring solutions provide precise time, low latency, and superior detection accuracy. However, they tend to be (Operating System) OS-dependent and have additional training and deployment complexities. This outlines the gap present in theoretical literature and practical edge monitoring systems that are scalable and applicable. Integrated observability is lacking in the current MEC scheduling models. Observability is the ability to monitor the system&#x2019;s behavior, patterns of performance, the use of resources, and responsiveness. The distributed and heterogeneous nature of MEC systems does make real-time monitoring more difficult than in centralized cloud systems. This lack of observability makes it difficult to monitor the system&#x2019;s performance. From the absence of observability, scheduling algorithms in MEC systems perform poorly. They are simply unable to handle the runtime scenarios, emerging bottlenecks, and provide consistent Quality of Service (QoS).</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Strengths, and limitations of existing edge monitoring solutions.</title>
</caption>
<table>
<colgroup>
<col align="center" width="40mm"/>
<col align="center" width="7mm"/>
<col align="center" width="30mm"/>
<col align="center" width="35mm"/>
<col align="center" width="30mm"/> </colgroup>
<thead>
<tr>
<th>Proposed Method/Reference</th>
<th>Year</th>
<th>Tool/Technique Used</th>
<th>Strengths</th>
<th>Limitations</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Secure IoT Service Architecture (BOE)</bold> [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>2020</td>
<td>Trust Evaluation &#x002B; Cloud/Edge Cooperation</td>
<td>Fast convergence; accurate malicious device isolation</td>
<td>IoT-focused; limited generalization to other domains</td>
</tr>
<tr>
<td><bold>DeSFAM (eBPF &#x002B; AI Threat Monitoring Framework)</bold> [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>2024</td>
<td>eBPF &#x002B; VAE &#x002B; Isolation Forest</td>
<td>96% detection accuracy; blocks real CVE attacks; minimal overhead</td>
<td>Requires kernel-level access; ML training cost</td>
</tr>
<tr>
<td><bold>The Emergence of Edge Computing</bold> [<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>2019</td>
<td>Cloudlets, Fog Computing Model</td>
<td>Foundational EC model; reduces latency; improves mobility support</td>
<td>Conceptual only; lacks empirical evaluation</td>
</tr>
<tr>
<td><bold>Edge Computing: Vision and Challenges</bold> [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>2017</td>
<td>Edge&#x2013;Cloud Hierarchy, Service Offloading</td>
<td>Identifies full EC challenge landscape; comprehensive vision</td>
<td>No implementation; no monitoring experiments</td>
</tr>
<tr>
<td><bold>Determining Edge Node Real-Time Capabilities (eBPF)</bold> [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>2022</td>
<td>eBPF, kprobes, High-precision Timing</td>
<td>Nanosecond precision; extremely low overhead</td>
<td>Linux-only; no multi-node evaluation</td>
</tr>
<tr>
<td><bold>Demystifying Performance of eBPF Network Applications</bold> [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>2025</td>
<td>eBPF Benchmarking Suite</td>
<td>Reveals performance bottlenecks; precise evaluation</td>
<td>No security or monitoring pipeline evaluation</td>
</tr>
<tr>
<td><bold>Real-Time Monitoring and Analysis (Industrial Edge)</bold> [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>2019</td>
<td>Edge&#x2013;Cloud Collaborative Sensing</td>
<td>Enhanced industrial stability; real-time diagnosis</td>
<td>Not validated in large-scale distributed factories</td>
</tr>
<tr>
<td><bold>Mobility &#x0026; Dependence-Aware QoS Monitoring (ghBSRM-MEC)</bold> [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>2021</td>
<td>Bayesian Monitoring &#x002B; KNN</td>
<td>Higher QoS accuracy under mobility; reduces monitoring deviation</td>
<td>Small dataset; moderate computational overhead</td>
</tr>
<tr>
<td><bold>Edge-Based Insulator Self-Explosion Monitoring</bold> [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>2021</td>
<td>SSD, MobileNet, Multi-model Fusion (YOLO, RCNN)</td>
<td>High recognition accuracy (&#x007E;95%); lowers communication cost</td>
<td>Domain-specific; limited scalability beyond grids</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To overcome these challenges, there is a need for a comprehensive scheduling framework that is supported by observability and capable of making intelligent task offloading decisions while considering dynamic network behavior and changing performance conditions. The motivation of this research is to develop a cost-aware and mobility-enhanced task scheduling approach that integrates latency modeling, energy consumption estimation, and transmission delay analysis. There are some limitations in the current monitoring solution as well, including scheduling methods, a lack of mobility support, no performance modelling, and poor system visibility. This work is an effort toward the development of more adaptive, efficient, transparent edge computing. The goal of this work is not only to improve theoretical understanding but also the design of next-generation MEC systems, which can provide reliable and real-time computation services for modern mobile, IoT, and distributed applications.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Research Objective</title>
<p>The objective of conducting this study is to offer a thorough insight into the observability frameworks and tools that have been specially developed for containerized microservices. The primary contribution of this study is the establishment of a thematic taxonomy, which would serve to introduce a sense of order to a field of study that, to the current state, has not had a systematic means of categorizing edge monitoring tools and applications developed for such purposes in place. In addition to offering a means of systematic categorization, the findings and methodology of this study will offer a comparison that reveals the true effects of such tools on system reliability and performance.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Data Collection Process</title>
<p>In the process, searches were conducted using different libraries i.e., IEEE Xplore, ACM Digital Library, and Google Scholar, utilizing targeted keywords such as:<list list-type="bullet">
<list-item>
<p>Observability in micro services</p></list-item>
<list-item>
<p>Containerized environments</p></list-item>
<list-item>
<p>Micro services Performance Monitoring in Kubernetes</p></list-item>
<list-item>
<p>Observability in Cloud-Native Applications</p></list-item>
</list></p>
<p>Additionally, industry-standard tools such as OpenTelemetry, Prometheus, and Grafana were selected due to their widespread adoption. The collected data informed the thematic taxonomy and facilitated the comparative analysis of observability solutions.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Selection of Observability Frameworks and Tools</title>
<p>In the selection process, researchers have suggested exploring other ranked frameworks. The technologies involved in these frameworks were capable of providing relevance to micro services and containerized environments. There are some industry standard solutions like Prometheus, Loki, Jaeger available, as well as tools made for cloud computing (or containerized composite apps) to solve new problems. The assessment was centered on three primary observability characteristics:<list list-type="bullet">
<list-item>
<p>Metrics collection</p></list-item>
<list-item>
<p>Distributed tracing</p></list-item>
<list-item>
<p>Log aggregations</p></list-item>
</list></p>
<p>These criteria provided comprehensive insights into each framework&#x2019;s role in enabling observability within microservices architectures.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Inclusion and Exclusion Criteria</title>
<p><bold>Inclusion Criteria</bold>
<list list-type="bullet">
<list-item>
<p>Studies on task scheduling, computation offloading, or performance optimization in MEC.</p></list-item>
<list-item>
<p>Research offering analytical models for latency, transmission delay, or energy consumption.</p></list-item>
<list-item>
<p>Work addressing mobility-aware decision-making or dynamic wireless conditions in edge systems.</p></list-item>
<list-item>
<p>Studies incorporating observability elements such as performance monitoring or resource tracking.</p></list-item>
<list-item>
<p>Research presenting cost models, optimization formulations, or mathematical frameworks for edge task execution.</p></list-item>
<list-item>
<p>Recent studies comparing local and edge execution performance.</p></list-item>
</list></p>
<p><bold>Exclusion Criteria</bold>
<list list-type="bullet">
<list-item>
<p>Studies limited to cloud computing with no edge involvement.</p></list-item>
<list-item>
<p>Research on mobile performance is lacking in scheduling or offloading aspects.</p></list-item>
<list-item>
<p>Models missing latency, energy, or transmission cost components.</p></list-item>
<list-item>
<p>Papers providing only conceptual architecture without analytical evaluation.</p></list-item>
<list-item>
<p>Outdated work (pre-2017) unless foundational to MEC.</p></list-item>
<list-item>
<p>Solutions designed for monolithic or centralized systems without mobility or distributed execution.</p></list-item>
</list></p>
<p>The review and selection process establishes a clear foundation by identifying relevant studies, frameworks, and evaluation criteria within the domain of edge computing and observability. Building upon this foundation, <xref ref-type="sec" rid="s3">Section 3</xref> introduces a structured taxonomy that organizes these selected works into coherent categories. This transition enables a more analytical perspective, where the diverse approaches identified earlier are thematically classified based on common parameters, thereby facilitating a deeper understanding of patterns, relationships, and gaps in edge monitoring solutions.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Taxonomy of Edge Monitoring</title>
<p>This section presents a thematic taxonomy of existing edge monitoring solutions discussed in the literature, with particular emphasis on their classification. The proposed classification is based on a set of parameters that frequently appear across the majority of related works, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The parameters that were chosen and used in this thematic taxonomy include Monitoring Intent, Telemetry Indicators, Observability Scope, Architectural Layers, Deployment Environment, Observability Toolchain, and Architecture Paradigms. Each dimension incorporates refined subtopics that consistently emerge across modern monitoring frameworks. Monitoring Intent captures goals related to performance optimization, real-time responsiveness, and security and reliability assurance. Telemetry Indicators encompass the key signals used to characterize system behaviour, including latency-related measures, resource consumption indicators, and quality-of-service metrics. Observability Scope distinguishes whether monitoring is focused on device-level behaviour, edge-node execution, or application-level operations. This is where the computing actually happens, in Architectural Layers. The precise computing happens in the architectural layers, be it fog, edge, or cloud. The deployment setting denotes the place of the physical hardware and where monitoring is to be performed, such as well-endowed cloud data centers or lean devices at the edge. For us to reason about and troubleshoot what&#x2019;s going on in these systems, we need our observability toolchain to gather all this data and let us query it. Architectural paradigms specify how such disparate systems are architected and constructed, for example, edge&#x2013;cloud symbiosis, Edge AI, or the increasing reliance on containerized micro services.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Conceptual comparison of cloud and edge computing architectures.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80115-fig-2.tif"/>
</fig>
<p>This taxonomy provides a clear and logical structure that brings together practical methods, such as eBPF-based tracing, on-device inference, and fault stress testing, with theoretical models. By bridging this gap, it helps us identify patterns in a systematic way and maintain a consistent focus, ultimately allowing different monitoring studies to connect and complement each other more effectively.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Monitoring Intent</title>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the monitoring intent is concerned with what should be measured in an edge computing system. In the scope of traditional computing, resources at edge environments are limited and workloads are changeable abruptly and cannot afford delay. If you&#x2019;re new to the concept of monitoring intention, its core is finding a way to safeguard your system and make it always available and secure. To provide such monitoring, it requires monitoring edge nodes in a continuous manner so that the nodes are kept live and viable despite sudden changes in the network.</p>

<sec id="s3_1_1">
<label>3.1.1</label>
<title>Performance Optimization</title>
<p>The main concern for performance optimization is to keep things stable and use fewer resources. Since the nodes on the edges have fewerresources in terms of processing capability, a more accurate measurement of CPU use and memory patterns as well as task processing, becomes essential. This will enable more informed workload distribution and model adjustment to avoid the bottlenecks that are characteristic of resource-constrained hardware, even as the system maintains a high level of efficiency due to the changeable workload.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Real-Time Responsiveness</title>
<p>In most edge use cases, having the ability to respond to local occurrences on the spot can be taken as an indicator of success. Real-time monitoring ensures that an immediate feedback cycle is achieved to fulfill the latency requirements of higher intensity by monitoring queue accumulation and event distribution at the time of occurrence. For instance, in applications such as industrial sensing or auto-automation, real-time observation of phenomena ensures that critical decisions, for instance, at the time of observation, are taken at an edge point and not after an entire round-trip to a cloud point.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Security &#x0026; Reliability Assurance</title>
<p>Security and reliability mainly focus on maintaining the stability, integrity, and reliability of edge systems. The goal of monitoring is to identify anomalies, detect unauthorized or hazardous behavior, and evaluate system strength at both the software and hardware levels. Since edge nodes are deployed in distributed and often unprotected environments, monitoring acts as a safeguard against resource misuse, abnormal patterns, and potential intrusions. Some reliable monitoring systems also detect early signs of degradation, sensor faults, or performance instabilities that could lead to failures. Considering all these together, this ensures continuous, safe, and predictable operation across decentralized edge infrastructures.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Telemetry Indicators</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> highlights the different types of telemetry indicators, focusing on the various ways to measure system behavior in edge environments. This includes the measurement of the various ways in which resources are utilized in the different distributed computing, communication, and resource computing. The various ways resources are used provide a foundation for measurement for the different ways in the edge distributed devices in the various ways of empirical, monitoring, managing, and optimizing each of the resource types. Given that edges operate under different types of variable load, the edges operate under telemetry, which is processed at different dead networks and at different heterogeneous embedded processing hardware constraints. This provides different ways to measure and evaluate the balance of computation, communication, resource utilization, and various processes.</p>

<sec id="s3_2_1">
<label>3.2.1</label>
<title>Latency and Timing Metrics</title>
<p>The various types of performance metrics cover a core category of telemetry, which includes processing delays, task completion times, throughput, CPU, and even metrics of utilization, memory, and resource metrics. The metrics help measure and evaluate the ability of the edge node to execute and perform various workloads. The metrics measure and evaluate the distribution of the resources in the system, and the metrics of resources. The utilization of the resources in the system, and the distributed resources among the systems, is effective and is in a limited state of resources.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Resource Utilization Metrics</title>
<p>The network and communication parameters detail telemetry involving delays in transmission, packet loss, jitter, stability of the connection, and quality of the link. These parameters are critical in edge systems, as data must move through temporary and highly variable wireless channels. Tracking communication in a system makes it possible to select a new route, reconfigure the distribution of workload, or shift to a different edge node in order to maintain the continuity of the operation. Because network conditions are the primary factor that affects the latency, energy usage, and dependability of the distributed computation, communication-related telemetry is very important for maintaining a particular level of quality of service, as it is directly correlated to the level of service offered.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>QoS/Anomaly Indicators</title>
<p>For telemetry related to security and anomaly, the primary focus is on curious system behavior, including the atypical invocation of system calls, unanticipated behavior or activity, or any of the patterns that are detached from normal operation. These symptoms may signal potential interruptions, resource underutilization, data manipulation, or other undesirable conditions that can compromise the reliability and security of the system. By continuously surveilling and evaluating these parameters, edge systems can issue warning messages, take remedial action, or contain affected components when the situation demands it. This type of telemetry is critical within distributed edge computing environments, where nodes are subjected to more risk.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Observability Scope</title>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the degree to which a monitoring system can be considered effective is heavily influenced by its Observability Scope, which is its boundaries of strategic focus that dictate what monitoring is done and what is filtered out as insignificant behavior. The heterogeneity of edge systems naturally makes a monitoring plan inadequate, as these networks comprise a great number of nodes and dynamic network conditions. By defining a scope on which a monitoring system focuses, it becomes possible to highlight data points and filter out noise that could overcome the edge system&#x2019;s limited hardware capabilities in a monitoring system.</p>

<sec id="s3_3_1">
<label>3.3.1</label>
<title>Device-Level &#x0026; Sensor-Level Monitoring</title>
<p>On the most comprehensive level, observability begins to concentrate on each individual node. This includes monitoring CPU cycles, memory pressure, and system calls to understand the translation of hardware events into system-level latency. In the context of edge nodes that need to run independently in various distant sites, observability is a lifesaver, and it helps to specifically isolate a problem of resource consumption or hardware decay that needs to be fixed prior to a local failure that could bring down a system.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Edge Node &#x0026; Network-Level Monitoring</title>
<p>Network-level monitoring at the system as a whole, instead of focusing on just individual devices. It shows how different devices communicate with each other over the network. This helps in checking connection quality, delays in communication, and changes in performance caused by movement or unstable connections. Because edge devices work together and share data, a stable and reliable network is important. If network connections are weak or unpredictable, it can affect the entire system, causing service interruptions or uneven workload distribution.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Application/Service-Level Monitoring</title>
<p>Application-level observability correlates raw system data to what users are actually experiencing. It facilitates demonstrating how an application will perform under real-time conditions in terms of response time, accuracy and quality of results. Particularly in the domain of medical applications or industrial robotics, even minor latencies can cause big problems. This is to monitor when apps behave in the edge environment and, based on a policy, automatically allocate resources to meet performance requirements.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Architectural Layers</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> categorizes architectural layers; the edge contains several layers, which range from small sensors at the bottom to fog nodes in the middle and a central cloud layer at the top. Each layer provides another perspective as to whether the system is functioning or not, and how trustworthy that layer might be. It is the combination of these layers that assists in determining where monitoring should take place, since not all monitoring can or should be executed from a battery-powered sensor, as it may not support the same degree of checking as a cloud server. So in this case, that monitoring becomes more flexible. It is able to easily detect problems at different layers and, at the same time, provide a high-level view that can reduce system circulation time.</p>

<sec id="s3_4_1">
<label>3.4.1</label>
<title>Edge Layer</title>
<p>At the device level, this is the most basic layer of computing infrastructure. It is made up of IoT sensors, embedded devices, mobile clients, and lightweight processors that run close to the physical environment. At this layer, monitoring is done through hardware behavior, sensor readings, real-time actions, and system activity. This layer is critical for applications that demand immediate response, such as health monitoring, autonomous movement, and industrial sensing. Because of limited resources and changing operating environments, monitoring device level is critical for avoiding localized failures, increasing data integrity, and detecting abnormal patterns at this level without sending the problems to the upper level.</p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Fog Layer</title>
<p>The edge layer contains fog nodes, gateways, micro-data centers, and servers that actually process data at the point of origin. Instead of merely relaying information, these nodes filter, aggregate, and process data as it flows through. To ensure this layer runs smoothly, we have oversight regarding resource usage, real-time behavior, and how well these nodes work with the cloud or other devices. At the end of the day, it is the edge that makes low-latency and mobile apps feasible. And by keeping monitoring, we can immediately spot bottlenecks, balancing work across the cluster to keep performance even when data flow gets unpredictable.</p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>Cloud Layer</title>
<p>The powerful back end of the system is out in the cloud, tackling tasks that require a lot of storage and heavy processing. While edge devices address snappy, real-time operations, the cloud works on more holistic concerns, including retaining long-term data storage, working with complex models, and managing general security. At the cloud level, you&#x2019;re not making small instantaneous changes to adjust something immediately but rather tracking long-term patterns and ensuring data is kept synchronized as the system scales&#x2014;and that the whole thing remains stable along the way.</p>
</sec>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Deployment Environment</title>
<p>A monitoring system does not exist in a vacuum but is affected by the world around it. This could be a paradigm for the development of the system itself because, irrespective of whether the system falls into the health sector, industry, or transportation network, the environment defines the conditions that must be met, as mentioned in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. This could be the sensitivity of the information being dealt with or the allowable system downtime, because the moment these physical world conditions are factored in, the entire scenario of the system changes because of the monitoring challenges presented by the environment.</p>

<sec id="s3_5_1">
<label>3.5.1</label>
<title>Kubernetes-Based Edge &#x0026; Cloud Clusters</title>
<p>The Kubernetes edge system is useful in health monitoring platforms, biomedical sensing systems, and smart homes, in which reliable data collection is necessary. Most of the data in the above fields is sensitive, such as health data, critical alerts, and interactions between sensors, edge systems, and cloud systems. The key objective of health monitoring in these areas is accuracy of data, ensuring low latencies in critical alerts, as well as minimizing reliance on the central systems. The observability requirements suit various sensors, network connectivity, and changing workload requirements.</p>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>Containerized Edge Nodes</title>
<p>Industrial setups and smart grids are not merely about data acquisition and monitoring, but also ensuring that safety-critical systems are working perfectly without a hitch. In these setups, a network of edge nodes is embedded within factory floors and power lines, where any form of delay, no matter how minute, can result in a serious malfunction. In these setups that are already dealing with very intensive chores, such as real-time image analysis and malfunction detection, there is a high requirement on the monitoring tools to be extremely light and precise regarding their usage tracking.</p>
</sec>
<sec id="s3_5_3">
<label>3.5.3</label>
<title>Lightweight Edge Devices</title>
<p>Edge computing for transportation and smart cities is a domain that interacts with mobile clients, has fluctuating channels, and mixed workloads. It is used for traffic surveillance, vehicular analytics, and public safety surveillance, where low latency and location awareness pose a prominent challenge. System monitoring, in these domains, has to detect issues that may arise from mobility, a fluctuating Internet, coupled with a steady flow of data from all distributed nodes. Observability in these systems is, hence, focused on monitoring latency, throughput, and mobility. The deployment context requires monitoring systems that are adaptive, lightweight, and capable of coordinating across wide geographic regions.</p>
</sec>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Observability Toolchain</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> identifies observability toolchains as a key dimension, encompassing technologies such as eBPF, AI-driven analytics, distributed tracing frameworks, and telemetry pipelines. It is the combination of tools and techniques that help us make sense of the things that are happening across the edge. For edge computing, which is the combination of many different devices and unpredictable workloads, you cannot simply apply a single tool. It requires a blend of fine-grained understanding and the ability to scale. A reliable toolchain permits kernel-level tracing and real-time analytics, which endows the system with the situational awareness to recognize and mitigate issues early on, without introducing too much overhead to any of the processes.</p>

<sec id="s3_6_1">
<label>3.6.1</label>
<title>eBPF-Based Tracing &#x0026; Kernel Instrumentation</title>
<p>eBPF is the foremost technology for unprecedented granularity. It permits the observation of kernel activity, system calls, data packets, and OS interactions without the need to modify the source code of the targeted application. This is particularly advantageous to containerized platforms. Observability at the source allows us to identify and remedy bottlenecks and timing issues within a negligible duration. In a setting where every millisecond counts, eBPF is the indispensable technology to achieve optimal performance.</p>
</sec>
<sec id="s3_6_2">
<label>3.6.2</label>
<title>Artificial Intelligence and Machine Learning (AI/ML) Driven Analysis &#x0026; Detection</title>
<p>In the systems that change rapidly, fixed rules and limits are not effective. This is where Artificial Intelligence and Machine Learning help. These technologies learn from system data to spot unusual behavior, predict possible failures, and improve how tasks are scheduled. Instead of waiting for a system to fail, machine learning allows problems to be detected early. This is especially important in areas like healthcare, where identifying a problem before it happens can prevent serious consequences.</p>
</sec>
<sec id="s3_6_3">
<label>3.6.3</label>
<title>Distributed Telemetry &#x0026; Edge Analytics</title>
<p>It is apparent that the AI or ML techniques being utilized focus on the numerous ways that data patterns are inspected and examined within the context of tools of prediction and decision-making. It learns from data that is collected by telemetry with the purpose of examining the various ways that irregularity may occur. Another aspect of this technology is that it supports the various ways that adaptation can occur with tools of monitoring in an environment that lacks sufficient static threshold levels.</p>
</sec>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Architecture Paradigm</title>
<p>As presented in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the architectural paradigm focuses on the multiple potentialities found in the design of edge monitoring systems for the framework or structure, which considers how different responsibilities in systems for computing, communication, and observability are distributed across different devices, nodes, and cloud systems. Architectural paradigms found in edge computing provide the building blocks necessary for different means by which latency and fault tolerance can be traded off for monitoring depth. A taxonomic framework is further described with regard to different architectural choices in these paradigms because it explains different means by which different designs meet monitoring requirements and different means by which resource processes can be organized. These concepts are critical for determining different means by which monitoring frameworks meet the different gaps found in operations.</p>

<sec id="s3_7_1">
<label>3.7.1</label>
<title>Edge&#x2013;Cloud Hybrid</title>
<p>The edge-cloud hybrid model describes the different ways monitoring and computation are allocated among dispersed, localized nodes and centralized systems. This model describes the trade-off between rapid responsiveness and extensive computing power. It uses the edge for instantaneous filtering and local decisions and the cloud for prolonged storage and distributed optimization. These hybrid architectures offer different ways to advance flexibility and enhance data transmission. This division of responsibility supports the various ways frameworks adapt to the different types of variable workloads and network conditions.</p>
</sec>
<sec id="s3_7_2">
<label>3.7.2</label>
<title>Edge AI</title>
<p>Distributed Edge AI is the combination of different methods of machine learning and analytics at the edge. This paradigm provides different ways for systems to process and respond to data at the edge without the time delays of the global network. Monitoring within this paradigm focuses on and analyzes the different ways telemetry and application outputs interface to keep each of the AI processes within reliable bounds. Providing this type of intelligence creates different ways to augment autonomy and decrease the multiple types of communicative friction. This approach is especially valuable in the diverse needs of context-aware monitoring in the industrial and healthcare domains.</p>
</sec>
<sec id="s3_7_3">
<label>3.7.3</label>
<title>Micro Services</title>
<p>The containerized micro services paradigm focuses on how functionality is modularized into different isolated components. This approach aligns with the ways edge and cloud-native frameworks operate in concert for resource management and inter-service coordination. Monitoring in this paradigm is centered around the diverse ways container performance is monitored and the patterns of communication, or lack thereof, in each of the distributed services. The modular approach creates an environment for the different ways fault isolation and dynamic scaling, or the lack thereof, are achieved. This makes it appropriate for the many different ways complex edge deployments need to scale.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Review and Comparative Analysis of State-of-the-Art Edge Monitoring Solutions</title>
<p>This section comprises two subsections. In <xref ref-type="sec" rid="s4_1">Section 4.1</xref>, we discuss the state of the art monitoring solutions, and in <xref ref-type="sec" rid="s4_2">Section 4.2</xref>, we discuss about the comparative analysis of the current state-of-the-art monitoring solutions using the thematic taxonomy we have developed in <xref ref-type="sec" rid="s3">Section 3</xref>.</p>
<sec id="s4_1">
<label>4.1</label>
<title>State-of-the-Art Edge Monitoring Solutions</title>
<p>Edge computing is a paradigm that has emerged to address the challenges posed by traditional cloud computing architectures. This taxonomy synthesizes insights from research papers to provide a broad overview of the challenges, solutions, applications, evaluation methodologies, integration with emerging technologies, and future directions in the field of edge computing.</p>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Foundational and Architectural Edge Frameworks</title>
<p>This category includes solutions that define the architectural foundations of edge computing, focusing on workload distribution, latency reduction, and integration with cloud environments. These works provide the underlying infrastructure and system design principles upon which monitoring and observability mechanisms are built, but they do not explicitly implement advanced monitoring techniques.</p>
<p>a. EC&#x2019;s Rise</p>
<p>In [<xref ref-type="bibr" rid="ref-31">31</xref>], Satyanarayanan discusses the different ways in which edge computing has evolved as a new technology to solve the issues of outdated cloud systems. The attention is drawn to the measurements of different manners in which massive amounts of data are governed in the Internet of Things (IoT). This evolution of placing compute and storage nodes, also known as &#x201C;cloudlets,&#x201D; closer to the data source offers a basis for different ways in which latency is overcome and privacy is ensured. The different manners in which real-time tracking can be supported by cloud computing systems are offered by different methods involving bandwidth and latency. This approach offers different manners in which sensitive data can be preprocessed by sensor data which is an important solution in addition to cloud systems to ensure different manners in which agility and efficient data management.</p>
<p>b. EC Emergence</p>
<p>In [<xref ref-type="bibr" rid="ref-32">32</xref>], the need for edge computing has been discussed in terms of it becoming a revolutionary technology that brings resources closer to different types of mobile devices as well as sensors. This technological change provides different means of efficiently achieving scalable cloud services for IoT. The different means by which the application workflow of cloudlets relies on edge analytics to improve bandwidth as well as the efficiency of the system have been described. Because of research work done by different research initiatives, edge computing stands ready to provide the means by which processing power gets closer to the end-user.</p>
<p>c. Secure IoT EC</p>
<p>In [<xref ref-type="bibr" rid="ref-12">12</xref>], Wang et al. introduce a secure IoT service architecture, which targets the different ways the computation workload is distributed across cloud and edge nodes. The proposed architecture targets the measurement of the different ways the network workload and sensitive tasks are distributed, with a focus on balancing tasks with a high latency rate at the edge nodes. The proposed system also targets the integration of all the different ways the balancing algorithms and cooperation methods ensure the optimal system functionality related to efficiency and security. The system enables different ways of improving system responsiveness and overcoming different ways of communication delays. The system uses all the different ways of ensuring data privacy and reliability.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>System-Level Observability and Lightweight Monitoring</title>
<p>This group focuses on efficient, low-overhead monitoring mechanisms designed for resource-constrained edge environments. It includes kernel-level observability (eBPF), real-time performance profiling, telemetry aggregation, and container-level monitoring. These solutions emphasize scalability, real-time data collection, and minimal system overhead, making them suitable for large-scale and dynamic edge deployments.</p>
<p>a. eBPF Cloud Apps</p>
<p>In [<xref ref-type="bibr" rid="ref-38">38</xref>], Fournier et al. review the different ways in which eBPF innovation is utilized inside cloud administrations, and underline its role in the numerous ways by which assets are arranged and secured. This study has sorted use cases of eBPF under various zones: networking, storage, and security, among others, to feature the different ways it has been embraced to accomplish execution changes. It gives a premise for the different ways by which analysts look at arrangement methods and execution increases. Various ways to seek machine learning and programmable storage have likewise been recognized through eBPF. On-path vs. off-path monitoring for different structure plans gave insight into the various ways in which complexity is estimated and tried in an execution.</p>
<p>b. MessTool Real Time (RT)</p>
<p>In [<xref ref-type="bibr" rid="ref-33">33</xref>], Gowtham et al. discuss the various ways real-time capabilities are achieved on edge servers running on Linux through the MessTool framework. This tool focuses on the measurement of the various ways industrial applications perform on edge nodes, providing a foundation for detailed insights into timing performance. The framework utilizes the various ways of eBPF probing for precise measurement and focuses on the different ways time-critical applications are profiled. It addresses the various challenges in monitoring computing nodes and the different types of remote I/O controllers. The architecture includes various components like the Monitoring Manager and Aggregator to handle the different ways distributed measurements are processed. This provides various ways to evaluate real-time monitoring within the hardware constraints of edge environments.</p>
<p>c. EC eBPF Observability</p>
<p>In [<xref ref-type="bibr" rid="ref-34">34</xref>], a novel eBPF-based monitoring architecture is presented. Such a system affords edge observability at a fine grain. This approach leverages eBPF to observe kernel-side events and resource utilization without added overhead on the target system. This enables the system to gather data with minimal latency and performance overhead, making it suitable for a low-power edge network. The architecture also enables adaptive tracing and continuous monitoring of system health, aiding in increasing reliability. Our test demonstrates that the system incurs a minimal amount of CPU overhead and maintains an event capturing precision at the microsecond scale.</p>
<p>d. EC Lightweight Telemetry</p>
<p>In [<xref ref-type="bibr" rid="ref-39">39</xref>], a distributed lightweight telemetry system is proposed to improve the various ways real-time observability is achieved in large-scale edge environments. The framework focuses on the measurement of the various ways local monitoring agents adaptively collect and compress system performance data. This provides a foundation for the various ways summarized metrics are forwarded to edge collectors while maintaining a limited state of resource utilization. In contrast with typical cloud-based telemetry, this architecture enables various kinds of hierarchical aggregation, such that the variety of real-time insights is kept even if constrained by hardware. Experimental results show how it is possible to lower the transmission latency and computational overhead, confirming that the system ensures accurate monitoring between edge networks in various ways.</p>
<p>e. EC Orchestration Monitoring</p>
<p>In [<xref ref-type="bibr" rid="ref-40">40</xref>], a microservice-oriented orchestration monitoring system is developed to integrate the various ways monitoring capabilities work within Kubernetes environments. The framework tracks the various ways container lifecycles and inter-service communications are measured, providing a foundation for the different ways resources are managed. Through the various ways of continuous observability, the system provides different ways to predict anomalies and enable proactive fault mitigation. This highlights the various ways orchestration-aware monitoring is essential for supporting the different types of scalable and autonomous edge infrastructures.</p>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Application-Specific and Domain-Oriented Monitoring</title>
<p>These solutions are designed for specific application domains such as healthcare, smart infrastructure, and intelligent transportation systems. They leverage edge computing to enable real-time data processing, reduce latency, and optimize bandwidth usage. While highly effective within their domains, these approaches are often not generalized for broader edge environments.</p>
<p>a. EC Smart Grids</p>
<p>In [<xref ref-type="bibr" rid="ref-41">41</xref>], the implementation of an edge computing framework is discussed in relation to different aspects of real-time monitoring in smart grid systems. It points out that real-time monitoring in smart grid systems takes place in different forms that are measured by the different aspects in which traditional real-time monitoring systems have shortcomings in power grid systems. The significance of different aspects in which edge computing enhances frame rate and lowers detection delay compared to cloud computing systems has been mentioned. Various aspects of solving the problem related to scheduling of monitoring systems and edge server systems have been explained in relation to designing an algorithm for an efficient solution. The different aspects in which the framework operates in smart grid systems have been explained by different experiments.</p>
<p>b. Insulator Detection</p>
<p>To improve the accuracy and responsiveness of this detection system of power grid inspection, an intelligent online monitoring system for insulator self-explosion is proposed in [<xref ref-type="bibr" rid="ref-37">37</xref>]. This system integrates deep learning and edge computing. The framework deploys edge nodes to process captured Unmanned Aerial Vehicle (UAV) image data locally before sending summarized results to the cloud. It uses lightweight Single Shot Detector (SSD) neural networks for feature extraction and classification of insulator defects. When compared to conventional cloud-only solutions, this architecture dramatically lowers computational costs and transmission latency. Additionally, the system uses distributed edge servers to manage parallel image inference tasks, increasing network reliability and fault detection effectiveness. The proposed edge-deep learning hybrid system thus represents a scalable and efficient monitoring framework for power grid asset management, supporting autonomous maintenance and fault prediction.</p>
<p>c. Traffic Monitoring</p>
<p>Tang et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] focuses on improving performance and responsiveness in such distributed systems using edge computing, in a hybrid edge-cloud setup with tasks running on both edge nodes and the cloud. Local processing, for tasks like motion detection and object tracking, takes place on an edge node. Longer-term storage and advanced processing tasks are carried out in the cloud. Edge processing serves to save bandwidth and reduce latency compared to cloud-only systems by eliminating the step of uploading video to the cloud for processing. This leads to a more uniform load distribution across the network, allows for more rapid data acquisition, and permits real-time feedback. These properties are corroborated by empirical studies showing that the framework is scalable and responsive to changing conditions, e.g., the time to download the data falls considerably when varying the sampling precision.</p>
</sec>
<sec id="s4_1_4">
<label>4.1.4</label>
<title>QoS-Aware, Mobility-Aware, and Performance Monitoring</title>
<p>This category focuses on maintaining system performance and service reliability under dynamic conditions. It includes QoS monitoring, mobility-aware frameworks, and fault analysis mechanisms. These solutions aim to ensure stable service delivery by monitoring latency, throughput, and system resilience, particularly in environments with fluctuating workloads and user mobility.</p>
<p>a. QoS Monitoring</p>
<p>Zhang et al. [<xref ref-type="bibr" rid="ref-42">42</xref>] proposed a novel approach called ghBSRM-MEC (graph-based Bayesian Service Relationship Model for Mobile Edge Computing) that focuses on monitoring Quality of Service (QoS) in the context of mobile edge computing. This approach tackles challenges related to user mobility and interdependencies among QoS parameters with the goal of minimizing monitoring inaccuracies. It proposes the idea of forming parent attributes to mitigate dependence between QoS attributes and switching edges in a dynamic way for monitoring purposes. The architecture of ghBSRM-MEC, the algorithmic formulation of the methodology, and the empirical assessment of the methodology include difficulties within traditional QoS monitoring, the importance of probabilistic monitoring approaches, and the need for effective QoS monitoring in mobile edge computing.</p>
<p>b. Fault Impact</p>
<p>In [<xref ref-type="bibr" rid="ref-43">43</xref>], they examine the effect that various resource-contending faults have on the latency of edge-computing processes. And also, it looks into multiple overloads inflicted, for instance, on object recognition and speech recognition applications. The goal is to find out which faults are the worst for meeting the requirements to improve fault tolerance on edge computing systems. The research method involves benchmarking applications on a Raspberry Pi computing device and then performing a fault injection experiment to measure the disruption&#x2019;s impact on the performance metrics, which is latency. The results of the study reveal resource-stressing edge faults that can be addressed to increase the trustworthiness of edge systems and systems for resourcedeficient area.</p>
<p>c. EC Fault Resilience</p>
<p>In [<xref ref-type="bibr" rid="ref-44">44</xref>], a fault-resilient edge framework is proposed to analyze the various ways system overloads and hardware faults impact application latency and availability. The study introduces a foundation for measurement through fault injection mechanisms that emulate the various ways CPU and memory stress occur at runtime. By profiling the various ways different fault types affect the system, the research identifies different ways to evaluate critical vulnerabilities in time-sensitive applications. Experimental results show how the various ways of dynamic load redistribution can recover lost performance after transient failures, providing various ways to enhance the reliability of self-healing edge infrastructures.</p>
<p>d. EC QoS Mobility</p>
<p>In [<xref ref-type="bibr" rid="ref-36">36</xref>], a mobility-aware QoS monitoring framework is proposed to ensure the various ways consistent service delivery is managed in mobile edge scenarios. The system continuously monitors the various ways latency, jitter, and throughput interact as users move between network cells. By using probabilistic modeling, the framework reduces errors in monitoring when users or devices are moving. This helps limit service disruptions and keeps the quality of service more stable. It results in smoother and more reliable connectivity, which is especially important in smart transportation systems.</p>
</sec>
<sec id="s4_1_5">
<label>4.1.5</label>
<title>AI-Driven and Intelligent Monitoring Systems</title>
<p>This group represents the integration of artificial intelligence into edge monitoring. These solutions enable intelligent anomaly detection, model performance tracking, and adaptive system behavior. They mark a transition toward autonomous and predictive observability, where monitoring systems can dynamically respond to changes in workload and data patterns.</p>
<p>a. EC ML Drift Detection</p>
<p>In [<xref ref-type="bibr" rid="ref-45">45</xref>], a monitoring solution for machine learning drift detection is implemented to focus on the various ways inference accuracy is sustained at the edge. The system tracks the various ways data distribution shifts and prediction errors are measured at edge nodes to detect early signs of model degradation. Upon detection, the framework provides various ways to trigger retraining workflows and synchronize models with the cloud. This approach ensures the various ways distributed edge ML models remain adaptive, improving the different ways decision reliability is maintained in dynamic and data-intensive scenarios.</p>
<p>b. Federated EC Privacy</p>
<p>In [<xref ref-type="bibr" rid="ref-46">46</xref>], a federated edge monitoring system is proposed that enables secure data analysis in distributed IoT settings, preserving the user&#x2019;s privacy. Sensors spot trouble locally and send only encrypted data to a centralized system. This method is used to protect sensitive information. By integrating encryption and secure aggregation, the system provides strong data confidentiality. The series of experiments performed to confirm that the proposed secure monitoring technique works successfully with a marginal accuracy drop and lower communication.</p>
<p>c. EC Dynamic eBPF-driven Syscall Filtering and Anomaly Mitigation (DeSFAM)</p>
<p>Zehra et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed DeSFAM, an adaptive monitoring framework designed to improve the various ways real-time observability and threat mitigation occur in containerized environments. By combining the various ways of AI-driven anomaly detection with eBPF-based kernel tracing, the system allows the monitoring layer to adapt to the various ways workloads shift. Experimental evaluations show the various ways DeSFAM achieves high detection accuracy in identifying malicious activities while reducing the various types of monitoring overhead. The framework provides various ways for edge-hybrid environments to outperform static systems in terms of the measurement of latency and adaptability.</p>
</sec>
<sec id="s4_1_6">
<label>4.1.6</label>
<title>Other Monitoring Solutions</title>
<p>Several edge monitoring solutions have been extensively discussed in various research papers, each addressing different aspects of edge computing environments. Astrolabe, for example, is a hierarchical monitoring system that organizes nodes into domains (zones) and uses on-the-fly aggregation to compute summaries of system data, enhancing scalability and efficiency [<xref ref-type="bibr" rid="ref-47">47</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>]. There are a number of systems available to manage a large, complex environment. SDIMS [<xref ref-type="bibr" rid="ref-49">49</xref>] (Scalable Distributed Monitoring System) provides a robust infrastructure for large network systems and distributed resources, ensuring effective monitoring across the networks. Fog/Edge Monitoring emphasizes scalability, non-intrusiveness, and locality. They often feature data aggregation, event-driven monitoring, and long-term storage capabilities that support fog and edge infrastructures [<xref ref-type="bibr" rid="ref-50">50</xref>&#x2013;<xref ref-type="bibr" rid="ref-52">52</xref>]. Cloud monitoring solutions specially designed for cloud environments that mainly focus on scalability, robustness, non-intrusive, and real-time monitoring to maintain smooth cloud operations [<xref ref-type="bibr" rid="ref-53">53</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>]. IoT monitoring solutions are designed for Internet of Things (IoT) devices, which are sensors and smart devices [<xref ref-type="bibr" rid="ref-55">55</xref>&#x2013;<xref ref-type="bibr" rid="ref-58">58</xref>], prioritizing real-time monitoring, data processing, and context awareness to handle the unique challenges of IoT ecosystems [<xref ref-type="bibr" rid="ref-59">59</xref>]. Some monitoring solutions are designed for applications that can adapt automatically. These systems track performance and resource usage to ensure applications run efficiently. Additionally, in [<xref ref-type="bibr" rid="ref-60">60</xref>&#x2013;<xref ref-type="bibr" rid="ref-63">63</xref>], these solutions focus on cloudlets, which are small servers placed close to users to reduce delays. Monitoring in these environments emphasizes fast response times and scalability. The influence of monitoring in fog and edge computing [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>] underscores the necessity for scalable, non-intrusive, and context-aware monitoring solutions. Moreover, holistic monitoring services for fog/edge infrastructures propose comprehensive solutions that integrate these critical features [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-66">66</xref>]. Recent advancements highlight the integration of AI-driven techniques into monitoring systems for complex distributed environments. Machine learning enhances anomaly detection and log analysis and reduces alert noise in Kubernetes-based systems. Additionally, AI-driven microservices monitoring frameworks leverage fine-grained telemetry and model-centric metrics to improve system visibility, root cause analysis, and overall observability in dynamic cloud-native environments [<xref ref-type="bibr" rid="ref-67">67</xref>,<xref ref-type="bibr" rid="ref-68">68</xref>]. Lastly, discussions on push vs. pull approaches in web-based network management highlight the importance of monitoring and control in optimizing edge computing environments [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>]. Collectively, these solutions illustrate the diversity and complexity of edge monitoring, emphasizing the crucial need for adaptable and efficient monitoring frameworks to support the dynamic nature of edge computing.</p>
<p>The discussion of state-of-the-art edge monitoring solutions highlights a wide range of architectures, tools, and application-specific implementations, each addressing different challenges in edge environments. Due to the observation that reported numbers from various works are obtained from different platforms with different datasets, workloads, and edge/fog/cloud deployment scenarios, direct comparison of these numbers may not be completely apples-to-apples. Thus, comparisons in this paper are conducted based on a common set of axes (for example, latency, overhead, network bandwidth consumption, detection accuracy) when available. This comparison will serve as a guideline to provide relative and trend analysis among existing approaches, based on both quantitative numbers and architectural qualitative analysis. While these studies provide valuable insights into design choices and practical deployments, a deeper understanding requires an analytical comparison across common dimensions. To achieve this, the following analysis evaluates these solutions using a unified framework, enabling a clearer identification of their strengths, limitations, and overall effectiveness in supporting scalable and reliable edge observability.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Comparative Analysis of State-of-the-Art Edge Monitoring Solutions</title>
<p>This sub-section provides a critical analysis of edge monitoring solutions based on the parameters set in the taxonomy in <xref ref-type="sec" rid="s3">Section 3</xref>. This comparison is categorized into qualitative and quantitative aspects of existing edge monitoring solutions. The monitoring solutions presented in the table were analyzed, and their distinguishing attributes and features are illustrated in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of state-of-the-art monitoring solutions.</title>
</caption>
<table>
<colgroup>
<col align="center" width="18mm"/>
<col align="center" width="4mm"/>
<col align="center" width="4mm"/>
<col align="center" width="4mm"/>
<col align="center" width="4mm"/>
<col align="center" width="5mm"/>
<col align="center" width="4mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="4mm"/>
<col align="center" width="4mm"/>
<col align="center" width="4mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="4mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/> </colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Paper Reference</th>
<th align="center" colspan="3">Monitoring Intent</th>
<th align="center" colspan="3">Telemetry Indicators</th>
<th align="center" colspan="3">Observability Scope</th>
<th align="center" colspan="3">Architectural Layers</th>
<th align="center" colspan="3">Deployment Environment</th>
<th align="center" colspan="3">Observability Toolchain</th>
<th align="center" colspan="3">Architecture Paradigms</th>
</tr>
<tr>
<th>PO</th>
<th>RTR</th>
<th>SRA</th>
<th>LTM</th>
<th>RUM</th>
<th>QAI</th>
<th>DLSM</th>
<th>ENNM</th>
<th>ASLM</th>
<th>EL</th>
<th>FL</th>
<th>CL</th>
<th>KECC</th>
<th>DCEN</th>
<th>ELED</th>
<th>EBT</th>
<th>AIMD</th>
<th>DTEA</th>
<th>HEC</th>
<th>EAI</th>
<th>MS</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>EC&#x2019;s Rise</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>EC Emergence</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>Secure IoTEC</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>eBPF Cloud Apps</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>MEC Smart Grids</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>MestTool IRT</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>QoS Monitoring</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
</tr>
<tr>
<td><bold>Insulator Detection</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>Traffic Monitoring</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>Fault Impact</bold></td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>Lightweight Telemetry</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>EC eBPF Observability</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>EC Fault Resilience</bold></td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
</tr>
<tr>
<td><bold>EC ML Drift Detection</bold></td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>EC QoS Mobility</bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold>Federated EC Privacy</bold></td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
</tr>
<tr>
<td><bold>Orchestration Monitoring</bold></td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
</tr>
<tr>
<td><bold><italic>EC DeSFAM</italic></bold></td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2716;</td>
<td>&#x2714;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: <bold>PO</bold>: Performance Optimization, <bold>RTR</bold>: Real-Time Response, <bold>SRA</bold>: Security Risk Assessment, <bold>LTM</bold>: Low-Level Telemetry Metrics, <bold>RUM</bold>: Real User Monitoring, <bold>QAI</bold>: Quality of Analytics Insights, <bold>DLSM</bold>: Device-Level System Monitoring, <bold>ENNM</bold>: End-to-End Network Monitoring, <bold>ASLM</bold>: Application &#x0026; Service-Level Monitoring, <bold>EL</bold>: Edge Layer, <bold>FL</bold>: Fog Layer, <bold>CL</bold>: Cloud Layer, <bold>KECC</bold>: Kubernetes Edge Cloud Computing, <bold>DCEN</bold>: Data Center Environment, <bold>ELED</bold>: Edge-Level Embedded Devices, <bold>EBT</bold>: Edge-Based Telemetry, <bold>AIMD</bold>: AI-Driven Monitoring &#x0026; Diagnostics, <bold>DTEA</bold>: Distributed Trace &#x0026; Event Analysis, <bold>HEC</bold>: Hybrid Edge Computing, <bold>EAI</bold>: Edge AI, <bold>MS</bold>: Microservices.</p>
</table-wrap-foot>
</table-wrap>
<p>Edge monitoring is taking charge of modern distributed systems in terms of reliability, responsiveness, and security. As edge computing spreads into areas like healthcare and industrial automation, monitoring systems need to handle a wide range of devices while also working within limited resources. To provide a unified understanding of how research approaches this, this section presents a detailed analysis of the various ways state-of-the-art frameworks are built. The analysis synthesizes how each solution uses telemetry, architectural patterns, and monitoring goals to achieve system visibility.</p>
<p>Across the surveyed literature, the various ways monitoring solutions demonstrate their primary intent are quite clear. Several frameworks focus on the measurement of the various ways performance is optimized by reducing response latency and balancing computation. General studies emphasize the different ways proximity-based computation lowers delays in various latency-sensitive domains like vehicular networks [<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. More targeted systems leverage the various ways edge-side deep learning minimizes transmission delays [<xref ref-type="bibr" rid="ref-37">37</xref>]. While other tools, such as MessTool, focus on the various ways real-time responsiveness is measured through nanosecond-level timing [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. Security-centric frameworks expand this intent by focusing on the various ways reliability and threat mitigation are handled through risk assessment and dynamic balancing [<xref ref-type="bibr" rid="ref-12">12</xref>] distributed IoT deployments, while eBPF-based adaptive systems strengthen detection accuracy and minimize attack surfaces through kernel-level observability and anomaly identification [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>The variability in how telemetry is acquired illustrates the various ways methodological diversity exists. eBPF-based systems prioritize the various ways kernel-level telemetry captures system calls, and packet flows to provide a foundation for measurement with minimal overhead [<xref ref-type="bibr" rid="ref-33">33</xref>], distributed observability [<xref ref-type="bibr" rid="ref-34">34</xref>], and adaptive threat mitigation [<xref ref-type="bibr" rid="ref-28">28</xref>]. Conversely, deep learning frameworks in smart grids or manufacturing rely on the various ways sensor-level signals, such as image streams or voltage patterns, are processed locally [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. Health-oriented systems monitor the various ways biomedical signals support real-time diagnosis [<xref ref-type="bibr" rid="ref-32">32</xref>]. Different monitoring approaches focus on different goals. Some QoS-driven systems mainly track network metrics such as delay and jitter to check whether the network is stable [<xref ref-type="bibr" rid="ref-36">36</xref>]. This shows that data collection is usually designed to match the specific needs of each application domain.</p>
<p>A comparison of monitoring is done to cover the various ways of observability that are achieved across different layers. Some solutions focus on device-level monitoring, where raw data from sensors is processed; this would be placed as close to the device as possible [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. In contrast, other frameworks focus on the different ways of edge-node observability that capture metrics from the operating system and kernel [<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. There are also hybrid systems that monitor application behavior and service interactions across multiple layers [<xref ref-type="bibr" rid="ref-35">35</xref>]. Together, the approach in [<xref ref-type="bibr" rid="ref-36">36</xref>] highlights the need for monitoring all layers at the same time to stay accurate in changing environments.</p>
<p>Many systems also use layered architectures to spread computation efficiently. For example, smart-grid systems often use a hierarchy where edge devices handle quick analysis, and the cloud requires heavier processing [<xref ref-type="bibr" rid="ref-35">35</xref>]. Healthcare monitoring systems for elderly users use multistage pipelines that send data only when an unusual change is detected. eBPF-based systems focus on monitoring directly at the edge to gain detailed system insights without relying on other nodes [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>Mobility-aware systems mainly monitor the network layer to keep services stable as users move [<xref ref-type="bibr" rid="ref-36">36</xref>]. The differences among these approaches demonstrate how architectural choices are shaped by latency demands, resource constraints, and domain-level requirements. Deployment environments also vary widely. Healthcare systems often rely on embedded devices and smart-home nodes to reduce dependence on the cloud [<xref ref-type="bibr" rid="ref-12">12</xref>]. Vision-based uses powerful GPU-enabled devices to accelerate real-time detection [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. eBPF-based monitoring systems typically run on Linux systems that support high-performance data collection and observability [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. These factors illustrate the different ways edge architectures must adapt to the various types of hardware constraints and domain-specific requirements.</p>
<p>Finally, monitoring tools and architectural designs differ across systems. eBPF-based tools focus on low-overhead monitoring by working at the kernel level [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. AI-based approaches use compressed models so they can run on limited hardware. Despite these strengths, there are various ways limitations remain, such as the lack of end-to-end observability across all layers. These gaps reveal the various ways unified and security-aware frameworks are still needed to provide multi-layer visibility in the different types of resource-constrained edge environments.</p>
<p><italic>Evaluation of Observability Solutions: Tools, Features, and Performance Metrics</italic></p>
<p>Evaluating existing monitoring and observability solutions is important for understanding how well they work in real systems. How they performed in the real world and how suitable they are for modern edge-based microservices. Previous research has introduced many solutions that focus on areas such as performance monitoring, detecting anomalies, identifying faults, and adapting to changing conditions across edge and cloud environments. However, these solutions differ widely in terms of goals, the type of data they collect, the technologies they use, and how thoroughly they are evaluated. To clearly highlight these differences, a qualitative comparison of key edge monitoring and observability solutions is shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Qualitative comparison of edge monitoring solutions.</title>
</caption>
<table>
<colgroup>
<col align="center" width="40mm"/>
<col align="center" width="33mm"/>
<col align="center" width="26mm"/>
<col align="center" width="50mm"/> </colgroup>
<thead>
<tr>
<th>Paper Reference</th>
<th>Observability Purpose</th>
<th>Data Evaluated</th>
<th>Technology Used</th>
</tr>
</thead>
<tbody>
<tr>
<td>EC&#x2019;s Rise</td>
<td>PA, AH</td>
<td>Metrics, latency trends</td>
<td>Edge-cloud architecture, virtualization</td>
</tr>
<tr>
<td>EC Emergence</td>
<td>PA, AH</td>
<td>Metrics, system behavior</td>
<td>MEC architecture, distributed systems</td>
</tr>
<tr>
<td>Secure IoTEC</td>
<td>AD, RCA</td>
<td>Logs, security events</td>
<td>IoT security frameworks, encryption</td>
</tr>
<tr>
<td>eBPF Cloud Apps</td>
<td>PA, AD</td>
<td>Traces, syscall data</td>
<td>eBPF, Linux kernel</td>
</tr>
<tr>
<td>MEC Smart Grids</td>
<td>PA</td>
<td>Sensor data, grid metrics</td>
<td>Smart grid IoT, edge analytics</td>
</tr>
<tr>
<td>MestTool IRT</td>
<td>PA, AD</td>
<td>Fault traces, metrics</td>
<td>Real-time fault detection algorithms</td>
</tr>
<tr>
<td>QoS Monitoring</td>
<td>PA, RCA</td>
<td>QoS metrics</td>
<td>QoS models, MEC</td>
</tr>
<tr>
<td>Insulator Detection</td>
<td>AD</td>
<td>Image data</td>
<td>Computer vision, edge AI</td>
</tr>
<tr>
<td>Traffic Monitoring</td>
<td>PA, AD</td>
<td>CV event data</td>
<td>Computer vision, edge processing</td>
</tr>
<tr>
<td>Fault Impact</td>
<td>PA, AD</td>
<td>Impact metrics</td>
<td>Edge analytics</td>
</tr>
<tr>
<td>Lightweight Telemetry</td>
<td>PA</td>
<td>Telemetry metrics</td>
<td>Lightweight probes</td>
</tr>
<tr>
<td>EC eBPF Observability</td>
<td>PA, AD</td>
<td>Syscalls, traces</td>
<td>eBPF, kernel tracing</td>
</tr>
<tr>
<td>EC Fault Resilience</td>
<td>PA, RCA</td>
<td>Fault logs, impact metrics</td>
<td>Resilience modeling frameworks</td>
</tr>
<tr>
<td>EC ML Drift Detection</td>
<td>AD</td>
<td>Model drift metrics</td>
<td>ML models, drift analysis</td>
</tr>
<tr>
<td>EC QoS Mobility</td>
<td>PA</td>
<td>Mobility metrics</td>
<td>MEC mobility models</td>
</tr>
<tr>
<td>Federated EC Privacy</td>
<td>RCA, AD</td>
<td>Federated data</td>
<td>Federated learning, privacy models</td>
</tr>
<tr>
<td>Orchestration Monitoring</td>
<td>PA, RCA</td>
<td>Orchestration logs/metrics</td>
<td>Docker/K8s orchestration</td>
</tr>
<tr>
<td>EC DeSFAM</td>
<td>PA, AD, RCA</td>
<td>Syscall traces</td>
<td>eBPF, VAE &#x002B; Isolation Forest</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: PA &#x003D; Performance Analysis, AH &#x003D; Anomaly Handling, AD &#x003D; Anomaly Detection, RCA &#x003D; Root Cause Analysis.</p>
</table-wrap-foot>
</table-wrap>
<p>This comparison focuses on the various ways the primary observability purpose, the types of data evaluated, and the technologies employed interact within each approach. As shown in the synthesized data, early edge-centric studies primarily emphasize the various ways performance analysis and architectural understanding are achieved through the measurement of latency trends, system behavior, and resource utilization. These works largely rely on the various ways edge&#x2013;cloud architectures and distributed systems models address scalability and the different challenges of the various ways observability requirements change depending on the different application contexts.</p>
<p>Advanced approaches further integrate the various ways machine learning models and privacy-preserving mechanisms support the different processes of adaptive monitoring. Overall, the research highlights a clear progression from the various ways coarse-grained monitoring was once handled toward the different ways fine-grained, intelligent, and application-aware observability is achieved today. Following this qualitative assessment, a quantitative performance comparison of the various ways state-of-the-art solutions perform is provided in the subsequent analysis in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Quantitative comparison of edge monitoring and observability solutions.</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="35mm"/>
<col align="center" width="30mm"/>
<col align="center" width="30mm"/>
<col align="center" width="30mm"/>
<col align="center" width="30mm"/> </colgroup>
<thead>
<tr>
<th>Paper</th>
<th>Latency/Response Time</th>
<th>CPU/Memory Overhead</th>
<th>Bandwidth/ Transmission</th>
<th>Accuracy/Detection Metrics</th>
<th>Other Metrics/Notes</th>
</tr>
</thead>
<tbody>
<tr>
<td>Demystifying eBPF Performance</td>
<td>Microsecond-level syscall tracing; sub-microsecond observation latency</td>
<td>CPU overhead typically &#x003C; 5%</td>
<td>N/A</td>
<td>N/A</td>
<td>High-precision kernel instrumentation</td>
</tr>
<tr>
<td>Real-Time Monitoring &#x0026; Analysis (MessTool IRT)</td>
<td>UDP delay: 18&#x2013;44 &#x03BC;s; Modbus-TCP RTT &#x003C; 2 ms; kernel array read 3.9 ns; syscall 485 ns</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Provides nanosecond-level measurements</td>
</tr>
<tr>
<td>Mobility &#x0026; QoS Monitoring (ghBSRM-MEC)</td>
<td>Classifier training time: &#x2248; 2.17 s for 2000 samples</td>
<td>N/A</td>
<td>N/A</td>
<td>Reduces QoS-related monitoring inaccuracy</td>
<td>Uses probabilistic ratio &#x002B; KNN fallback</td>
</tr>
<tr>
<td>Edge Computing Vision &#x0026; Challenges</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Conceptual paper; no numerical benchmarks</td>
</tr>
<tr>
<td>Smart Traffic Monitoring System</td>
<td>Real-time responsiveness; edge inference reduces latency vs. cloud-only</td>
<td>N/A</td>
<td>Bandwidth greatly reduced (only processed features sent)</td>
<td>High detection precision reported</td>
<td>Edge&#x2013;cloud hybrid architecture</td>
</tr>
<tr>
<td>Real-Time Edge Computing Framework</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Provides architectural guidelines</td>
</tr>
<tr>
<td>DeSFAM (AI &#x002B; eBPF Adaptive Security)</td>
<td>Enforcement latency: 0.540 &#x03BC;s; policy updates 65&#x2013;82 &#x03BC;s</td>
<td>CPU &#x002B; 1.74%; memory reduced &#x007E;6.1% (&#x2248;50&#x2013;55 MB)</td>
<td>Monitoring cost reduced 32%</td>
<td>Accuracy 96%; F1 0.92; Precision 94%; Recall 90%; blocked 14/14 CVE attacks</td>
<td>Syscall reduction rate: 78.6%</td>
</tr>
<tr>
<td>Ultra-Lightweight Co-Inference Model</td>
<td>Low detection latency; training time few seconds per epoch</td>
<td>Compression ratio 0.341; quantized model runs at 6&#x2013;8 bits</td>
<td>N/A</td>
<td>F1: 0.88&#x2013;0.94 (dataset-dependent)</td>
<td>Works on extremely constrained devices</td>
</tr>
<tr>
<td>Insulator Self-Explosion Monitoring</td>
<td>Edge inference latency &#x2248; 30 ms (lightweight SSD model)</td>
<td>Edge device power: 25.6 W; standby 0.5 W</td>
<td>Large reduction due to local inference</td>
<td>Accuracy 95.75%; VAL 94.68%; AUC 0.9153</td>
<td>Uses NVIDIA Jetson TX2</td>
</tr>
<tr>
<td>eBPF Edge Node Real-Time Capability</td>
<td>Kernel tracing latency in microseconds</td>
<td>Almost zero overhead; minimal syscall perturbation</td>
<td>N/A</td>
<td>N/A</td>
<td>Shows eBPF suitability for real-time edge tasks</td>
</tr>
<tr>
<td>Emergence of Edge Computing</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Conceptual; describes ecosystem trends</td>
</tr>
<tr>
<td>Resource-Stressing Fault Analysis</td>
<td>Reports application latency variation under fault injection</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Identifies worst fault types for workloads</td>
</tr>
<tr>
<td>Secure IoT Service Architecture (Cloud &#x002B; Edge)</td>
<td>N/A</td>
<td>N/A</td>
<td>Communication reduced through hybrid routing</td>
<td>N/A</td>
<td>Provides secure multi-layer architecture</td>
</tr>
<tr>
<td>Extended BPF: Application Perspective</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>Conceptual paper on eBPF; no numerical evaluation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This analysis examines the various ways key metrics such as latency, CPU and memory overhead, and bandwidth consumption are measured across different frameworks. As illustrated in the performance data, kernel-level instrumentation techniques provide a foundation for the various ways microsecond-level latency is achieved with minimal performance overhead. This makes them highly suitable for the various ways real-time observability is maintained in resource-constrained edge nodes. These monitoring frameworks demonstrate the various ways latency is significantly reduced compared to cloud-only pipelines by performing the various processes of local processing and inference directly at the edge, thereby lowering communication overhead.</p>
<p>Solutions that integrate the various ways of artificial intelligence exhibit high detection accuracy and provide different ways for effective anomaly identification while maintaining a limited state of resource utilization. In particular, the various ways lightweight inference models and adaptive security mechanisms are combined demonstrate a favorable balance between the measurement of detection performance and computational cost.</p>
<p>However, the quantitative results also reveal the various ways many conceptual and architectural studies lack the necessary empirical benchmarking, which limits the different ways they can be applied to performance-critical deployments. Furthermore, some solutions report the various ways strong accuracy metrics are achieved but require specialized hardware or kernel-level access, which can restrict the various ways portability and scalability.</p>
<p>Together, these findings present a more comprehensive assessment that integrates the many ways in which qualitative architectural insights and quantitative performance evidence complement each other. As a side effect of this double analysis, we demonstrate the nature of different kinds of limitations in state-of-the-art approaches, and also peel off how large-scale validation and multi-node evaluation are still not enough. These gaps provide a foundation for the need for an observability framework that delivers the various ways of fine-grained visibility and adaptability across the different types of edge microservice environments.</p>
<p>Additionally, domain-specific solutions&#x2014;including smart grids, traffic monitoring, and healthcare that demonstrate.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Open Research Issues and Challenges in Edge Computing</title>
<p>Edge computing is a speedily growing field with significant potential to transform the way data is handled and delivered. However, it also presents numerous research challenges and issues that must to be addressed for its effective execution and optimization. Here are some of the key open research issues and challenges in edge computing.</p>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> presents a consolidated view of the major open research issues and challenges in edge computing identified through the literature review and analytical evaluation conducted in this study. The diagram highlights six interrelated challenge domains, including comprehensive observability across edge layers, unified telemetry collection and cross-layer data fusion, scalability and real-time processing constraints, adaptability to mobility and dynamic edge environments, privacy and security challenges across edge layers, and emerging research opportunities and future directions. These challenges reflect the complexity of edge computing systems, which are characterized by various challenges such as heterogeneity, decentralization, and strict latency requirements. The conceptual roadmap provides a foundation for the various ways each challenge is examined in depth, particularly regarding its impact on the different ways edge observability and performance monitoring are evaluated across smart scheduling mechanisms.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Key open research challenges in edge computing include observability, telemetry, scalability, mobility, and security.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80115-fig-3.tif"/>
</fig>
<sec id="s5_1">
<label>5.1</label>
<title>Comprehensive Observability across Edge Layers</title>
<p>Even though edge computing approaches have emerged, the problem of achieving end-to-end visibility has proven to be difficult. To date, much of the research is siloed, looking at the health of the system at the infrastructure layer or monitoring the cloud, rather than providing a unified view over the entire data flow, as performance problems are rarely contained within a single component. These problems could be at the operating system, network, or application layer. To detect problems in real time, a monitoring framework should be multi-layered and should be able to analyze and reason on real-time information from all layers [<xref ref-type="bibr" rid="ref-69">69</xref>]. At the same time, observability should be provided without overloading resource-constrained edge devices, which are typically battery-powered and computationally limited. Solving this problem would enable accurate observability for mobile edge applications without sacrificing performance and battery lifetime.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Unified Telemetry Collection and Cross-Layer Data Fusion</title>
<p>Edge data may include low-level telemetry streams such as kernel traces, but also high-level semantics such as mobility. Due to the wide variety in the format and granularity, we often see these monitoring pipelines become siloed, which limits their usefulness in analyzing the edges. Without a common way to collect and fuse this data in real time, accurate root cause analysis and diagnosis of small performance drifts remain beyond the capabilities of current tools. Future research needs integrated architectures that can fuse this multi-source data in real time to address the limitations of current solutions [<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-69">69</xref>]. We need a platform that reliably collects data and has the ability to correlate data to see where jitters or delays in scheduling or processing are occurring.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Scalability and Real-Time Processing Constraints</title>
<p>As edge deployments grow in scale, observability systems face a tradeoff between the telemetry emitted by the edge nodes and resource consumption on each node, which is typically very low. Many applications that require real-time monitoring have strict latency requirements that are not compatible with customary monitoring approaches that are too &#x201C;heavy&#x201D;. The problem is that the high frequency and accuracy of the trace, i.e., monitoring, can interfere with the task&#x2019;s performance itself [<xref ref-type="bibr" rid="ref-40">40</xref>]. To solve this problem, ultra-low-overhead observability with sufficient precision, without sacrificing responsiveness, is likely needed. Adaptive sampling and event-driven tracing appear to be the most promising approaches in this regard. Thus, frameworks like this could be used to provide high granularity at high load while imposing little or no overhead when the system is not stressed.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Adaptability to Mobility and Dynamic Edge Environments</title>
<p>Mobility may add a type of uncertainty that makes it difficult to measure network latency and link stability. Most of the existing solutions assume some other kind of stable condition, and therefore are not able to quickly adapt to mobility [<xref ref-type="bibr" rid="ref-70">70</xref>]. This entails considering diverse mobility-aware observability formulations that combine real-time and predictive perspectives, as well as considering various forms of proactive changes in the scheduling and tracing depth relative to the different forms of changes in the environment.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Privacy and Security Challenges across Edge Layers</title>
<p>The deeper the observability mechanisms are, the more types of interaction involving privacy and security issues have to be considered. The telemetries collected from the devices can reveal the different types of sensitive user activities and location patterns, which call for the various types of strong privacy-preserving mechanisms. Different types of instrumentation techniques that involve the kernel can shed more light on different types of attack surface, which can be addressed by developing different types of secure observability frameworks with encrypted transport, as well as by employing different types of anonymization to maintain the different types of analytical value [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Emerging Research Opportunities and Future Directions</title>
<p>Aside from the above-mentioned challenges, there are many intersections between observability and mobility-aware edge computing, such as self-adaptive observability systems and telemetry control systems that are able to follow dynamic workloads. Interestingly, a future research direction can include investigating the type of causal relations that exist over an entire pipeline and the performance anomaly metrics these relations influence. Future work could additionally consider the core observability and type of schedulers used, adaptive task offloading, as well as resource allocation achievable.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>Edge computing is increasingly becoming a core paradigm for latency-sensitive and data-intensive applications. Its proximity to the end users presents unique monitoring and observability challenges in guaranteeing system reliability over a highly distributed and heterogeneous environment. In contrast to homogeneous, elastic infrastructures underlying conventional cloud-centric approaches, edge environments are characterized by limited computational resources, limited energy resources, shared and constrained networking resources, and stringent real-time guarantees. In those regards, a conceptualization of monitoring and observability in the context of the edge is described.</p>
<p>This work provides a thorough taxonomy of the state-of-the-art edge monitoring and observability solutions based on their optimization objectives, observability scope, telemetry characteristics, and deployment architectures. The former is characterized by tracing the source code at the kernel level with relatively low overhead, and the latter is typified by learning-based observability that is semantically richer but computationally more expensive. The differences between existing solutions emphasize the need for dynamic observability frameworks that consider edge resource constraints.</p>
<p>From the comparative analysis have been identified in current research. The comparative analysis reveals a notable number of solutions and proposals that focus solely on one aspect, either the effectiveness of anomaly detection or performance monitoring, without cross-layer visibility or a holistic view. This impedes correlating system behavior across services, infrastructure, and network layers for root cause analysis. The diversity of deployment environments, ranging from more powerful fog nodes to much more constrained edge devices, still poses portability and scalability challenges. Many solutions include certain domain-specific assumptions, which limit generalizability with insufficiently addressed issues like resource overhead, mobility awareness, and diversified telemetry integration.</p>
<p>From a design perspective, the proposed taxonomy offers important guidance for future edge monitoring systems by pinpointing neglected dimensions and the most promising combinations of techniques. More specifically, such a system integration strategy should represent future research priorities centered on integrating cross-layer observability with lightweight telemetry mechanisms to provide unified system/service/network visibility. The coupling of kernel instrumentation under a low footprint (e.g., eBPF-based techniques) with intelligent adaptive analytics is a promising candidate in this regard. In the same way, using mobility-aware monitoring and standardized telemetry pipelines can bring great improvements to how adaptable and scalable the system is in different dynamic environments. These insights indicate that less explored dimensions and promising combinations of techniques provide valuable design guidance for future edge monitoring systems.</p>
<p>Overall, the findings emphasize that effective edge observability requires a holistic approach that goes beyond isolated metrics, logs, and traces that can be correlated across multiple layers of the system stack. These findings emphasize that effective edge observability requires a holistic approach that goes beyond isolated metrics, logs, and traces to enable their correlation across multiple layers of the system stack. This work serves to inform design priorities based on relevant gaps identified in existing research to accomplish robust, scalable, and intelligent observability solutions in the future on the evolving edge computing landscape.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their sincere appreciation to the FAST School of Computing, National University of Computer and Emerging Sciences, Karachi, 75030, Pakistan, for the institutional support and resources that enabled this research.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Hamza Ahmed: Contributed to conceptualization, literature review, and initial manuscript drafting. Hassan Jamil Syed: Involved in designed the methodology, performed data analysis, reviewed and edited the manuscript, and supervised the project. Aqsa Aslam: Provide the key supervison and project supervision, help in study validation and provide the final manuscript revision. Sehar Zehra: Assisted in software review, figure and table preparation. Ummay Faseeha: Provided additional support in data verification, literature review, and manuscript formatting. Nurzati Iwani Othman: Contributed to conceptual guidance, provided critical insights into the research direction, and supported manuscript review and refinement to enhance the overall quality of the study. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jararweh</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Doulat</surname> <given-names>A</given-names></string-name>, <string-name><surname>AlQudah</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>E</given-names></string-name>, <string-name><surname>Al-Ayyoub</surname> <given-names>M</given-names></string-name>, <string-name><surname>Benkhelifa</surname> <given-names>E</given-names></string-name></person-group>. <article-title>The future of mobile cloud computing: integrating cloudlets and mobile edge computing</article-title>. In: <conf-name>Proceedings of the 2016 23rd International Conference on Telecommunications (ICT); 2016 May 16&#x2013;18</conf-name>; <publisher-loc>Thessaloniki, Greece</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/ict.2016.7500486</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choudhury</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sarma</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Gulvanskii</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kaplun</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dutta</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Leveraging federated learning and edge computing for pandemic-resilient healthcare</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>20497</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-00199-9</pub-id>; <pub-id pub-id-type="pmid">40593962</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Doghman</surname> <given-names>F</given-names></string-name>, <string-name><surname>Moustafa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Khalil</surname> <given-names>I</given-names></string-name>, <string-name><surname>Sohrabi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Tari</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zomaya</surname> <given-names>AY</given-names></string-name></person-group>. <article-title>AI-enabled secure microservices in edge computing: opportunities and challenges</article-title>. <source>IEEE Trans Serv Comput</source>. <year>2023</year>;<volume>16</volume>(<issue>2</issue>):<fpage>1485</fpage>&#x2013;<lpage>504</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsc.2022.3155447</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Dynamic management network system of automobile detection applying edge computing</article-title>. <source>Int J Netw Manag</source>. <year>2023</year>;<volume>33</volume>(<issue>5</issue>):<fpage>e2231</fpage>. doi:<pub-id pub-id-type="doi">10.1002/nem.2231</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Syafrudin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fitriyani</surname> <given-names>NL</given-names></string-name>, <string-name><surname>Alfian</surname> <given-names>G</given-names></string-name>, <string-name><surname>Rhee</surname> <given-names>J</given-names></string-name></person-group>. <article-title>An affordable fast early warning system for edge computing in assembly line</article-title>. <source>Appl Sci</source>. <year>2019</year>;<volume>9</volume>(<issue>1</issue>):<fpage>84</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app9010084</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Stojmenovic</surname> <given-names>I</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>S</given-names></string-name></person-group>. <article-title>The fog computing paradigm: scenarios and security issues</article-title>. <source>Proc 2014 Fed Conf Comput Sci Inf Syst</source>. <year>2014</year>;<volume>2</volume>:<fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.15439/2014f503</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zissis</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lekkas</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Addressing cloud computing security issues</article-title>. <source>Future Gener Comput Syst</source>. <year>2012</year>;<volume>28</volume>(<issue>3</issue>):<fpage>583</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2010.12.006</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dustdar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zomaya</surname> <given-names>AY</given-names></string-name></person-group>. <article-title>Edge intelligence: the confluence of edge computing and artificial intelligence</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>7</volume>(<issue>8</issue>):<fpage>7457</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2020.2984887</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Raza</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Anwar</surname> <given-names>MR</given-names></string-name></person-group>. <article-title>A survey on vehicular edge computing: architecture, applications, technical issues, and future directions</article-title>. <source>Wirel Commun Mob Comput</source>. <year>2019</year>;<volume>2019</volume>(<issue>1</issue>):<fpage>3159762</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2019/3159762</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bulkan</surname> <given-names>U</given-names></string-name>, <string-name><surname>Dagiuklas</surname> <given-names>T</given-names></string-name>, <string-name><surname>Iqbal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Huq</surname> <given-names>KMS</given-names></string-name>, <string-name><surname>Al-Dulaimi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rodriguez</surname> <given-names>J</given-names></string-name></person-group>. <article-title>On the load balancing of edge computing resources for on-line video delivery</article-title>. <source>IEEE Access</source>. <year>2018</year>;<volume>6</volume>:<fpage>73916</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2018.2883319</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kyryk</surname> <given-names>M</given-names></string-name>, <string-name><surname>Pleskanka</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pleskanka</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nykonchuk</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Load balancing method in edge computing</article-title>. In: <conf-name>Proceedings of the 2020 IEEE 15th International Conference on Advanced Trends in Radioelectronics, Telecommunications and Computer Engineering (TCSET); 2020 Feb 25&#x2013;29</conf-name>; <publisher-loc>Lviv-Slavske, Ukraine</publisher-loc>. p. <fpage>978</fpage>&#x2013;<lpage>81</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bhuiyan</surname> <given-names>MZA</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>A secure IoT service architecture with an efficient balance dynamics based on cloud and edge computing</article-title>. <source>IEEE Internet Things J</source>. <year>2019</year>;<volume>6</volume>(<issue>3</issue>):<fpage>4831</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2018.2870288</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kiani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khreishah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ansari</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Smart traffic monitoring system using computer vision and edge computing</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2022</year>;<volume>23</volume>(<issue>8</issue>):<fpage>12027</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2021.3109481</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Bhuiyan</surname> <given-names>MZA</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Big data reduction for a smart city&#x2019;s critical infrastructural health monitoring</article-title>. <source>IEEE Commun Mag</source>. <year>2018</year>;<volume>56</volume>(<issue>3</issue>):<fpage>128</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mcom.2018.1700303</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dave</surname> <given-names>R</given-names></string-name>, <string-name><surname>Seliya</surname> <given-names>N</given-names></string-name>, <string-name><surname>Siddiqui</surname> <given-names>N</given-names></string-name></person-group>. <article-title>The benefits of edge computing in healthcare, smart cities, and IoT</article-title>. <comment>arXiv:2112.01250. 2021</comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hassan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yau</surname> <given-names>KA</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Edge computing in 5G: a review</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>127276</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2938534</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shou</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Toward edge intelligence: multiaccess edge computing for 5G and Internet of Things</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>7</volume>(<issue>8</issue>):<fpage>6722</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2020.3004500</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Soldani</surname> <given-names>D</given-names></string-name>, <string-name><surname>Nahi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bour</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jafarizadeh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Soliman</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Di Giovanna</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>eBPF: a new approach to cloud-native observability, networking and security for current (5G) and future mobile networks (6G and beyond)</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>57174</fpage>&#x2013;<lpage>202</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2023.3281480</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Mobile edge computing towards 5G: vision, recent progress, and open challenges</article-title>. <source>China Commun</source>. <year>2016</year>;<volume>13</volume>(<issue>Suppl 2</issue>):<fpage>89</fpage>&#x2013;<lpage>99</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cc.2016.7833463</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Celesti</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Abbas</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Edge computing in IoT-based manufacturing</article-title>. <source>IEEE Commun Mag</source>. <year>2018</year>;<volume>56</volume>(<issue>9</issue>):<fpage>103</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mcom.2018.1701231</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hassan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Gillani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>E</given-names></string-name>, <string-name><surname>Yaqoob</surname> <given-names>I</given-names></string-name>, <string-name><surname>Imran</surname> <given-names>M</given-names></string-name></person-group>. <article-title>The role of edge computing in Internet of Things</article-title>. <source>IEEE Commun Mag</source>. <year>2018</year>;<volume>56</volume>(<issue>11</issue>):<fpage>110</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mcom.2018.1700906</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>PAWSSP: a two-stage parallelism-aware algorithm for joint workflow scheduling and service placement in edge computing</article-title>. <source>IEEE Trans Serv Comput</source>. <year>2026</year>;<volume>19</volume>(<issue>1</issue>):<fpage>572</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsc.2025.3643326</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>An overview on edge computing research</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>85714</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2020.2991734</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>W</given-names></string-name></person-group>. <chapter-title>Challenges and opportunities in edge computing</chapter-title>. In: <source>Edge computing: a primer</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2018</year>. p. <fpage>59</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-02083-5_5</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gro&#x00DF;mann</surname> <given-names>M</given-names></string-name>, <string-name><surname>Schenk</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A comparison of monitoring approaches for virtualized services at the network edge</article-title>. In: <conf-name>Proceedings of the 2018 International Conference on Internet of Things, Embedded Systems and Communications (IINTEC); 2018 Dec 20&#x2013;21</conf-name>; <publisher-loc>Hamammet, Tunisia</publisher-loc>. p. <fpage>85</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Edge computing for Internet of everything: a survey</article-title>. <source>IEEE Internet Things J</source>. <year>2022</year>;<volume>9</volume>(<issue>23</issue>):<fpage>23472</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2022.3200431</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pydi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Iyer</surname> <given-names>GN</given-names></string-name></person-group>. <article-title>Analytical review and study on load balancing in edge computing platform</article-title>. In: <conf-name>Proceedings of the 2020 Fourth International Conference on Computing Methodologies and Communication (ICCMC); 2020 Mar 11&#x2013;13</conf-name>; <publisher-loc>Erode, India</publisher-loc>. p. <fpage>180</fpage>&#x2013;<lpage>87</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zehra</surname> <given-names>S</given-names></string-name>, <string-name><surname>Syed</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Samad</surname> <given-names>F</given-names></string-name>, <string-name><surname>Faseeha</surname> <given-names>U</given-names></string-name></person-group>. <article-title>DeSFAM: an adaptive eBPF and AI-driven framework for securing cloud containers in real time</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>139203</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3592192</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zehra</surname> <given-names>S</given-names></string-name>, <string-name><surname>Syed</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Samad</surname> <given-names>F</given-names></string-name>, <string-name><surname>Faseeha</surname> <given-names>U</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khurram Khan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Securing the shared kernel: exploring kernel isolation and emerging challenges in modern cloud computing</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>179281</fpage>&#x2013;<lpage>317</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3507215</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Container scheduling strategy based on image layer reuse and sequential arrangement in mobile edge computing</article-title>. <source>IEEE Trans Mobile Comput</source>. <year>2025</year>;<volume>24</volume>(<issue>9</issue>):<fpage>8700</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmc.2025.3557160</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Satyanarayanan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>The emergence of edge computing</article-title>. <source>Computer</source>. <year>2017</year>;<volume>50</volume>(<issue>1</issue>):<fpage>30</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mc.2017.9</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Edge computing: vision and challenges</article-title>. <source>IEEE Internet Things J</source>. <year>2016</year>;<volume>3</volume>(<issue>5</issue>):<fpage>637</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2016.2579198</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gowtham</surname> <given-names>V</given-names></string-name>, <string-name><surname>Keil</surname> <given-names>O</given-names></string-name>, <string-name><surname>Yeole</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schreiner</surname> <given-names>F</given-names></string-name>, <string-name><surname>Tschoke</surname> <given-names>S</given-names></string-name>, <string-name><surname>Willner</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Determining edge node real-time capabilities</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/ACM 25th International Symposium on Distributed Simulation and Real Time Applications (DS-RT); 2021 Sep 27&#x2013;29</conf-name>; <publisher-loc>Valencia, Spain</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/ds-rt52167.2021.9576132</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shahinfar</surname> <given-names>F</given-names></string-name>, <string-name><surname>Miano</surname> <given-names>S</given-names></string-name>, <string-name><surname>Panda</surname> <given-names>A</given-names></string-name>, <string-name><surname>Antichi</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Demystifying performance of eBPF network applications</article-title>. <source>Proc ACM Netw</source>. <year>2025</year>;<volume>3</volume>:<fpage>1</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3749216</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Long</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Real-time monitoring and analysis of track and field athletes based on edge computing and deep reinforcement learning algorithm</article-title>. <source>Alex Eng J</source>. <year>2025</year>;<volume>114</volume>(<issue>7</issue>):<fpage>136</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aej.2024.11.024</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mr krishnamoorthy</surname> <given-names>J</given-names></string-name>, <string-name><surname>Egneswari</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Mobility and dependence-aware QoS monitoring in mobile edge computing</article-title>. <source>J Adv Zool</source>. <year>2023</year>;<volume>44</volume>(<issue>3</issue>):<fpage>386</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.17762/jaz.v44i3.653</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Online monitoring method for insulator self-explosion based on edge computing and deep learning</article-title>. <source>CSEE J Power Energy Syst</source>. <year>2021</year>;<volume>8</volume>(<issue>6</issue>):<fpage>1684</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.17775/cseejpes.2020.05910</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fournier</surname> <given-names>G</given-names></string-name>, <string-name><surname>Afchain</surname> <given-names>S</given-names></string-name>, <string-name><surname>Baubeau</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Runtime security monitoring with eBPF</article-title>. <year>2021 [cited 2026 Jan 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.sstic.org/media/SSTIC2021/SSTIC-actes/runtime_security_with_ebpf/SSTIC2021-Article-runtime_security_with_ebpf-fournier_afchain_baubeau.pdf">https://www.sstic.org/media/SSTIC2021/SSTIC-actes/runtime_security_with_ebpf/SSTIC2021-Article-runtime_security_with_ebpf-fournier_afchain_baubeau.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Otero</surname> <given-names>M</given-names></string-name>, <string-name><surname>Garc&#x00ED;a</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Fernandez</surname> <given-names>P</given-names></string-name></person-group>. <article-title>An extensible lightweight framework for distributed telemetry of microservices</article-title>. <source>Sustain Comput Inform Syst</source>. <year>2025</year>;<volume>46</volume>(<issue>5</issue>):<fpage>101100</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.suscom.2025.101100</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Morabito</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ficara</surname> <given-names>A</given-names></string-name>, <string-name><surname>Celesti</surname> <given-names>A</given-names></string-name>, <string-name><surname>Villari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fazio</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Consensus-based distributed orchestration framework for microservices in edge computing clusters</article-title>. <source>Future Gener Comput Syst</source>. <year>2026</year>;<volume>176</volume>(<issue>1</issue>):<fpage>108221</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2025.108221</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Leung</surname> <given-names>VCM</given-names></string-name></person-group>. <article-title>An edge computing framework for real-time monitoring in smart grid</article-title>. In: <conf-name>Proceedings of the 2018 IEEE International Conference on Industrial Internet (ICII); 2018 Oct 21&#x2013;23</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/icii.2018.00019</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Mobility and dependence-aware QoS monitoring in mobile edge computing</article-title>. <source>IEEE Trans Cloud Comput</source>. <year>2021</year>;<volume>9</volume>(<issue>3</issue>):<fpage>1143</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tcc.2021.3063050</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arroba</surname> <given-names>P</given-names></string-name>, <string-name><surname>Buyya</surname> <given-names>R</given-names></string-name>, <string-name><surname>C&#x00E1;rdenas</surname> <given-names>R</given-names></string-name>, <string-name><surname>Risco-Mart&#x00ED;n</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Moya</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Sustainable edge computing: challenges and future directions</article-title>. <source>Softw Pract Exp</source>. <year>2024</year>;<volume>54</volume>(<issue>11</issue>):<fpage>2272</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1002/spe.3340</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Samanta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Esposito</surname> <given-names>F</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>TG</given-names></string-name></person-group>. <article-title>Fault-tolerant mechanism for edge-based IoT networks with demand uncertainty</article-title>. <source>IEEE Internet Things J</source>. <year>2021</year>;<volume>8</volume>(<issue>23</issue>):<fpage>16963</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2021.3075681</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>EdgeFD: an edge-friendly drift-aware fault diagnosis system for industrial IoT</article-title>. In: <conf-name>Proceedings of the 2023 IEEE 23rd International Conference on Communication Technology (ICCT); 2023 Oct 20&#x2013;22</conf-name>; <publisher-loc>Wuxi, China</publisher-loc>. p. <fpage>390</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icct59356.2023.10419797</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Vijayakumar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Karuppiah</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Privacy-preserving federated learning for Internet of medical things under edge computing</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2023</year>;<volume>27</volume>(<issue>2</issue>):<fpage>854</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2022.3157725</pub-id>; <pub-id pub-id-type="pmid">35259124</pub-id></mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nouri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bozga</surname> <given-names>M</given-names></string-name>, <string-name><surname>Molnos</surname> <given-names>A</given-names></string-name>, <string-name><surname>Legay</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bensalem</surname> <given-names>S</given-names></string-name></person-group>. <article-title><italic>ASTROLABE</italic>: a rigorous approach for system-level performance modeling and analysis</article-title>. <source>ACM Trans Embed Comput Syst</source>. <year>2016</year>;<volume>15</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2885498</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Van Renesse</surname> <given-names>R</given-names></string-name>, <string-name><surname>Birman</surname> <given-names>KP</given-names></string-name>, <string-name><surname>Vogels</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Astrolabe: a robust and scalable technology for distributed system monitoring, management, and data mining</article-title>. <source>ACM Trans Comput Syst</source>. <year>2003</year>;<volume>21</volume>(<issue>2</issue>):<fpage>164</fpage>&#x2013;<lpage>206</lpage>. doi:<pub-id pub-id-type="doi">10.1145/762483.762485</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yalagandula</surname> <given-names>P</given-names></string-name>, <string-name><surname>Dahlin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A scalable distributed information management system</article-title>. <source>SIGCOMM Comput Commun Rev</source>. <year>2004</year>;<volume>34</volume>(<issue>4</issue>):<fpage>379</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1145/1030194.1015509</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brand&#x00F3;n</surname> <given-names>&#x00C1;</given-names></string-name>, <string-name><surname>P&#x00E9;rez</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Montes</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sanchez</surname> <given-names>A</given-names></string-name></person-group>. <article-title>FMonE: a flexible monitoring solution at the edge</article-title>. <source>Wirel Commun Mob Comput</source>. <year>2018</year>;<volume>2018</volume>(<issue>1</issue>):<fpage>2068278</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2018/2068278</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rac</surname> <given-names>S</given-names></string-name>, <string-name><surname>Brorsson</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Cost-aware service placement and scheduling in the edge-cloud continuum</article-title>. <source>ACM Trans Archit Code Optim</source>. <year>2024</year>;<volume>21</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3640823</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Taherizadeh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Taylor</surname> <given-names>I</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Stankovski</surname> <given-names>V</given-names></string-name></person-group>. <chapter-title>A network edge monitoring approach for real-time data streaming applications</chapter-title>. In: <source>Economics of grids, clouds, systems, and services</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2017</year>. p. <fpage>293</fpage>&#x2013;<lpage>303</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-61920-0_21</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Syed</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Sadiq</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ibrahim</surname> <given-names>AO</given-names></string-name>, <string-name><surname>Alohaly</surname> <given-names>M</given-names></string-name>, <string-name><surname>Elsadig</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Exploring performance degradation in virtual machines sharing a cloud server</article-title>. <source>Appl Sci</source>. <year>2023</year>;<volume>13</volume>(<issue>16</issue>):<fpage>9224</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app13169224</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Syed</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Gani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nasaruddin</surname> <given-names>FH</given-names></string-name>, <string-name><surname>Naveed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>AIA</given-names></string-name>, <string-name><surname>Khurram Khan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>CloudProcMon: a non-intrusive cloud monitoring framework</article-title>. <source>IEEE Access</source>. <year>2018</year>;<volume>6</volume>:<fpage>44591</fpage>&#x2013;<lpage>606</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2018.2864573</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iqbal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Naqvi</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Naveed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Usman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>HA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LDMRes-Net: a lightweight neural network for efficient medical image segmentation on IoT and edge devices</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2023</year>;<volume>28</volume>(<issue>7</issue>):<fpage>3860</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2023.3331278</pub-id>; <pub-id pub-id-type="pmid">37938951</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Korenivska</surname> <given-names>OL</given-names></string-name>, <string-name><surname>Benedytskyi</surname> <given-names>VB</given-names></string-name>, <string-name><surname>Andreiev</surname> <given-names>OV</given-names></string-name>, <string-name><surname>Medvediev</surname> <given-names>MG</given-names></string-name></person-group>. <article-title>A system for monitoring the microclimate parameters of premises based on the Internet of Things and edge devices</article-title>. <source>J Edge Comp</source>. <year>2023</year>;<volume>2</volume>(<issue>2</issue>):<fpage>125</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.55056/jec.614</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>An</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Edge computing on IoT for machine signal processing and fault diagnosis: a review</article-title>. <source>IEEE Internet Things J</source>. <year>2023</year>;<volume>10</volume>(<issue>13</issue>):<fpage>11093</fpage>&#x2013;<lpage>116</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2023.3239944</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rajagopal</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Supriya</surname> <given-names>M</given-names></string-name>, <string-name><surname>Buyya</surname> <given-names>R</given-names></string-name></person-group>. <article-title>FedSDM: federated learning based smart decision making module for ECG data in IoT integrated Edge&#x2013;Fog&#x2013;Cloud computing environments</article-title>. <source>Internet Things</source>. <year>2023</year>;<volume>22</volume>(<issue>3</issue>):<fpage>100784</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2023.100784</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Taherizadeh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Taylor</surname> <given-names>I</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Stankovski</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Monitoring self-adaptive applications within edge computing frameworks: a state-of-the-art review</article-title>. <source>J Syst Softw</source>. <year>2018</year>;<volume>136</volume>(<issue>9</issue>):<fpage>19</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jss.2017.10.033</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Habib ur Rehman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jayaraman</surname> <given-names>P</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>S</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Medhat Gaber</surname> <given-names>M</given-names></string-name></person-group>. <article-title>RedEdge: a novel architecture for big data processing in mobile edge computing environments</article-title>. <source>J Sens Actuator Netw</source>. <year>2017</year>;<volume>6</volume>(<issue>3</issue>):<fpage>17</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jsan6030017</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mehmood</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khalid</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kostakos</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gilman</surname> <given-names>E</given-names></string-name>, <string-name><surname>Pirttikangas</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A novel Edge architecture and solution for detecting concept drift in smart environments</article-title>. <source>Future Gener Comput Syst</source>. <year>2024</year>;<volume>150</volume>(<issue>5</issue>):<fpage>127</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2023.08.023</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Bhuiyan</surname> <given-names>MZA</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A novel trust mechanism based on fog computing in sensor&#x2013;cloud system</article-title>. <source>Future Gener Comput Syst</source>. <year>2020</year>;<volume>109</volume>(<issue>6</issue>):<fpage>573</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2018.05.049</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>DG</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>JX</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A novel edge computing architecture based on adaptive stratified sampling</article-title>. <source>Comput Commun</source>. <year>2022</year>;<volume>183</volume>(<issue>5</issue>):<fpage>121</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comcom.2021.11.012</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bonomi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Milito</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Addepalli</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Fog computing and its role in the internet of things</article-title>. In: <conf-name>Proceedings of the First Edition of the MCC Workshop on Mobile Cloud Computing; 2012 Aug 17</conf-name>; <publisher-loc>Helsinki, Finland</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1145/2342509.2342513</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Faseeha</surname> <given-names>U</given-names></string-name>, <string-name><surname>Jamil Syed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Samad</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zehra</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Observability in microservices: an in-depth exploration of frameworks, challenges, and deployment paradigms</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>1</issue>):<fpage>72011</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3562125</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Abderrahim</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ouzzif</surname> <given-names>M</given-names></string-name>, <string-name><surname>Guillouard</surname> <given-names>K</given-names></string-name>, <string-name><surname>Francois</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lebre</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A holistic monitoring service for fog/edge infrastructures: a foresight study</article-title>. In: <conf-name>Proceedings of the 2017 IEEE 5th International Conference on Future Internet of Things and Cloud (FiCloud); 2017 Aug 21&#x2013;23</conf-name>; <publisher-loc>Prague, Czech Republic</publisher-loc>. p. <fpage>337</fpage>&#x2013;<lpage>44</lpage>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gaddam</surname> <given-names>MK</given-names></string-name></person-group>. <article-title>Architecting observability for AI-driven microservices at scale</article-title>. In: <conf-name>2025 3rd International Conference on Intelligent Cyber Physical Systems and Internet of Things (ICoICI); 2025 Sep 17&#x2013;19</conf-name>; <publisher-loc>Coimbatore, India</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/icoici65217.2025.11252857</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nimmagadda</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Applying AI/ML to Kubernetes logging and monitoring in enhancing observability through intelligent systems</article-title>. <source>Eur J Comput Sci Inf Technol</source>. <year>2025</year>;<volume>13</volume>(<issue>49</issue>):<fpage>141</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.37745/ejcsit.2013/vol13n49141152</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hudson</surname> <given-names>N</given-names></string-name>, <string-name><surname>Khamfroush</surname> <given-names>H</given-names></string-name>, <string-name><surname>Baughman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lucani</surname> <given-names>DE</given-names></string-name>, <string-name><surname>Chard</surname> <given-names>K</given-names></string-name>, <string-name><surname>Foster</surname> <given-names>I</given-names></string-name></person-group>. <article-title>QoS-aware edge AI placement and scheduling with multiple implementations in FaaS-based edge computing</article-title>. <source>Future Gener Comput Syst</source>. <year>2024</year>;<volume>157</volume>(<issue>8</issue>):<fpage>250</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2024.03.035</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Resource allocation for twin maintenance and computing task processing in digital twin vehicular edge computing network</article-title>. <comment>arXiv:2407.07575. 2024</comment>.</mixed-citation></ref>
</ref-list>
</back></article>