<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">62819</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.062819</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Literature Review on Model Conversion, Inference, and Learning Strategies in EdgeML with TinyML Deployment</article-title>
<alt-title alt-title-type="left-running-head">A Literature Review on Model Conversion, Inference, and Learning Strategies in EdgeML with TinyML Deployment</alt-title>
<alt-title alt-title-type="right-running-head">A Literature Review on Model Conversion, Inference, and Learning Strategies in EdgeML with TinyML Deployment</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Arif</surname><given-names>Muhammad</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>mahamid@uqu.edu.sa</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Rashid</surname><given-names>Muhammad</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science and Artificial Intelligence, Umm Al-Qura University</institution>, <addr-line>Makkah Al-Mukarama, 21955</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer and Network Engineering, Umm Al-Qura University</institution>, <addr-line>Makkah Al-Mukarama, 21955</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Muhammad Arif. Email: <email>mahamid@uqu.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>26</day><month>03</month><year>2025</year>
</pub-date>
<volume>83</volume>
<issue>1</issue>
<fpage>13</fpage>
<lpage>64</lpage>
<history>
<date date-type="received">
<day>28</day>
<month>12</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>2</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_62819.pdf"></self-uri>
<abstract>
<p>Edge Machine Learning (EdgeML) and Tiny Machine Learning (TinyML) are fast-growing fields that bring machine learning to resource-constrained devices, allowing real-time data processing and decision-making at the network&#x2019;s edge. However, the complexity of model conversion techniques, diverse inference mechanisms, and varied learning strategies make designing and deploying these models challenging. Additionally, deploying TinyML models on resource-constrained hardware with specific software frameworks has broadened EdgeML&#x2019;s applications across various sectors. These factors underscore the necessity for a comprehensive literature review, as current reviews do not systematically encompass the most recent findings on these topics. Consequently, it provides a comprehensive overview of state-of-the-art techniques in model conversion, inference mechanisms, learning strategies within EdgeML, and deploying these models on resource-constrained edge devices using TinyML. It identifies 90 research articles published between 2018 and 2025, categorizing them into two main areas: (1) model conversion, inference, and learning strategies in EdgeML and (2) deploying TinyML models on resource-constrained hardware using specific software frameworks. In the first category, the synthesis of selected research articles compares and critically reviews various model conversion techniques, inference mechanisms, and learning strategies. In the second category, the synthesis identifies and elaborates on major development boards, software frameworks, sensors, and algorithms used in various applications across six major sectors. As a result, this article provides valuable insights for researchers, practitioners, and developers. It assists them in choosing suitable model conversion techniques, inference mechanisms, learning strategies, hardware development boards, software frameworks, sensors, and algorithms tailored to their specific needs and applications across various sectors.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Edge machine learning</kwd>
<kwd>tiny machine learning</kwd>
<kwd>model compression</kwd>
<kwd>inference</kwd>
<kwd>learning algorithms</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Deep Learning (DL) has emerged as a new Machine Learning (ML) paradigm, capable of automatically learning complex data representations at multiple levels of abstraction [<xref ref-type="bibr" rid="ref-1">1</xref>]. At the same time, the processing and communication capabilities of embedded devices have tremendously increased [<xref ref-type="bibr" rid="ref-2">2</xref>]. As a result, the term Internet of Things (IoT) has emerged that refers to a network of interconnected embedded devices [<xref ref-type="bibr" rid="ref-3">3</xref>]. However, IoT devices generate vast amounts of data and have limited resources. Therefore, cloud computing is integrated into IoT frameworks to provide the required computational and storage capacity. While the cloud offers essential processing power, cloud-stored data may not always be secure [<xref ref-type="bibr" rid="ref-4">4</xref>]. Consequently, IoT frameworks utilize edge computing, which relies on devices that can perceive their environment and process data locally. In this context, Edge Machine Learning (EdgeML) extends edge computing by directly integrating ML and DL capabilities into edge devices. This recent ML paradigm shifts all or part of the machine learning computation from the cloud to edge devices [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates a typical EdgeML architecture with three layers. The edge layer (front end) is equipped with sensors and signal-processing algorithms. This layer is responsible for data collection and processing. As a result, inference and training of local models take place. It connects to the edge computing layer (near end) via wireless communication. The edge computing layer acts as an intermediary, linking edge devices to the cloud. It supports inference and learning of new data in collaboration with edge devices. It may involve splitting the model across multiple devices for collaborative (distributed) inference or running the entire model on a single device. If an edge device lacks sufficient resources, it can collaborate with other edge devices or the edge server [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. Its key advantages include enhanced data privacy and security, as data is not transmitted to a centralized server. Additionally, the failure of a single device has minimal impact on the learning process, and scalability is easily managed by engaging multiple devices as needed. Strategies for distributed learning on edge devices include federated learning, model splitting/partitioning, and hierarchical clustering learning [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>A typical three-layer edge machine learning architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_62819-fig-1.tif"/>
</fig>
<p>While the edges layer and edge computing layer perform local and distributive model training, a cloud computing layer (far end) contains extensive computational resources for global model training and running high-demand algorithms using pre-processed data. Moreover, pre-trained models can be fine-tuned for specific tasks but often require model conversion to reduce complexity for edge device deployment [<xref ref-type="bibr" rid="ref-8">8</xref>]. Consequently, there are various distributed model design and deployment strategies. These distributed or collaborative strategies must be compared in terms of latency, privacy, security, reliability, energy consumption, and computational capabilities.</p>
<p>From the above discussion, it can be concluded that the EdgeML framework involves three main steps. The first step is model conversion or compression, which creates smaller models or converts complex models into simpler ones through pruning, quantization, and knowledge distillation. The model&#x2019;s complexity is tailored to the edge device&#x2019;s resources. The second step is optimized inference that minimizes resource usage on edge devices. Finally, protocols such as collaborative learning across multiple devices should be implemented to update the model if further learning is needed.</p>
<p>In addition to three major steps in EdgeML, another trend is performing complex processing tasks entirely on edge devices [<xref ref-type="bibr" rid="ref-2">2</xref>]. TinyML (Tiny Machine Learning) enables ML and DL algorithms to run on tiny devices like microcontrollers [<xref ref-type="bibr" rid="ref-9">9</xref>]. TinyML architecture focuses on designing memory-efficient models for edge devices, ensuring high performance. It has led to the concept of the Internet of Intelligent Things (IoIT), which combines embedded hardware, wireless networking, and artificial intelligence. TinyML frameworks rely on specialized sensors, software frameworks, and tools for developing, training, and deploying models on resource-constrained devices [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<sec id="s1_1">
<label>1.1</label>
<title>Motivation for the Review and Limitations of Existing Reviews</title>
<p>Given the variety of model conversion techniques in EdgeML, and the importance of scalable and efficient inference mechanisms and learning strategies for large-scale deployments, conducting a literature review on these topics can enhance understanding of performance improvements in speed, latency, and energy efficiency. Additionally, it can offer a comprehensive overview of TinyML applications across different domains, exploring the latest advancements in hardware, software, and sensor technologies. The limitations of state-of-the-art review articles have been highlighted in <xref ref-type="table" rid="table-1">Table 1</xref>. These limitations reveal that there is no comprehensive literature review that covers model conversion, inference, and learning in EdgeML and TinyML deployment at the same time. Therefore, by synthesizing existing research, the literature review can identify knowledge gaps and suggest future research directions in the most promising areas of EdgeML and TinyML.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of state-of-the-art review articles on EdgeML and TinyML</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ref.</th>
<th>Year</th>
<th>Focus</th>
<th>Limitations</th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="4">Edge ML</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-5">5</xref>]</td>
<td>2023</td>
<td><list list-type="bullet">
<list-item>
<p>Examines various edge computing paradigms from various perspectives</p></list-item>
<list-item>
<p>Provides an overview of EdgeML for limited computing resources</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Does not review state-of-the-art techniques on model conversion, inference mechanisms, and learning strategies</p></list-item>
<list-item>
<p>Does not discuss hardware and software frameworks for the deployment of ML and DL on edge devices</p></list-item>
</list>
</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Emphasizes the evolution of EdgeML</p></list-item>
<list-item>
<p>Pinpoints research opportunities in EdgeML to unify ML and edge computing</p></list-item>
</list></td>
<td/>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>2023</td>
<td><list list-type="bullet">
<list-item>
<p>Addresses key issues in EdgeML and summarizes optimization techniques</p></list-item>
</list></td>
<td/>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>2023</td>
<td><list list-type="bullet">
<list-item>
<p>Highlights the potential of edge computing</p></list-item>
<list-item>
<p>Describes lightweight ML frameworks</p></list-item>
<list-item>
<p>Explores privacy issues</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Does not review techniques to create lightweight ML models and learning mechanisms</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Compression techniques of DNN</p></list-item>
<list-item>
<p>Explore the applicability of compressed models in visual applications</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Does not cover the learning and inference aspects of EdgeML on resource-limited architectures</p></list-item>
</list></td>
</tr>
<tr>
<td colspan="4"><bold>TinyML</bold></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Classifies model optimization techniques</p></list-item>
<list-item>
<p>Hardware and software frameworks Discuss TinyML educational resources</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Does not classify hardware and software frameworks as well as model conversion, inference, and learning strategies</p></list-item>
</list>
</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Ahistorical evolution of IoIT and reviews TinyML models on embedded devices</p></list-item>
</list></td>
<td/>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Edge computing platforms and their role in the deployment of EdgeML workflows</p></list-item>
<list-item>
<p>Reviews frameworks and libraries for ML and models on constrained edge devices</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Only limited to the deployment of EdgeML workflows and does not discuss the deployment of TinyML workflow</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>2024</td>
<td><list list-type="bullet">
<list-item>
<p>Systematically reviews TinyML models on embedded devices with hardware and software frameworks</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Only limited to TinyML deployment and does not review model conversion, inference, and learning strategies</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Contributions</title>
<p>The limitations of state-of-the-art review articles, highlighted in <xref ref-type="table" rid="table-1">Table 1</xref>, have been rectified by performing this literature review. Particularly, it has explored the answers to the following five research questions:</p>

<p><bold>Research Question 1:</bold> What are the state-of-the-art model conversion techniques used in EdgeML, and how do they impact the performance, efficiency, and deployment of machine learning models on resource-constrained devices?</p>
<p><bold>Research Question 2:</bold> What are the current state-of-the-art inference mechanisms used in EdgeML, and how do they compare in performance and efficiency?</p>
<p><bold>Research Question 3</bold>: How do different learning strategies impact the performance and efficiency of ML models deployed in edge computing environments?</p>
<p><bold>Research Question 4:</bold> What are the key challenges and problems and the associated ML/DL models in deploying TinyML on resource-constrained devices in various sectors?</p>
<p><bold>Research Question 5:</bold> What are the latest advancements in hardware, software frameworks, and sensors designed explicitly for TinyML frameworks?</p>
<p>It can be observed from the above research questions that the contribution of this literature review is twofold: (1) a critical review of model conversion techniques, inference mechanisms, and learning strategies in EdgeML, (2) identification of hardware development boards, software frameworks, sensors, ML/DL models and applications across various sectors for the deployment of TinyML models on resource-constrained devices.</p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> provides an overview of the literature review on model conversion, inference, and learning in EdgeML with TinyML deployment in different sectors. The process starts by developing a review protocol, selecting 90 research articles from renowned databases, categorized into two main areas: (1) model conversion, inference, and learning in EdgeML and (2) model deployment in six different sectors on resource-constrained devices with TinyML. The first category includes 60 articles, divided into model conversion (40 articles), inference (10 articles), and learning (10 articles). The second category comprises 30 articles, divided into six sectors with five articles each.</p>
<p>It can be seen in <xref ref-type="table" rid="table-2">Table 2</xref> that most research in EdgeML centers on model conversion, adapting resource-intensive models for deployment on resource-limited devices. This focus stems from the widespread use of powerful deep-learning models in various real-world applications. However, research into learning and inference mechanisms on edge devices remains challenging and needs further exploration.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Overview of the literature review on model conversion, inference, and learning in EdgeML with TinyML deployment in different sectors</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td colspan="10"><bold>Selection of 90 research articles according to a review protocol</bold></td>
</tr>
<tr>
<td colspan="4">Model conversion, learning, and inference in EdgeML (60 articles)</td>
<td colspan="6">Model deployment with TinyML in (30 articles)</td>
</tr>
<tr>
<td>IEEE (18)</td>
<td>Elsevier (34)</td>
<td>ACM (4)</td>
<td>Springer (4)</td>
<td>IEEE (15)</td>
<td colspan="2">Elsevier (6)</td>
<td colspan="2">ACM (6)</td>
<td>Springer (3)</td>
</tr>
<tr>
<td colspan="10"><bold>Classification of selected research works into subcategories (total &#x003D; 90)</bold></td>
</tr>
<tr>
<td>Model conversion (40)</td>
<td colspan="2">Inference mechanisms (10)</td>
<td>Learning strategies (10)</td>
<td>Smart Agri.(5)</td>
<td>Health &#x0026; En.(5)</td>
<td>Vehic.&#x0026; Aut.(5)</td>
<td>Indus.&#x0026; Rob.(5)</td>
<td>Ener.(5)</td>
<td>Secu.(5)</td>
</tr>
<tr>
<td colspan="10"><bold>Synthesis of qualitative and quantitative analysis</bold></td>
</tr>
<tr>
<td><list list-type="bullet">
<list-item>
<p>Pruning</p></list-item>
<list-item>
<p>Quantization</p></list-item>
<list-item>
<p>Factorization</p></list-item>
<list-item>
<p>Knowledge Distillation</p></list-item>
</list></td>
<td colspan="2">
<list list-type="bullet">
<list-item>
<p>On-device</p></list-item>
<list-item>
<p>Distributed</p></list-item>
</list>
</td>
<td><list list-type="bullet">
<list-item>
<p>On-device</p></list-item>
<list-item>
<p>Continual</p></list-item>
<list-item>
<p>Federated</p></list-item>
</list></td>
<td colspan="6">
<list list-type="bullet">
<list-item>
<p>Identification and comparison of Hardware Development Boards</p></list-item>
<list-item>
<p>Identification of Software Frameworks</p></list-item>
<list-item>
<p>Identification of 30 applications (examples) in 6 sectors with models and sensors</p></list-item>
</list>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The first category&#x2019;s comprehensive analysis covers four model compression techniques (pruning, quantization, low-rank factorization, and knowledge distillation), inference mechanisms (on-device inference and split-model inference), and learning strategies (on-device learning, continual learning, and federated learning). The second category&#x2019;s synthesis provides an in-depth discussion on hardware development boards and software frameworks for deploying TinyML models, identifying 30 applications with corresponding ML/DL models and sensors across six major sectors.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Organization</title>
<p>This article is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> defines categories and reviews protocol development. The summary of major findings on model compression techniques, inference mechanisms, and learning strategies is presented in <xref ref-type="sec" rid="s3">Section 3</xref>, while the results of TinyML deployment on resource-constrained devices are presented in <xref ref-type="sec" rid="s4">Section 4</xref>. Discussion of results from <xref ref-type="sec" rid="s3">Sections 3</xref> and <xref ref-type="sec" rid="s4">4</xref>, addressing research questions, is presented in <xref ref-type="sec" rid="s5">Section 5</xref>. A discussion of important aspects and limitations of the research is presented in <xref ref-type="sec" rid="s6">Section 6</xref>. Exploration of challenges and future research directions are presented in <xref ref-type="sec" rid="s7">Section 7</xref>. Finally, <xref ref-type="sec" rid="s8">Section 8</xref> concludes the article.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Methodology of the Literature Review</title>
<p>In order to obtain the answers to research questions formulated in the introductory part of this article, literature review process guidelines provided in [<xref ref-type="bibr" rid="ref-17">17</xref>] have been used. A typical literature review procesds defines categories and develops a review protocol for selecting research articles. Therefore, <xref ref-type="sec" rid="s2_1">Section 2.1</xref> describes the essential background of different categories. Subsequently, various steps of the review protocol are described in <xref ref-type="sec" rid="s2_2">Section 2.2</xref>.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Background on Categories</title>
<p>The research articles, selected according to the review protocol, are categorized into two types: (1) model conversion, inference, and learning (2) model deployment on resource-constrained devices.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Model Conversion, Inference, and Learning</title>
<p>As the Introduction section mentions, edge devices have computational power, memory, and energy resource constraints. Therefore, it is essential to design ML and DL models considering these limitations. There are two main approaches to achieve this: (a) develop a model tailored to the edge device&#x2019;s limitations and train it on a large dataset to ensure satisfactory performance across various environments, and (b) adopt a large pre-trained model to make it suitable for edge devices. Moreover, due to limited resources, model inference requires device-specific or problem-specific strategies. Inference can be adapted based on available resources and task requirements. Furthermore, retraining models with new data is essential in some scenarios, optimizing performance over time. Therefore, the following subsections provide essential background on model conversion, inference, and continuous learning.</p>
<p>The main goal of model conversion is to reduce model size without compromising performance. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows various stages of model conversion, which can be applied individually or in combination. After training the original model on sensor data, parameters can be quantized and pruned. Other techniques to reduce model size include knowledge distillation and low-rank factorization. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the training data is collected from various sensors based on task requirements. A conventional model is developed using knowledge of the problem, data, and performance needs. DL models, popular for tasks like image processing and object detection, have numerous parameters and require significant computational and memory resources. These models are typically trained on cloud platforms or high-performance machines and then converted into edge device-friendly models using pruning. Pruning removes insignificant or redundant parameters, creating sparser models without significantly reducing performance.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Block diagram of model conversion and compression</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_62819-fig-2.tif"/>
</fig>
<p>While pruning removes less important connections in a model without compromising performance, quantization maps higher bit-width values to lower bit-width values, with accuracy depending on optimal quantization levels. Parameters like weights, biases, and activation functions in deep neural networks can be quantized. On edge devices, gradient and error values may also be quantized. Quantization can be applied to pre-trained models to assess accuracy or performance compromise. Fixed-precision quantization maps 32-bit floating points to k-bit integers, where k is less than 32 bits. For k-bit quantization, the procedure is defined as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mi>r</mml:mi><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>Z</mml:mi></mml:math></disp-formula>where <italic>Z</italic> is an integer zero point value, <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> is the step size, and <italic>r</italic> is the 32-bit floating point value ranging from &#x03B1; to &#x03B2;. Two simple methods of k-bit quantization are MaxRange method and MinPQE [<xref ref-type="bibr" rid="ref-6">6</xref>]. The MaxRange method selects the step size based on the real value range as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>In contrast, the MinPQE method optimizes step sizes for each neural network layer based on quantization error, improving performance and reducing error through quantization granularity. For instance, 8-bit quantization on the NasNet&#x2013;A model increases classification error from 4% to 5%. Extreme quantization uses less than 4-bit quantization. The step size, or scaling factor, is crucial for minimizing quantization loss, with much research focused on its optimization. Post-quantization training can enhance accuracy, and pre-training quantization benefits edge devices with limited resources. Low-rank decomposition can reduce the number of parameters in deep neural networks by approximating the weight matrix to a low-rank structure. However, it may cause significant performance degradation, making it less suitable for model compression.</p>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> explains two types of quantization strategies in model compression. In the quantization-aware training of the model, a quantization policy is defined as one that quantizes the weight, activation function, filters, etc. The model&#x2019;s learning depends on its performance based on the model parameters and the current quantization policy. Hence, the model parameters and the quantization mechanism are updated simultaneously or in batch mode. Meanwhile, in post-training quantization, a model is trained on the dataset first, and then the model parameters are quantized.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Block diagram of quantization process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_62819-fig-3.tif"/>
</fig>
<p>A complex teacher model with many parameters is trained on a dataset in knowledge distillation. Once trained and tested, its knowledge is transferred to a smaller, simpler student model. Student models are lightweight, can be deployed on resource-constrained devices, and can be customized according to hardware requirements. This process creates a simpler model with a smaller memory footprint and lower computational requirements. There are three main categories for transferring knowledge: response-based, features-based, and relation-based knowledge distillation [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Inference Strategies</p>
<p>Once deployed on an edge device, a machine learning model infers decisions based on its task. Due to limited resources, different inference techniques are used for efficiency [<xref ref-type="bibr" rid="ref-19">19</xref>]. Lightweight models like TinyML or SqueezeNet can perform direct inference on the edge device, offering low latency, no communication requirement, high privacy, and data security. For heavier computation, the model can be split into parts to run on multiple edge devices or servers or in collaboration with the cloud. It preserves data privacy by preprocessing on the edge device and sharing features with the server. Partitioning-based inference methods include data-based partitioning, where input data is divided among devices, and model partitioning, where the model is split among devices or servers.</p>
<p>Learning Strategies</p>
<p>Deep neural network-based solutions on edge or IoT devices are crucial for automation, offering high latency, privacy, and reduced communication bandwidth. High latency enables faster decision-making, which is essential for many applications. Edge machine learning often uses computational offloading to manage limited resources, transferring tasks to robust edge servers or cloud platforms. Task offloading decisions depend on the network connection, latency requirements, and DL model structure. Various strategies exist for offloading tasks. Federated learning (FL) enhances data privacy on edge devices by training models collaboratively with other devices or servers [<xref ref-type="bibr" rid="ref-20">20</xref>]. Data remains on edge devices, and only model parameters are shared, reducing network bandwidth utilization and latency while improving model generalization on heterogeneous data. Each device trains on a subset of data using stochastic gradient descent method variants. The main server initializes tasks, assigns them to edge devices, and aggregates optimized model parameters to create a global model. This process iterates to minimize global loss, with the server updating and distributing the global model parameters.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Model Deployment on Resource-Constrained Embedded Devices</title>
<p>TinyML enables the deployment of ML and DL models on embedded devices like microcontrollers and single-board computers. This technology makes low-power edge devices &#x201C;smart,&#x201D; allowing them to perform complex tasks independently. Implementing ML and DL models on these devices is more challenging than the conventional EdgeML approach, where models are hosted in the cloud and run on powerful computers.</p>
<p>Hardware Requirements</p>
<p>It is challenging for a hardware platform to simultaneously meet energy, cost, and processing efficiency. Some platforms may have enough processing power but lack energy efficiency or cost-effectiveness, and <italic>vice versa</italic>. Choosing the right device for a specific application is crucial in the TinyML paradigm. Devices with higher computational power that still consider energy efficiency and cost are called High-end TinyML architectures. Similarly, devices with lower computational power, suitable for small tasks only, are called Low-end TinyML architectures [<xref ref-type="bibr" rid="ref-2">2</xref>]. Examples of high-end TinyML architectures are single-board computers (SBCs), which run entire operating systems and TinyML models in real-time. They are used as sensor nodes and gateways in IoT. On the other hand, low-end TinyML architectures target energy efficiency and cost-effectiveness, making them suitable for more straightforward tasks. This class includes microcontroller units (MCUs) with low-power processors, limited RAM, and simple interfaces.</p>
<p>Software Frameworks</p>
<p>Software frameworks for model deployment offer pre-built tools and libraries, allowing developers to focus on model optimization rather than low-level hardware details. These frameworks optimize ML and DL models for resource-constrained devices, reducing model size and improving inference speed. They enable models to run on different hardware with minimal changes, making it easier to scale TinyML applications across various devices. Commonly used frameworks include Edge Impulse, TensorFlow, and TensorFlow Lite.</p>
<p>Model Types</p>
<p>Different machine learning models are tailored for specific tasks based on the application problem. For example, CNNs (Convolutional Neural Networks) identify objects, faces, and scenes in images. Similarly, time-series models like RNNs (Recurrent Neural Networks) and LSTMs (Long Short-Term Memory Networks) handle sequential data and capture temporal dependencies. Clustering models like K-Means and DBSCAN group similar data points without predefined labels used in anomaly detection.</p>
<p>Applications</p>
<p>TinyML models are ideal for battery-operated devices and real-time decision-making applications. They require small, inexpensive hardware, reducing deployment costs. Their adaptability across different devices and platforms makes them suitable for scalable applications. Local processing enhances privacy and security by keeping data on the device. Consequently, TinyML models can be deployed in various fields, such as healthcare, agriculture, smart homes, and industrial automation.</p>
<p>Sensors</p>
<p>Different sensors are used in TinyML applications to capture various data types for specific tasks&#x2014;for example, cameras for image classification and accelerometers for activity recognition. The right sensor ensures accurate and relevant data, improving model performance and reliability. Different environments require different sensors, like soil moisture sensors for agriculture and temperature sensors for smart homes. Power-efficient sensors are suitable for battery-operated devices, maintaining overall power efficiency. Ensuring sensor compatibility with hardware and software platforms ensures seamless integration and better performance of the TinyML system.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Review Protocol Development</title>
<p>The development of a review protocol is essential in a typical literature review process. The developed protocol in this section contains all the required steps. These steps are: criteria for selecting and rejecting studies (<xref ref-type="sec" rid="s2_2_1">Section 2.2.1</xref>), search process (<xref ref-type="sec" rid="s2_2_2">Section 2.2.2</xref>), and data extraction &#x0026; synthesis (<xref ref-type="sec" rid="s2_2_3">Section 2.2.3</xref>).</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Selection and Rejection Criteria</title>
<p><list list-type="roman-lower">
<list-item>
<p><bold>Subject-Relevant:</bold> Select research only if it is relevant to our research context and supports the answers to our research questions.</p></list-item>
<list-item>
<p><bold>Publication Date (2018&#x2013;2025):</bold> Select research published between 2020 and 2025 to include the latest studies. Reject any research published before 2018.</p></list-item>
<list-item>
<p><bold>Publisher:</bold> Select research published in one of the four renowned scientific databases: IEEE, Springer, Elsevier, and ACM.</p></list-item>
<list-item>
<p><bold>Crucial Effects</bold>: Select research that significantly affects IIoT development through the EdgeML and TinyML approach.</p></list-item>
<list-item>
<p><bold>Results-Oriented</bold>: Select results-oriented research with proposals and outcomes supported by solid facts and experimentation.</p></list-item>
<list-item>
<p><bold>Repetition:</bold> Avoid including identical research within the same context.</p></list-item>
</list></p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Search Process</title>
<p>The selection and rejection criteria outlined in <xref ref-type="sec" rid="s2_2_1">Section 2.2.1</xref> indicate that we have chosen four scientific databases (IEEE, ELSEVIER, SPRINGER, and ACM). These databases include high-impact journals and conference proceedings. Using search terms like Edge AI, TinyML, and EdgeML, we applied a &#x201C;2018&#x2013;2025&#x201D; filter. The search terms and results for each database are summarized in <xref ref-type="table" rid="table-3">Table 3</xref>. If a search term has produced thousands of results, we used advanced search options (e.g., &#x201C;where abstract contained&#x201D;, &#x201C;where title contained&#x201D;) provided by these databases to get more precise results. Finally, <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows various steps during the selection of research articles. We specified search terms in four scientific databases and analyzed approximately 43,651 results based on our selection and rejection criteria. We discarded 26,991 studies by title, 9173 by abstract, and 6495 after a general review. Finally, we performed a detailed review of the remaining 992 articles and selected 90 research articles.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Details of search terms and search results in selected databases</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>S. no.</th>
<th>Search term</th>
<th colspan="4">Number of search results</th>
</tr>
<tr>
<th/>
<th/>
<th>IEEE</th>
<th>Elsevier</th>
<th>Springer</th>
<th>ACM</th>
</tr>
</thead>
<tbody>
<tr>
<td>1.</td>
<td>Edge AI</td>
<td>4077</td>
<td>1012</td>
<td>51</td>
<td>586</td>
</tr>
<tr>
<td>2.</td>
<td>EdgeML</td>
<td>1741</td>
<td>857</td>
<td>2</td>
<td>373</td>
</tr>
<tr>
<td>3.</td>
<td>TinyML</td>
<td>497</td>
<td>59</td>
<td>31</td>
<td>254</td>
</tr>
<tr>
<td>4.</td>
<td>TinyML hardware</td>
<td>161</td>
<td>13</td>
<td>27</td>
<td>216</td>
</tr>
<tr>
<td>5.</td>
<td>TinyML applications</td>
<td>273</td>
<td>36</td>
<td>31</td>
<td>250</td>
</tr>
<tr>
<td>6.</td>
<td>TinyML software</td>
<td>79</td>
<td>8</td>
<td>24</td>
<td>197</td>
</tr>
<tr>
<td>7.</td>
<td>Edge machine learning</td>
<td>5899</td>
<td>2155</td>
<td>717</td>
<td>1189</td>
</tr>
<tr>
<td>8.</td>
<td>Internet of intelligent things</td>
<td>7847</td>
<td>1794</td>
<td>42</td>
<td>565</td>
</tr>
<tr>
<td>9.</td>
<td>TinyML model</td>
<td>416</td>
<td>44</td>
<td>32</td>
<td>4</td>
</tr>
<tr>
<td>10.</td>
<td>TinyML microcontroller</td>
<td>212</td>
<td>23</td>
<td>100</td>
<td>18</td>
</tr>
<tr>
<td>11.</td>
<td>Edge quantization</td>
<td>820</td>
<td>227</td>
<td>549</td>
<td>154</td>
</tr>
<tr>
<td>12.</td>
<td>Edge pruning</td>
<td>750</td>
<td>176</td>
<td>858</td>
<td>320</td>
</tr>
<tr>
<td>13.</td>
<td>Edge low-rank factorization</td>
<td>17</td>
<td>4</td>
<td>223</td>
<td>74</td>
</tr>
<tr>
<td>14.</td>
<td>Edge knowledge distillation</td>
<td>348</td>
<td>91</td>
<td>421</td>
<td>66</td>
</tr>
<tr>
<td>15.</td>
<td>Edge federated learning</td>
<td>2650</td>
<td>412</td>
<td>504</td>
<td>295</td>
</tr>
<tr>
<td>16.</td>
<td>Edge continual learning</td>
<td>94</td>
<td>21</td>
<td>215</td>
<td>35</td>
</tr>
<tr>
<td>17.</td>
<td>Edge model splitting</td>
<td>521</td>
<td>212</td>
<td>1537</td>
<td>145</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Search process (various steps during the selection of research articles)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_62819-fig-4.tif"/>
</fig>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Data Extraction and Synthesis</title>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, data extraction and synthesis are conducted for selected studies to address our research questions. For data extraction (serial numbers 2 to 6), we gather essential details from each study to ensure they meet the selection and rejection criteria. We perform a detailed analysis for data synthesis (serial numbers 7 to 10), thoroughly examining and categorizing each study. Each study is meticulously reviewed to fit the corresponding category. Statistics of research articles according to the publication year are provided in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Details of data extraction and synthesis</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>S. no.</th>
<th>Description</th>
<th>Details</th>
</tr>
</thead>
<tbody>
<tr>
<td>1.</td>
<td>Bibliographic information</td>
<td>Title, author, publication year, publisher details, and type of research (i.e., journal or conference)</td>
</tr>
<tr>
<td align="center" colspan="3">Data extraction</td>
</tr>
<tr>
<td>2.</td>
<td>Overview</td>
<td>The basic proposal and objective</td>
</tr>
<tr>
<td>3.</td>
<td>Results</td>
<td>Results acquired from the selected research</td>
</tr>
<tr>
<td>4.</td>
<td>Data collection</td>
<td>Quantitative or qualitative</td>
</tr>
<tr>
<td>5.</td>
<td>Assumption</td>
<td>Assumptions (if any) to validate the results</td>
</tr>
<tr>
<td>6.</td>
<td>Validation</td>
<td>Validation method used to validate its proposal</td>
</tr>
<tr>
<td align="center" colspan="3">Data synthesis</td>
</tr>
<tr>
<td>7.</td>
<td>Model conversion techniques</td>
<td>Table 5: Results of structured pruning methods<break/>Table 6: Results of quantization methods<break/>Table 7: Results on quantization-aware training<break/>Table 8: Results of low-rank factorization<break/>Table 9: Results of knowledge distillation</td>
</tr>
<tr>
<td>8.</td>
<td>Inference mechanisms</td>
<td>Table 10: Results of inference strategies</td>
</tr>
<tr>
<td>9.</td>
<td>Learning strategies</td>
<td>Table 11: Results of continual learning methods<break/>Table 12: Results of federated learning methods</td>
</tr>
<tr>
<td>10.</td>
<td>TinyML sectors, applications, and models</td>
<td>Table 13: Applications and machine learning models in six identified sectors</td>
</tr>
<tr>
<td>11.</td>
<td>Hardware boards</td>
<td>Table 14: Summary of development boards</td>
</tr>
<tr>
<td>12.</td>
<td>Software frameworks</td>
<td>Table 15: Distribution of selected research works according to three software frameworks</td>
</tr>
<tr>
<td>13.</td>
<td>Sensors</td>
<td>Table 16: Sensors used for the deployment</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Statistics of research articles according to the publication year</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_62819-fig-5.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Results on Model Conversion, Inference and Learning</title>
<p>Based on the methodology of <xref ref-type="sec" rid="s2">Section 2</xref>, the selected research studies have been classified into two major categories: (1) model conversion, inference, and learning in EdgeML (2) model deployment on resource-constrained devices with TinyML. Consequently, the results on model conversion (<xref ref-type="sec" rid="s3_1">Section 3.1</xref>), efficient inference (<xref ref-type="sec" rid="s3_2">Section 3.2</xref>), and continuous training strategies (<xref ref-type="sec" rid="s3_3">Section 3.3</xref>) are presented in this section, while the results of model deployment will be presented in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Model Conversion Methods</title>
<p>Deploying large models on resource-constrained edge devices is impractical. Two options are available: (1) Customize a small DL model with fewer layers and full precision parameters to fit within the edge device&#x2019;s memory and require less computational power, (2) Convert and compress an existing model by removing insignificant parameters and using quantized parameters with lower precision or bit-width to achieve similar performance. The learning process can incorporate quantized gradients if training occurs on edge devices. The following subsections will detail various techniques for model conversion and compression.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Pruning Methods</title>
<p>Pruning methods are divided into structured and unstructured methods. Unstructured pruning creates irregular, sparse models by pruning any weight without constraints, followed by retraining to mitigate performance loss. However, this method requires specific hardware and software redesigns for efficiency. Therefore, Li et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a flexible rate filter pruning method (structured) using a loss-aware process. Structured pruning removes redundant filters and channels to make the models lighter in size. Filter norms can help identify insignificant filters. For example, low norm filters exhibit low activation in feature maps and can be pruned. The sensitivity of each filter to model performance can be measured to prune unimportant filters [<xref ref-type="bibr" rid="ref-22">22</xref>]. Structured pruning is straightforward for DL models with simple convolutional layers. However, the pruning strategy must consider the consistency of feature maps and residual flows for residual block-based networks. Due to the complexity of DL models in pattern recognition, most research focuses on structured pruning methodologies, as shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Results of structured pruning methods on CNN architectures</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col/>
<col align="center"/>
<col/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ref.</th>
<th align="center">Dataset/<break/>Model</th>
<th>Pruning method</th>
<th align="center">Baseline accuracy</th>
<th>Accuracy</th>
<th align="center">Pruned parameters</th>
<th align="center">Model size reduction</th>
<th align="center">Flops saved</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>CIFAR-10<break/>ResNet-110</td>
<td>Flexible pruning</td>
<td>94.2%</td>
<td>94.2%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>64%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>CIFAR-10<break/>VGG-16</td>
<td>Filters</td>
<td>93.6%</td>
<td>93.3%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>52%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>CIFAR-110<break/>VGG-16</td>
<td>Filters</td>
<td>73.4%</td>
<td>73.6%</td>
<td>6.4 M</td>
<td>56%</td>
<td>44%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>CIFAR-110<break/>ResNet-164</td>
<td>Filters</td>
<td>76.8%</td>
<td>76.8%</td>
<td>0.8 M</td>
<td>30%</td>
<td>51%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>CIFAR-10<break/>VGG-16</td>
<td>Hybrid</td>
<td>93.7%</td>
<td>93.5%</td>
<td>0.76 M</td>
<td>94%</td>
<td>73%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>CIFAR-10<break/>ResNet-56</td>
<td>Hybrid</td>
<td>93.9%</td>
<td>90%</td>
<td>0.08 M</td>
<td>90%</td>
<td>87%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>CIFAR-100<break/>VGG-16</td>
<td>Filters</td>
<td>73.9%</td>
<td>73.6%</td>
<td>&#x2013;</td>
<td>50%</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>CIFAR-100<break/>ResNet50</td>
<td>Filters</td>
<td>75.9%</td>
<td>74%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>51%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>CIFAR-10<break/>VGG-16</td>
<td>Filters</td>
<td>93.9%</td>
<td>93.4%</td>
<td>&#x2013;</td>
<td>93%</td>
<td>81%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>CIFAR-10<break/>ResNet-110</td>
<td>Filters</td>
<td>93.2%</td>
<td>93.2%</td>
<td></td>
<td>62%</td>
<td>62%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>CIFAR-100<break/>VGG-16</td>
<td>Filters</td>
<td>73.6%</td>
<td>73.4%</td>
<td>0.85 M</td>
<td>94%</td>
<td>84%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>CIFAR-100<break/>ResNet-50</td>
<td>Filters</td>
<td>71%</td>
<td>70.9%</td>
<td>18 M</td>
<td>76%</td>
<td>76%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>CIFAR-10<break/>VGG-16</td>
<td>Filters</td>
<td>93.9%</td>
<td>92.8%</td>
<td>1.15 M</td>
<td>92%</td>
<td>84%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>CIFAR-10<break/>ResNet-110</td>
<td>Filters</td>
<td>93.5%</td>
<td>93.4%</td>
<td>0.52 M</td>
<td>69.9%</td>
<td>73%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>CIFAR-10<break/>VGG-16</td>
<td>Weights</td>
<td>93.6%</td>
<td>93.2%</td>
<td>&#x2013;</td>
<td>77%</td>
<td>64%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>CIFAR-10<break/>ResNet-110</td>
<td>Weights</td>
<td>92.5%</td>
<td>92.3%</td>
<td>&#x2013;</td>
<td>46%</td>
<td>49%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In CNNs, filters in convolutional layers are pruned to improve efficiency. Various methods have been proposed, such as Capped L1-norm balances regularization and filter selection [<xref ref-type="bibr" rid="ref-23">23</xref>]. Yu et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] treat pruning as a nonconvex-constrained optimization problem, using information from a Taylor-based approximation to quantify filter contributions and prune the network. Another method, proposed by He et al. [<xref ref-type="bibr" rid="ref-22">22</xref>], prunes filters layer by layer using a standard loss function like cross-entropy, followed by fine-tuning to improve performance. This process is repeated for each layer until the entire network is pruned. The goal of these methods [<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>] is to optimize the model&#x2019;s performance while reducing complexity.</p>
<p>In addition to the works in [<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>], some filter pruning methods are based on the statistics of the feature map generated by the entire training dataset. In this context, Mondal et al.&#x2019;s method [<xref ref-type="bibr" rid="ref-25">25</xref>] evaluates filter importance for each class separately using the normalized L1-norm of the feature map. Filters generating dominant patterns for a class are not pruned. Similarly, Sarvani et al.&#x2019;s knowledge transfer method [<xref ref-type="bibr" rid="ref-26">26</xref>] uses a customized regularizer function to transfer knowledge from pruned filters to important ones, minimizing information loss by adjusting L1-norm values. The mutual information theory method, proposed by Lu et al. [<xref ref-type="bibr" rid="ref-27">27</xref>], prunes filters in two steps. The first step calculates the class relevance of each filter using conditional mutual information on mini-batches. The second step averages class synthesis evaluation criteria (relevance and redundancy) across mini-batches and prunes filters with lesser contributions.</p>
<p>While methods in [<xref ref-type="bibr" rid="ref-25">25</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>] aim to optimize model performance by selectively pruning filters based on their importance and relevance, Wavelet Transform-Based pruning uses cosine similarity and energy-weighted components of high and low frequencies to determine the importance score of each feature map [<xref ref-type="bibr" rid="ref-28">28</xref>]. A multi-objective evolutionary framework by Chung et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] balances performance and efficiency for edge devices using a fitness function based on the pruned model&#x2019;s flops ratio and error rate. It employs three pruning criteria: L1-norm-based, percentage of zero activations, and gradient-based. The converged pruned architectures provide a Pareto front for solutions balancing accuracy and inference time. These methods [<xref ref-type="bibr" rid="ref-28">28</xref>] and [<xref ref-type="bibr" rid="ref-29">29</xref>] aim to optimize model performance while considering hardware constraints and efficiency.</p>
<p>To summarize, <xref ref-type="table" rid="table-5">Table 5</xref> highlights the performance of different pruning methods and their performance on various architectures and problems. The tables show that pruning the weights or filters can reduce the computational cost of the model without sacrificing accuracy. For a higher reduction of Flops (78%), a drop of around 3% accuracy is observed. Structured pruning (<xref ref-type="table" rid="table-5">Table 5</xref>) is based on the selection or pruning of the filters present in deep learning architectures. Such a pruning method deals with the weight matrix sparsity in a better and more optimized way when considering hardware implementation.</p>

</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Quantization Methods</title>
<p>State-of-the-art model conversion methods using quantization are shown in <xref ref-type="table" rid="table-6">Table 6</xref>. Various quantization methods used in the research works of <xref ref-type="table" rid="table-6">Table 6</xref> are mixed precision quantization [<xref ref-type="bibr" rid="ref-30">30</xref>], vector quantization [<xref ref-type="bibr" rid="ref-31">31</xref>], quantization-loss aware algorithm [<xref ref-type="bibr" rid="ref-32">32</xref>], smart-DNN&#x002B; [<xref ref-type="bibr" rid="ref-33">33</xref>], multi-branch topology [<xref ref-type="bibr" rid="ref-34">34</xref>], ultra-low bit quantization [<xref ref-type="bibr" rid="ref-35">35</xref>], Quantized MobileNetV2 [<xref ref-type="bibr" rid="ref-36">36</xref>], tiny CNN for fall prediction [<xref ref-type="bibr" rid="ref-37">37</xref>], EtinyNet [<xref ref-type="bibr" rid="ref-38">38</xref>], Weight quantization based on clustering [<xref ref-type="bibr" rid="ref-39">39</xref>], layer-wise quantization [<xref ref-type="bibr" rid="ref-40">40</xref>] and extreme quantization [<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Results of post-training quantization methods (&#x002A;DSC &#x003D; dice similarity coefficient)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">Dataset/Model</th>
<th align="center">Bits<break/>W/A</th>
<th align="center">Baseline accuracy</th>
<th align="center">Accuracy</th>
<th align="center">FP32 model</th>
<th align="center">Compressed model</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>UFPR<break/>ResNet-50</td>
<td>2</td>
<td>99.6%</td>
<td>99.92%</td>
<td>328.5 M</td>
<td>20.5 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>CIFAR-10<break/>VGG-Like</td>
<td>3</td>
<td>93.5%</td>
<td>93%</td>
<td>20.44 M</td>
<td>1.94 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>ILSVRC-12<break/>ResNet-18</td>
<td>3/3</td>
<td>69.2%</td>
<td>68.6%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>CIFAR-100<break/>VGG-16</td>
<td>2</td>
<td>66%</td>
<td>66%</td>
<td>138 M</td>
<td>23 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>CIFAR-100<break/>ResNet-20</td>
<td>4</td>
<td>68.7%</td>
<td>69.4%</td>
<td>0.27 M</td>
<td>0.19 MB (Memory)</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>BRATS<break/>Residual UNet</td>
<td>2</td>
<td>0.8418 DSC&#x002A;</td>
<td>0.8363 DSC&#x002A;</td>
<td>5.9 M</td>
<td>0.25 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>LiTS<break/>3D UNet</td>
<td>2</td>
<td>0.7843 DSC&#x002A;</td>
<td>0.7509</td>
<td>94.8 M</td>
<td>3.1 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>MobileNetV2<break/>Tongue Dataset</td>
<td>8</td>
<td>99%</td>
<td>98%</td>
<td>26 M (Flash memory)</td>
<td>5.4 (Flash memory)</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>SisFall<break/>TinyCNN</td>
<td>8</td>
<td>98.4%</td>
<td>98.3%</td>
<td>Parameter size &#x003D; 546</td>
<td>Parameter size &#x003D; 546</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>ETinyNet</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>66.5%</td>
<td>&#x2013;</td>
<td>477 K</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>1D MINIST<break/>LeNet5</td>
<td>Clustering</td>
<td>83.7%</td>
<td>83.5%</td>
<td>114.06 M</td>
<td>70.27 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>LiTs, BraTS2020<break/>3D Unmet</td>
<td>4/4</td>
<td>78.36%, 84.8%</td>
<td>78.26%, 84.4%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>YOLOv5l<break/>MPQ-YOLOl</td>
<td>1/1 &#x002B; 4/4</td>
<td>87%</td>
<td>74%</td>
<td>178.4 M</td>
<td>12.6 M</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Kolf et al.&#x2019;s mixed precision quantization [<xref ref-type="bibr" rid="ref-30">30</xref>] reduces model memory footprint by first quantizing a 32-bit floating point model to an 8-bit integer model, then iteratively training and reducing weights to as low as two bits. As a result, the method achieves a 16-fold reduction in memory with minimal accuracy loss on models like ResNet18, ResNet50, and MobileFaceNet. Vector Quantization [<xref ref-type="bibr" rid="ref-31">31</xref>] balances quantization loss and model accuracy, achieving 5 to 16 times size reduction with minimal accuracy loss on different models. The quantization-loss-aware algorithm [<xref ref-type="bibr" rid="ref-32">32</xref>], uses Taylor&#x2019;s expansion to quantize deep neural network weights to low bit-widths for stable convergence. Smart-DNN&#x002B; [<xref ref-type="bibr" rid="ref-33">33</xref>] provides a layer-wise quantization protocol to compress models from full-precision to binary quantization. The multi-branch topology [<xref ref-type="bibr" rid="ref-34">34</xref>] avoids quantization error due to bit-width switching by using fixed 2-bit weights in each branch and combining branches to achieve the desired bit-width.</p>
<p>Ultra-low bit quantization [<xref ref-type="bibr" rid="ref-35">35</xref>] employs an adaptive quantizer with two tunable parameters to minimize quantization error. Both weights and activation functions are quantized, and full-precision weights are discarded after training. It achieves a comparable segmentation accuracy with one- and two-bit quantization. In [<xref ref-type="bibr" rid="ref-36">36</xref>], a 8-bit quantized MobileNetV2 model has been deployed on a Cam H7 Plus embedded device to classify oral cavity cancer. Tiny CNN for fall prediction [<xref ref-type="bibr" rid="ref-37">37</xref>] is trained on wearable sensor data, achieving over 98% accuracy on two datasets after being quantized to 8 bits. Xu et al.&#x2019;s EtinyNet [<xref ref-type="bibr" rid="ref-38">38</xref>] is a seven-layer CNN with an extremely tiny backbone for visual processing, operating at 160 mW and processing 30 frames per second (FPS).</p>
<p>Automated quantization and retraining of the deep neural network models using multi-objective optimization (NSGA-II) is proposed in [<xref ref-type="bibr" rid="ref-39">39</xref>]. The authors have used two objectives, accuracy and model size, to find the Pareto front of the solutions. Quantization is done through vector quantization by representing the weight space into sub-regions, and a centroid of every sub-region is selected as a quantized weight representation. Results of various models and datasets showed a decrease in model size by at least 30% without significant loss of accuracy. Zhang et al. [<xref ref-type="bibr" rid="ref-40">40</xref>] have proposed a layer-wise quantization framework with an alternating direction method of multipliers to achieve fast convergence during deep neural network training. They have applied the proposed method for volumetric medical image segmentation. A weight regularization term is added to the quantization objective to retain the knowledge of the full precision weights.</p>
<p>The MPQ-YOLO [<xref ref-type="bibr" rid="ref-41">41</xref>] is an ultra-low mixed quantization of the YOLO model for edge devices, combining 1-bit backbone and 4-bit head quantization with a dedicated training policy. The backbone, containing convolutional layers, is highly compressed with 1-bit quantization, while the head, sensitive to quantization error, uses 4-bit quantization. When compressing the model with extreme quantization, a compromise is desired between efficiency and the model&#x2019;s accuracy or performance [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<p>To summarize, quantization methods map full-precision values to the nearest quantized value, but clustering-based quantization assigns similar weight values to a single cluster center. As evident from <xref ref-type="table" rid="table-6">Table 6</xref>, It is challenging to decide which quantization method is better considering the accuracy and model compression ratio tradeoff. Regularization, which prevents overfitting in neural networks, must be modified for quantized networks. Post-training quantization can significantly drop model performance, so retraining or fine-tuning is necessary. Therefore, quantization-aware training simulates the effect of low precision during training, allowing the model to learn parameters that improve performance and reduce quantization errors simultaneously. Quantization to a lower number of bits may reduce the model size considerably. However, activation functions associated with each layer of the deep neural networks are difficult to quantize due to their nonlinear nature. A combination of the quantization of weights and the activation function achieves better results in terms of performance. Moreover, the quantization levels depend on the model architecture and the problem it is solving. A combination of weight and activation function quantization performs better in accuracy in many models. Therefore, the results of quantization-aware training are shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Results of quantization-aware training methods</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">Dataset/Model</th>
<th align="center">Bits W/A</th>
<th align="center">Baseline accuracy</th>
<th align="center">Accuracy</th>
<th align="center">FP32 model</th>
<th align="center">Compressed model</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>CIFAR-10<break/>GXNOR-Nets</td>
<td>3</td>
<td>92.88%</td>
<td>92.5%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>CIFAR-10<break/>DenseNet</td>
<td>2</td>
<td>94.3%</td>
<td>94%</td>
<td>7 M</td>
<td>0.49 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>CIFAR-10<break/>ResNet-20</td>
<td>2</td>
<td>91.8%</td>
<td>90.9%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>CIFAR-10<break/>ResNet50</td>
<td>2.68/4</td>
<td>76%</td>
<td>75.4%</td>
<td>97.28 M</td>
<td>9.45 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>CIFAR-10<break/>ResNet-18</td>
<td>8</td>
<td>88%</td>
<td>79%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>ImageNet<break/>ResNet-18</td>
<td>3/3</td>
<td>70.2%</td>
<td>69.2%</td>
<td>11.6 M</td>
<td>&#x007E;45% decrease</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>ImageNet<break/>ResNet-18</td>
<td>5/5</td>
<td>70.2%</td>
<td>70.4%</td>
<td>11.6 M</td>
<td>&#x007E;35% decrease</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-7fn1" fn-type="other">
<p>Note: W: Weights, A: Activation.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The work in [<xref ref-type="bibr" rid="ref-42">42</xref>] has introduced a multi-step activation function quantization method and a derivative approximation technique for backpropagation in discrete deep neural networks. This algorithm quantizes both activation functions and weights to ternary values, creating a binary sparse network. A symmetric mixture of Gaussian modes [<xref ref-type="bibr" rid="ref-43">43</xref>] is a soft quantization-aware method that uses low-bit fixed quantization. It involves training with real-valued weights and generating posterior distributions for post-quantization. In [<xref ref-type="bibr" rid="ref-44">44</xref>], a training mechanism in a finite weight search space is presented. A Hessian-based mixed precision quantization-aware training method is used to optimize the search for the best bit configuration [<xref ref-type="bibr" rid="ref-45">45</xref>], employing a Pareto frontier method based on the average Hessian trace for different configurations. The quantized process does not consider energy consumption in quantization-aware training. Hence, Hamming Weight-based Energy Aware Quantization (HAMQ) is proposed in [<xref ref-type="bibr" rid="ref-46">46</xref>]. Considering the Compute-in-memory (CIM) architecture for edge devices with limited resources, HAMQ provides better quantization for energy efficiency. Jung et al. [<xref ref-type="bibr" rid="ref-47">47</xref>] proposed a trainable quantizer based on a quantization-interval-learning (QIL) framework. They obtained the optimal values of quantization intervals by minimizing the task loss of the network. Using the QIL framework, better quantization levels of weights and activation functions are achieved on ResNet variants without degrading the ImageNet dataset&#x2019;s accuracy, as shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>

</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Model Compression Using Low-Rank Factorization</title>
<p><xref ref-type="table" rid="table-8">Table 8</xref> provides the results of low-rank factorization methods. The tensor decomposition method leverages shared tensor structures and parameters across DNN layers [<xref ref-type="bibr" rid="ref-48">48</xref>]. The optimization achieves 93% accuracy with an eight-fold compression on ResNet-18 using CIFAR-10. It is important to note that matrix and tensor decomposition identify and remove redundant parameters, decomposing larger weight matrices into smaller, storage-friendly ones. On the other hand, convolutional layers are factorized into depth-wise or pointwise convolutions, reducing the computational time during inference.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Results of low-rank factorization methods to compress the model</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">Factorization</th>
<th align="center">Model</th>
<th align="center">Dataset</th>
<th align="center">Compression</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>Tensors decomposition</td>
<td>ResNet-18</td>
<td>CIFAR-10</td>
<td>8 times</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>Sparse low-rank factorization</td>
<td>VGG-16, VGG-19</td>
<td>CIFAR-10</td>
<td>C&#x002A; &#x003D; 3.6</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>Tensor decomposition</td>
<td>VGG16, VGG19, ResNet-50</td>
<td>CIFAR-100</td>
<td>Compression ratio 98%, 97%, 95%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
<td>Joint matrix decomposition</td>
<td>ResNet-50<break/>ResNet-34</td>
<td>CIFAR-10</td>
<td>6 times<break/>22 times</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-52">52</xref>]</td>
<td>Deep compression</td>
<td>VGG-16</td>
<td>ImageNet</td>
<td>15 times</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-8fn1" fn-type="other">
<p>Note: &#x002A;C &#x003D; Sparsified Compression ratio (ratio between truncated SVD and SLR).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The low-rank factorization is effective for compressing fully connected layers in deep neural networks and can act as a regularization method to improve performance. Higher-order singular value decomposition (SVD) can compress convolutional layers. Therefore, sparse low-rank factorization, using SVD and truncating SVD matrices, achieves good compression ratios on VGG16, VGG19, and Lenet5 [<xref ref-type="bibr" rid="ref-49">49</xref>]. However, it is challenging to deploy bigger models due to mobile devices&#x2019; limited storage, computational power, and energy. Hence, the Tensor decomposition method can compress the network parameters [<xref ref-type="bibr" rid="ref-50">50</xref>]. A rank decomposition algorithm decomposes the tensor into a limited number of principal vectors. Hence, all the convolutional layers of the DNN can be decomposed and compressed. The original parameters can be reproduced and applied from these principal vectors in DNN architecture. When applied to VGG16, VGG19, and ResNet50, a compression ratio of 98%, 97%, and 95% is achieved, respectively, with insignificant loss of accuracy on the CIFAR100 dataset. Due to the similarity among weight tensors of the different layers within a neural network model, simultaneous tensor decomposition can compress the model without considerably decreasing performance. The authors developed two tensor decompositions for fully or partially structure-sharing cases.</p>
<p>Chen et al. [<xref ref-type="bibr" rid="ref-51">51</xref>] proposed joint matrix decomposition to improve the compression ratio. They have suggested that compressing the convolutional layers separately produces less efficient CNN. Hence, three different joint matrix decomposition methods are tested on three CNN architectures. After compression, finetuning recovers some of the accuracy loss. On ResNet-34, a good compression ratio is achieved with an accuracy loss of less than 1%. However, the compression ratio on ResNet-50 is low (x6.2), with an accuracy loss of 3%. A deep compression technique is proposed in [<xref ref-type="bibr" rid="ref-52">52</xref>] based on global average pooling, iterative filter pruning, applying truncated SVD on the fully connected layers, and quantization. A pre-trained VGG16 model is used for the experiments on the ImageNet dataset. They have achieved a compression rate of 60 times (VGG16: 138 M, VGG16 compressed: 8.66 M parameters) with degradation in the classification accuracy of 0.85%.</p>
</sec>
<sec id="s3_1_4">
<label>3.1.4</label>
<title>Model Compression Using Knowledge Distillation</title>
<p><xref ref-type="table" rid="table-9">Table 9</xref> shows results on knowledge distillation methods. The background on knowledge distillation techniques has been provided in <xref ref-type="sec" rid="s2_1_1">Section 2.1.1</xref>. The work in [<xref ref-type="bibr" rid="ref-53">53</xref>] has used Efficient-Net-B0, with some modifications, as a teacher model to train a lightweight student model using a knowledge distillation technique. The student model is a simplified Efficient-Net-B0 model with a frozen convolution head and MBConvBlock-25. A response-based knowledge distillation method is used to calculate distillation loss. A distillation loss in the response-based knowledge distillation is calculated based on the difference between the logit of the teacher and student models. The student model is trained along with the teacher model using distillation loss and ground truth labels. If the gap between the teacher and student models is too big, then knowledge distillation may fail, and the student model may not follow the teacher model [<xref ref-type="bibr" rid="ref-54">54</xref>]. Hence, the teacher model can be simplified by using Tucker decomposition of the tensors to ensure that the student model follows the teacher model during knowledge distillation. Using VGG16 as a teacher model and replacing the convolutional layers with the tucket decomposition layers, a simpler model can distill the knowledge to the student model (LeNet in this paper). The decomposition rank is 16 to maximize the accuracy improvement of the student model.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Results of knowledge distillation methods used to compress the models</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col/>
<col/>
<col/>
<col/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th>Ref.</th>
<th align="center">Dataset</th>
<th colspan="4">Teacher model</th>
<th colspan="3">Student model</th>
</tr>
<tr>
<th></th>
<th></th>
<th align="center">Model</th>
<th>Accuracy</th>
<th>Size</th>
<th>Flops</th>
<th>Size</th>
<th align="center">Accuracy</th>
<th>Flops</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>BOSSbase</td>
<td>Efficient-Net-B0</td>
<td>97.8%</td>
<td>5.3 M</td>
<td>0.56 B</td>
<td>3.6 M</td>
<td>98%</td>
<td>0.34 B</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-54">54</xref>]</td>
<td>UTKinect-Action3D</td>
<td>VGG16</td>
<td>99%</td>
<td>138 M</td>
<td>15 B</td>
<td>4.3 M</td>
<td>98.7%</td>
<td>4.4 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-55">55</xref>]</td>
<td>DFU dataset</td>
<td>InceptionV3</td>
<td>98%</td>
<td>28.5 M</td>
<td>5.71 M</td>
<td>0.49 M</td>
<td>96%</td>
<td>0.42 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>FaceForensics&#x002B;&#x002B;</td>
<td>XceptionNet</td>
<td>96.3%</td>
<td>20.8 M</td>
<td>6 B</td>
<td>2.74 M</td>
<td>96%</td>
<td>1.2 B</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-57">57</xref>]</td>
<td>CIFAR-100</td>
<td>ResNet56</td>
<td>72.3%</td>
<td>0.85 M</td>
<td>&#x2013;</td>
<td>0.27 M</td>
<td>72%</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>URBAN100</td>
<td>RCAN</td>
<td>29.09 PSNR</td>
<td>15.44 M</td>
<td>35.3 G</td>
<td>4.28 M</td>
<td>32.85 PSNR</td>
<td>9.79 G</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>Histopathologic cancer dataset</td>
<td>ResNet50</td>
<td>95.75%</td>
<td>26 M</td>
<td>&#x2013;</td>
<td>11 M</td>
<td>95.8%<break/>96.6% (Ensemble)</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>Vaihingen</td>
<td>ResNet50</td>
<td>88%</td>
<td>24 M</td>
<td>4.1 B</td>
<td>14 M</td>
<td>84%</td>
<td>116 B</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Amjad et al. [<xref ref-type="bibr" rid="ref-55">55</xref>] proposed a lightweight model called DFU-LWNet and trained it by InceptionV3 (teacher model) through knowledge distillation. The proposed model contains three convolutional layers with the max-pooling layers, and convolutional modules are borrowed from the Efficient-Net model. Finally, a customized classifier is added. Response-based knowledge distillation is used to transfer knowledge. Comparable classification accuracy is achieved with a DFU-LWNet with only 0.48 M parameters compared to the InceptionV3 model with 28.5 M parameters. Xu et al. [<xref ref-type="bibr" rid="ref-56">56</xref>] used a feature-based knowledge distillation technique to train the student model using cross-entropy loss, knowledge distillation loss, and gradient-guided feature distillation loss. XceptionNet is used as a teacher model and trained on the deepfake video dataset FaceForensics&#x002B;&#x002B; dataset. In the gradient-guided feature loss (feature-based distillation), the intermediate layers of the teacher and student models are mapped to define the distillation target. Gradient-guided weights define the importance of different channels in the feature maps. A decayed teaching strategy I used to modify the gradient-guided weights. A comparable accuracy is achieved by the student model with only 2.74 M parameters compared to the teacher model (20.8 M parameters). Usually, the SoftMax scaling factor (fixed temperature value) does not change for the data samples and considers all the samples of equal difficulty level. Hence, Ham et al. [<xref ref-type="bibr" rid="ref-57">57</xref>] explored the effect of data difficulty level on knowledge distillation. The model distills the knowledge based on three difficulty levels. The difficulty level is estimated through the Euclidean distance between the teacher&#x2019;s and pruned teacher&#x2019;s predictions. They have tested their method on various combinations of teacher-student models for CIFAR-100 and FGVR datasets.</p>
<p>A hybrid knowledge distillation from intermediate layers of the teacher and student model is used to create a single image super-resolution [<xref ref-type="bibr" rid="ref-58">58</xref>]. For this purpose, auxiliary up-samplers are added to the teacher and student models to create intermediate super-resolution images. Once the up-samplers of the teacher model are trained, the frequency similarity matrix (using discrete wavelet transform) and adaptive channel fusion are used to distill the knowledge and update the student model&#x2019;s up-samplers parameters. The models are trained using the DIV2K dataset and tested on various datasets, including BSD100 and Urban100. The comparable peak signal-to-noise ratio (PSNR) achieves a compression ratio of four times in 2X and 4X super-resolution. Niyaz et al. [<xref ref-type="bibr" rid="ref-59">59</xref>] used an interesting concept of knowledge distillation from one teacher model to multiple student models with collaborative learning among the student models.</p>
<p>Furthermore, they have compared the offline KD (training the teacher model first) and online KD (both teacher and students trained simultaneously) effect on the classification performance. An ensemble makes the final prediction of the prediction of the student models. Moreover, different learning styles, final prediction with one student, and intermediate layer features with other students are also investigated. Results of ResNet50 (teacher) and ResNet18 (students) are given in the table below for the histopathologic cancer detection dataset.</p>
<p>A multidimensional KD approach is adopted to improve the capability of transferring knowledge from the teacher to the student model [<xref ref-type="bibr" rid="ref-60">60</xref>]. ResNet-50 is the backbone of the teacher model, and MobileNet V2 is the backbone of the student model. Outputs of five different channels from the teacher and student models are fused, multiscale information is extracted, and feature-based distillation loss is calculated. Apart from this distillation loss, five other distillation losses, namely, inter-layer relation-based distillation loss, intra-layer feature-based distillation loss, wavelet transform response-based distillation loss, and logits distillation loss. They tested their methodology on two datasets, Potsdam and Vaihingen, and compared them with other published methods to prove the efficacy of the proposed method.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Inference Strategies</title>
<p>Once the model is deployed on the edge device, the inference mechanism has two options. Either the inference is performed entirely on the device, with the results utilized and communicated by the edge device (on-device inference), or the inference is carried out in collaboration with other edge devices or edge servers (distributed inference). <xref ref-type="sec" rid="s3_2_1">Sections 3.2.1</xref> and <xref ref-type="sec" rid="s3_2_2">3.2.2</xref> describe on-device inference and distributed inference, respectively. A summary of different inference mechanisms is presented in <xref ref-type="table" rid="table-10">Table 10</xref>.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Comparison of inference strategies used in the edge ML</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">Inference strategy</th>
<th align="center">Dataset</th>
<th align="center">Model</th>
<th align="center">Performance</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>On-device inference</td>
<td>KTH, UCI</td>
<td>CNN (7.7 K parameters)</td>
<td>Inference &#x003C; 3 s<break/>Energy &#x003D; 250 mJ</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
<td>Model split<break/>(Multiple points)</td>
<td>&#x2013;</td>
<td>GoogleNet</td>
<td>Latency &#x003D; 2.058 s<break/>Energy &#x003D; 7616 mJ</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>Model split<break/>(Layers based)</td>
<td>Customized</td>
<td>AlexNet</td>
<td>Save inference time by 12% to 66%<break/>Best latency &#x003D; 1.45 s</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>Task-aware<break/>Splitting</td>
<td>Metal casting dataset</td>
<td>CNN</td>
<td>Better latency under different bandwidth</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
<td>Optimal<break/>Splitting</td>
<td>&#x2013;</td>
<td>ResNet, AlexNet, VGG</td>
<td>Latency and energy consumption under different settings</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-66">66</xref>]</td>
<td>Model<break/>Partitioning</td>
<td>ImageNet</td>
<td>VGG-16, ResNet-34, MobileNetV1</td>
<td>84% reduction in Latency, 14x speedup</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-67">67</xref>]</td>
<td>Model split<break/>(Layers based)</td>
<td>ModelNet40</td>
<td>VGG-16</td>
<td>Inference latency compared with different Inference schemes</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-68">68</xref>]</td>
<td>Horizontal<break/>Partitioning</td>
<td>&#x2013;</td>
<td>AlexNet, VGG16-BN, ConvNext</td>
<td>More robust at the expense of maximum memory and energy</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-69">69</xref>]</td>
<td>Energy-aware<break/>Model split</td>
<td>&#x2013;</td>
<td>VGG-16, MobileNetV2</td>
<td>Load reduction in ESS &#x003D; 15%&#x2013;20%, LSS &#x003D; 60%, Energy saving 18%, 52%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-70">70</xref>]</td>
<td>Privacy-aware<break/>Model split</td>
<td>CIFAR-10, MNIST</td>
<td>Customized CNN</td>
<td>Strong privacy, better runtime on datasets</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>On-Device Inference</title>
<p>The on-device inference is feasible when the small model size and the edge device have sufficient computational and energy resources. TinyML models, typically with parameter sizes under 1 M, often make on-device inference a viable option. For instance, Nooruddin et al. [<xref ref-type="bibr" rid="ref-61">61</xref>] introduced a two-stream multi-resolution fusion method for human activity recognition from video data. They utilized a customized CNN model with quantization-based conversion. This compressed model was deployed on three tiny edge devices and tested on the KTH and UCF11 datasets. The proposed model achieved 98% accuracy, with 7.7 K parameters requiring 575 KB of memory. The inference time on the Arduino Nano 33 BLE Sense device was under 3 s, with a power consumption of 250 mJ.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Distributed Inference</title>
<p>A comparison of state-of-the-art inference methods is provided in <xref ref-type="table" rid="table-10">Table 10</xref>.</p>

<p>The DeepWear framework in [<xref ref-type="bibr" rid="ref-62">62</xref>] offloads deep learning tasks from wearable sensors to handheld devices via Bluetooth, eliminating the need for an internet connection. This approach enhances model performance and reduces the energy footprint of edge devices. The authors have explored various model-splitting strategies, examining their impact on latency and energy consumption. The model has achieved an inference speedup of two to three times and energy savings of 18% to 32%.</p>
<p>In [<xref ref-type="bibr" rid="ref-63">63</xref>], a DNNOff strategy is proposed, which comprises three components. The extraction component extracts the structure and parameters of the DNN mode, an offloading mechanism, and an estimation model that defines the offloading strategy on edge devices. In the adaptive offloading scheme, a decision must be made for each layer on whether the computation will be on the edge of the server. A random forest regression is used to predict the execution time of each layer. The offloading schemes are tested on the AlexNet model, splitting the model into edge server, cloud server, and edge devices. Gautam et al. [<xref ref-type="bibr" rid="ref-64">64</xref>] proposed a task-aware DNN splitting scheme for EdgeML smart manufacturing. The model contains sensing and edge layers. The edge layer consists of an edge computing node and an edge server. Each layer of DNN is a potential splitting point, and the splitting policy is based on the task execution time, bandwidth between the device and the edge server, and profiling parameters. An optimal splitting policy is adopted based on minimizing average execution time. Energy optimization of the edge devices is not considered [<xref ref-type="bibr" rid="ref-65">65</xref>].</p>
<p>Self-aware model partitioning [<xref ref-type="bibr" rid="ref-66">66</xref>] considers the model partitioning by the edge device according to its computational resources status or time constraint on the inference. It engages a subset of available devices with sufficient resources to collaborate in the inference task by offloading partial inference tasks. Collaborative inference with two to four devices shows an improvement in inference performance. Another work on collaborative inference on edge devices focuses on a selective scheme that reduces data redundancy and bandwidth resource availability [<xref ref-type="bibr" rid="ref-67">67</xref>]. Multi-view images have a lot of spatial correlation taken by different edge devices. So, multi-view classification from edge devices can collaborate with centralized servers, splitting the inference tasks. In the selective ensemble inference, each edge device decides whether it provides its inference to the server for ensemble decision or not. The above methods focused on vertical partitioning of the model in which an entire layer is assigned to a node. Hence, different layers are assigned to different nodes. This type of partitioning produces high throughput but has a higher failure risk. In case of failure of one node, the whole inference procedure is affected as the entire layer is comprised.</p>
<p>Guo et al. [<xref ref-type="bibr" rid="ref-68">68</xref>] proposed the RobustDiCE method for robust distribution of the tasks among the inference nodes. This method divides and distributes every model layer among the inference nodes (horizontal partitioning). In case of a node failure, it is easy to recover the inference flow of the model. Hence, in this scheme, neurons of every layer are evenly distributed and assigned to the edge devices taking part in the inference. The method is evaluated on AlexNet, VGG16, and ConvNext models against failures of the devices. Experimental results showed that accuracy does not drop significantly for one device failure. However, more than one device failure affects the accuracy by more than 20%. In [<xref ref-type="bibr" rid="ref-69">69</xref>], authors proposed two strategies of model splitting for battery-operated IoT devices and regular-powered IoT devices. In the early split strategy, the maximum part of the model is offloaded to the edge server, saving the energy consumption of the battery-operated edge devices.</p>
<p>In contrast, in the late-split strategy, layers of the model requiring heavy computation are offloaded to the edge server. Wang et al. [<xref ref-type="bibr" rid="ref-70">70</xref>] proposed a privacy-preserving protocol for the edge device and edge server collaborative inference in the industrial Internet of Things. Two edge servers participate in the inference and calculate the model output without knowing the data or the model. The results showed that the proposed method can achieve a good tradeoff between latency and throughput.</p>
<p>To summarize the results, the most effective inference strategy is the model splitting among the edge devices. Much research has been done on optimizing. Hence, the inference strategy may vary according to the available resources, inference demand, and energy consumption.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Learning Strategies</title>
<p>We have classified our results into (a) online learning strategies, (b) continual learning or life-long learning, and (c) federated Learning. <xref ref-type="sec" rid="s3_3_1">Sections 3.3.1</xref>&#x2013;<xref ref-type="sec" rid="s3_3_3">3.3.3</xref> provide results on online learning strategies, continual learning or life-long learning, and federated Learning. <xref ref-type="table" rid="table-11">Tables 11</xref> and <xref ref-type="table" rid="table-12">12</xref> summarize continual learning and federated Learning, respectively.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Comparison of continual learning methods in the edge ML</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">Model dataset</th>
<th align="center">TinyModel</th>
<th align="center">Baseline accuracy</th>
<th align="center">TinyModel accuracy</th>
<th align="center">Memory requirement</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-71">71</xref>]</td>
<td>Customized MLP<break/>Three diseases</td>
<td>Full precision</td>
<td>&#x2013;</td>
<td>92.8%</td>
<td>Buffer size: 2.8 MB<break/>Model size: 0.35 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-72">72</xref>]</td>
<td>MobileNetV2 OpenLORIS</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>97.6%</td>
<td>5.90 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-73">73</xref>]</td>
<td>MobileNetV1<break/>Core50</td>
<td>Quantized UNIT-8/UINT-7</td>
<td>77%</td>
<td>68%<break/>75%</td>
<td>&#x003C;4 MB<break/>&#x003C;40 MB</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-74">74</xref>]</td>
<td>MCUNet multiple</td>
<td>UINT-8</td>
<td>73.3%</td>
<td>73.7%</td>
<td>0.48 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-75">75</xref>]</td>
<td>ResNet-20<break/>CIFAR-10</td>
<td>3-bit gradient quantization</td>
<td>90%</td>
<td>89.8%</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Comparison of federated learning methods for Edge ML</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Ref.</th>
<th align="center">FL method</th>
<th align="center">Dataset</th>
<th align="center">No. of nodes</th>
<th align="center">Sample per node</th>
<th align="center">Performance</th>
<th align="center">Time complexity</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-76">76</xref>]</td>
<td>DeFL</td>
<td>Credit card</td>
<td>8000</td>
<td>&#x007E;31</td>
<td>89.5%</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-77">77</xref>]</td>
<td>DeFL</td>
<td>CIFAR-10</td>
<td>&#x007E;80</td>
<td>&#x2013;</td>
<td>&#x007E;80%</td>
<td>Model: 10.7 s<break/>W/update: 0.015 s</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-78">78</xref>]</td>
<td>HetFL</td>
<td>CIFAR-100</td>
<td>10</td>
<td>10</td>
<td>76%</td>
<td>FLOPs 11.5 M</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-79">79</xref>]</td>
<td>HierFL</td>
<td>MNIST</td>
<td>100/3&#x002A;&#x002A;</td>
<td>&#x2013;</td>
<td>97.9%</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-80">80</xref>]</td>
<td>HetFL</td>
<td>CIFAR-10</td>
<td>50</td>
<td>1000</td>
<td>64.5%</td>
<td>60 s (GPU: RTX3080)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-12fn1" fn-type="other">
<p>Note: DeFL: Decentralized FL, HetFL: Heterogenous FL, HierFL: Hierarchical FL. &#x002A;&#x002A; Nodes/servers.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>On-Device Learning</title>
<p>In offline strategies, large datasets are collected and used to train deep learning or machine learning models on high-performance computing resources. Once trained, models are reduced in size using techniques like quantization and pruning to fit edge devices, often called TinyML. Model inference can be performed on the edge device alone through model splitting, task offloading, or distributed inference across multiple edge devices. In decentralized learning, edge devices collaborate to train the model. On-device learning is beneficial when new data is continuously available, requiring continuous model updates. It necessitates reliable and continuous wireless communication between edge devices, typically within fixed topologies.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Continual Learning or Life-Long Learning</title>
<p>Continual learning, or life-long learning, involves continuously recording data and learning in a non-stationary environment without losing previously acquired knowledge. It can be done on edge devices, enhancing security and privacy, though memory constraints make it challenging. Various learning methods have been compared in <xref ref-type="table" rid="table-11">Table 11</xref>.</p>

<p>A continual learning framework for disease detection using wearable medical sensors has been proposed [<xref ref-type="bibr" rid="ref-71">71</xref>]. Authors have used a data preservation method to retain the most informative previously learned data. Subsequently, a multilayer perceptron (MLP) was trained and achieved an average accuracy of 92.8% on the edge device. Similarly, the latent replay concept was proposed to address catastrophic forgetting in continual learning by storing activations of initial model layers instead of raw data [<xref ref-type="bibr" rid="ref-72">72</xref>]. The work in [<xref ref-type="bibr" rid="ref-73">73</xref>] has adapted this for 8-bit quantized models, calling it quantized latent replay-based continual learning. They found a tradeoff between latency layers and accuracy. The work in [<xref ref-type="bibr" rid="ref-74">74</xref>] has explored on-device training on Cortex-M MCUs using MCUNet, employing dynamic sparse gradient updates for fully quantized training. Finally, the work in [<xref ref-type="bibr" rid="ref-75">75</xref>] has proposed a framework for continual learning in noisy, dynamic environments, using selective experience replay and low bit-width quantization, achieving minimal accuracy loss (0.02%) with ResNet-20 on CIFAR-10 under heavy image degradation.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Federated Learning</title>
<p>Architectures in this category vary based on aggregation methods and task assignments to edge devices. In centralized FL, a server connects to all edge devices or nodes, selects nodes, and aggregates the model, suitable for a limited number of nodes. In hierarchical FL, instead of a centralized server, various nodes act as aggregation nodes, with edge devices connected to these aggregation nodes. In decentralized FL, all nodes participate in training and aggregating the global model. Similarly, heterogeneous FL addresses variations in data distribution, communication environments, device hardware, and model architectures.</p>
<p><xref ref-type="table" rid="table-12">Table 12</xref> compares state-of-the-art federated learning methods in terms of various performance attributes. A deep autoencoder for FL uses non-iterative training to reduce training time [<xref ref-type="bibr" rid="ref-76">76</xref>]. In a multi-node environment, nodes send local model information via the Queuing Telemetry Transport (MQTT) broker protocol, which updates and aggregates this information. This method has proven effective in accuracy, latency, and energy consumption on several datasets. Similarly, Zhang et al. [<xref ref-type="bibr" rid="ref-77">77</xref>] have designed an efficient FL mechanism for edge devices successfully applied to MNIST and CIFAR-10 datasets. Yang et al. [<xref ref-type="bibr" rid="ref-78">78</xref>] have proposed a resource-efficient heterogeneous federated continual learning algorithm for edge devices, reducing resource consumption by dividing the model into adapter and retainer sub-models. Qiang et al. [<xref ref-type="bibr" rid="ref-79">79</xref>] have introduced a multi-layer federated edge learning framework using edge servers between devices and the cloud to reduce latency and energy consumption. Similarly, Cao et al. [<xref ref-type="bibr" rid="ref-80">80</xref>] have proposed feature-space and output-space alignments for aggregating local models to minimize performance loss due to data heterogeneity. They found that for the CIFAR100 dataset, achieving a target accuracy of 35 requires 30 communication rounds with significant data distribution deviation and 20 rounds with more uniform data distribution.</p>

<p>It is important to note that aggregating local model updates is crucial in federated learning. The issues in aggregating local model updates include imbalanced, varied quality, non-uniformly distributed data subsets among FL clients, and the heterogeneity of edge devices affecting global model training efficiency. Moreover, the convergence is also challenging due to device heterogeneity and asynchronous updates. However, standard aggregation averages model updates, outliers, or malicious updates can corrupt the global model. Therefore, clipping model updates within a defined range can counter outliers [<xref ref-type="bibr" rid="ref-81">81</xref>]. Furthermore, momentum-based aggregation, where edge devices send the momentum term and local model update, speeds up the convergence process [<xref ref-type="bibr" rid="ref-82">82</xref>].</p>
<p>Another critical factor in FL methods is the contribution of local model updates to the global model for defining edge device performance. Therefore, weights are assigned to each contributing edge device based on its reliability and representation in global model training. Similarly, incorporating predictive uncertainty during aggregation can enhance the global model&#x2019;s generalization capability [<xref ref-type="bibr" rid="ref-83">83</xref>]. Furthermore, quantizing local updates improves communication efficiency and energy management for edge devices. While homogeneous quantization simplifies aggregation, real-world scenarios often involve heterogeneous quantization, where devices have varying precision levels. Federated learning with heterogeneous quantization assigns different weights to edge devices to account for quantization errors [<xref ref-type="bibr" rid="ref-84">84</xref>]. For more details on aggregation and learning strategies in federated learning, refer to a comprehensive survey [<xref ref-type="bibr" rid="ref-85">85</xref>].</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results on Model Deployment on Resource-Constrained Devices Using TinyML</title>
<p><xref ref-type="sec" rid="s3">Section 3</xref> critically reviews model conversion, inference mechanisms, and learning strategies. This section reviews state-of-the-art model deployment techniques on resource-constrained devices using TinyML. Moreover, the results have been classified according to different sectors, as TinyML has numerous applications in various sectors. Therefore, we have identified six major sectors where TinyML-based model deployment has shown promising results. This classification aims to highlight the target problems (practical examples) in each sector, along with the corresponding ML/DL models.
<list list-type="roman-lower">
<list-item>
<p><bold>Smart agriculture, or smart farming</bold> sector, involves monitoring/collecting real-time data about various parameters (corps, livestock, soil quality, etc.) in farming and optimizing it to increase yields.</p></list-item>
<list-item>
<p><bold>Medical or healthcare with environmental safety</bold> sector targets monitoring vital signs, using wearable devices to diagnose early health anomalies. Moreover, it also includes monitoring environmental conditions (such as air quality, water quality, and weather patterns) to enable pre-emptive actions.</p></list-item>
<list-item>
<p><bold>Vehicles or the automotive</bold> sector collect and process real-time data for multiple driver assistance systems, enhancing vehicle safety and efficiency.</p></list-item>
<list-item>
<p><bold>The industrial and robotics sectors</bold> focus on monitoring equipment and facilities for signs of wear and tear to enable predictive maintenance, reducing downtime and maintenance costs.</p></list-item>
<list-item>
<p><bold>The energy sector</bold> mainly contains applications for managing renewable energy from different sources.</p></list-item>
<list-item>
<p><bold>Secure smart cities and the consumer electronics</bold> sectors include data collection from smart cameras and sensors to detect unusual activities or security breaches in real time, enhancing public safety. In addition to the security of smart cities, this sector also covers the security of smart homes using various consumer electronic devices. Other applications of TinyML in smart cities are discussed in the environmental sector (such as monitoring air quality, noise levels, and other environmental factors) and the vehicle sector (such as optimizing traffic flow).</p></list-item>
</list></p>
<p>The achieved results have been synthesized in different ways to perform a critical analysis and comparative evaluation of the methodologies. The synthesis results are shown in <xref ref-type="table" rid="table-13">Tables 13</xref>&#x2013;<xref ref-type="table" rid="table-16">16</xref>. <xref ref-type="table" rid="table-13">Table 13</xref> shows applications (practical problems) and associated ML/DL models used for each practical problem in different sectors. Similarly, <xref ref-type="table" rid="table-14">Table 14</xref> summarizes fifteen hardware development boards and compares their features for TinyML deployment. The development boards in <xref ref-type="table" rid="table-14">Table 14</xref> have been extracted from the selected research works. It implies that <xref ref-type="table" rid="table-14">Table 14</xref> compares selected research works regarding hardware development. In addition to the comparison in terms of hardware development, <xref ref-type="table" rid="table-15">Table 15</xref> classifies and compares the selected research studies in terms of three major software frameworks. Finally, <xref ref-type="table" rid="table-16">Table 16</xref> details the sensors used in all the chosen research work.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Applications and machine learning models in six identified sectors</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Sector name</th>
<th align="center">Applications (practical problems)</th>
<th align="center">Model type</th>
<th align="center">Year</th>
<th align="center">Ref.</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Smart agriculture or farming</td>
<td>Watering process automation</td>
<td>Not specified</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-87">87</xref>]</td>
</tr>
<tr>
<td>Maize leaf disease detection and classification</td>
<td>CNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-88">88</xref>]</td>
</tr>
<tr>
<td>Soil quality monitoring and management</td>
<td>CNN, RNN, Q-learning</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
</tr>
<tr>
<td>Maize leaf disease detection and classification</td>
<td>CNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
</tr>
<tr>
<td>Disease detection and classification in plants</td>
<td>CNN</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
</tr>
<tr>
<td rowspan="5">Smart healthcare</td>
<td>Face mask detection on faces</td>
<td>CNN</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
</tr>
<tr>
<td>Recognition of human activities using wearable sensors</td>
<td>CNN</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-93">93</xref>]</td>
</tr>
<tr>
<td>Energy-efficient healthcare decision support system</td>
<td>RF, SVM, DT</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-94">94</xref>]</td>
</tr>
<tr>
<td>Human activity recognition</td>
<td>CNN, LSTM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-95">95</xref>]</td>
</tr>
<tr>
<td>Real-time blood pressure (bp) estimation</td>
<td>CNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
</tr>
<tr>
<td rowspan="5">Smart automotive or smart vehicles</td>
<td>To process vehicular data on edge devices for fuel consumption prediction</td>
<td>AutoCloud &#x002B; TEDA</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td>Outlier detection and correction</td>
<td>TEDA-RLS</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
</tr>
<tr>
<td>Intrusion detection system</td>
<td>CNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-99">99</xref>]</td>
</tr>
<tr>
<td>To enhance the cybersecurity in electric vehicle</td>
<td>MLP, RF</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-100">100</xref>]</td>
</tr>
<tr>
<td>Real-time driver behavior analysis</td>
<td>AutoCloud &#x002B; TEDA</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-101">101</xref>]</td>
</tr>
<tr>
<td rowspan="5">Industrial automation</td>
<td>Identifying anomalies using autoencoders</td>
<td>Neural Networks</td>
<td>2021</td>
<td>[<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
</tr>
<tr>
<td>Operational efficiency and safety</td>
<td>CNN</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
<tr>
<td>Reactive and dynamic online control for robots</td>
<td>ADMM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>Vibration-based fault diagnosis of machines</td>
<td>CNN</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
</tr>
<tr>
<td>Identification of various bolt defects in steel structures</td>
<td>FOMO</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
</tr>
<tr>
<td rowspan="5">Energy sector</td>
<td>Real-time fault diagnosis and classification of defects in PV modules</td>
<td>CNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
</tr>
<tr>
<td>Fault detection in PV modules</td>
<td>DCNN</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-108">108</xref>]</td>
</tr>
<tr>
<td>Forecasting solar energy yield</td>
<td>LSTM, BiGRU, BiLSTM, BiRNN</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-109">109</xref>]</td>
</tr>
<tr>
<td>Energy consumption prediction on mobile devices</td>
<td>LSTM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-110">110</xref>]</td>
</tr>
<tr>
<td>Upgradation of energy distribution panel</td>
<td>LSTM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
</tr>
<tr>
<td rowspan="5">Consumer electronics and security</td>
<td>Detection of jamming attacks in wireless networks</td>
<td>CNN</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-112">112</xref>]</td>
</tr>
<tr>
<td>Monitoring the condition of handheld power tools</td>
<td>CNN</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-113">113</xref>]</td>
</tr>
<tr>
<td>To detect and combat various Wi-Fi attacks</td>
<td>DNN, LSTM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-114">114</xref>]</td>
</tr>
<tr>
<td>Real-time health monitoring and intrusion detection</td>
<td>Naive Bayes and SVM</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-115">115</xref>]</td>
</tr>
<tr>
<td>Energy harvesting resource allocation for Unmanned Aerial Vehicles (UAV)-assisted TinyML consumer electronics</td>
<td>Naive Bayes, SVM, RF, KNN</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-116">116</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Summary of development boards in selected research works</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">S. no.</th>
<th align="center">Development boards</th>
<th align="center">MCU</th>
<th align="center">Core (CPU)</th>
<th align="center">Speed (MHz)</th>
<th align="center">Flash (MB)</th>
<th align="center">SRAM (KB)</th>
<th align="center">Ref.</th>
</tr>
</thead>
<tbody>
<tr>
<td>1.</td>
<td>Wio terminal</td>
<td>ATSAMD51</td>
<td>ARM Cortex-M4F</td>
<td>120</td>
<td>1</td>
<td>256</td>
<td>[<xref ref-type="bibr" rid="ref-87">87</xref>,<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
</tr>
<tr>
<td>2.</td>
<td>Arduino Nano 33 BLE Sense</td>
<td>nRF52840</td>
<td>ARM Cortex-M4</td>
<td>64</td>
<td>1</td>
<td>256</td>
<td>[<xref ref-type="bibr" rid="ref-88">88</xref>,<xref ref-type="bibr" rid="ref-90">90</xref>&#x2013;<xref ref-type="bibr" rid="ref-92">92</xref>,<xref ref-type="bibr" rid="ref-95">95</xref>&#x2013;<xref ref-type="bibr" rid="ref-98">98</xref>,<break/> <xref ref-type="bibr" rid="ref-102">102</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>,<break/> <xref ref-type="bibr" rid="ref-107">107</xref>,<xref ref-type="bibr" rid="ref-114">114</xref>]</td>
</tr>
<tr>
<td>3.</td>
<td>ESP32 development board</td>
<td>ESP32</td>
<td>Xtensa LX6, dual core,<break/>32-bit</td>
<td>240</td>
<td>4</td>
<td>520</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>&#x2013;<xref ref-type="bibr" rid="ref-98">98</xref>,<xref ref-type="bibr" rid="ref-100">100</xref>,<xref ref-type="bibr" rid="ref-101">101</xref>,<break/> <xref ref-type="bibr" rid="ref-105">105</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>]</td>
</tr>
<tr>
<td>4.</td>
<td>STM32F746 discovery kit</td>
<td>STM32F746NGH6</td>
<td>ARM Cortex-M7</td>
<td>216</td>
<td>1</td>
<td>320</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
</tr>
<tr>
<td>5.</td>
<td>Espress ESP-EYE</td>
<td>ESP32</td>
<td>Tensilica LX6</td>
<td>216</td>
<td>1</td>
<td>320</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
</tr>
<tr>
<td>6.</td>
<td>Arduino Uno R3</td>
<td>ATmega328P</td>
<td>8-bit AVR</td>
<td>16</td>
<td>32 KB</td>
<td>2</td>
<td>[<xref ref-type="bibr" rid="ref-94">94</xref>]</td>
</tr>
<tr>
<td>7.</td>
<td>Raspberry Pi 3 Model B</td>
<td>Single board computer</td>
<td>ARM Cortex A-53</td>
<td>1.2 GHz</td>
<td>16 GB</td>
<td>1 GB</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>,<xref ref-type="bibr" rid="ref-112">112</xref>]</td>
</tr>
<tr>
<td>8.</td>
<td>Raspberry Pi Pico</td>
<td>RP2040 chip</td>
<td>ARM Cortex M0&#x002B;</td>
<td>133</td>
<td>2</td>
<td>264 KB</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td>9.</td>
<td>Siemens board</td>
<td>Not mentioned</td>
<td>ARM Cortex-M4</td>
<td>90</td>
<td>1</td>
<td>320 KB</td>
<td>[<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
<tr>
<td>10.</td>
<td>Teensy 4.1</td>
<td>Not mentioned</td>
<td>ARM Cortex-M7</td>
<td>600</td>
<td>7.75</td>
<td>512 kB</td>
<td>[<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>11.</td>
<td>Crazyflie 2.1 Quadrotor</td>
<td>Not mentioned</td>
<td>ARM Cortex-M4</td>
<td>168</td>
<td>1</td>
<td>192 KB</td>
<td>[<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>12.</td>
<td>Arduino Nicla vision</td>
<td>STM32H747AII6</td>
<td>Dual Arm Cortex M7</td>
<td>480</td>
<td>2</td>
<td>1</td>
<td>[<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
</tr>
<tr>
<td>13.</td>
<td>Arduino Nano 33 IoT</td>
<td>&#x2013;</td>
<td>ARM Cortex-M0&#x002B; 32-bit SAMD21</td>
<td>48</td>
<td>256 KB</td>
<td>32 KB</td>
<td>[<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
</tr>
<tr>
<td>14.</td>
<td>Raspberry Pi 4</td>
<td>Broadcom BCM2711</td>
<td>ARM Cortex-A72</td>
<td>1.5 GHz</td>
<td>None</td>
<td>Up to 8 GB</td>
<td>[<xref ref-type="bibr" rid="ref-108">108</xref>]</td>
</tr>
<tr>
<td>15.</td>
<td>nRF52832</td>
<td>&#x2013;</td>
<td>ARM Cortex-M4</td>
<td>64 MHz</td>
<td>512 KB</td>
<td>64 KB</td>
<td>[<xref ref-type="bibr" rid="ref-113">113</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Summary of software platforms in selected research works</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">S. no.</th>
<th align="center">Software<break/>platform</th>
<th align="center">Supported development boards</th>
<th align="center">Language for deployment</th>
<th align="center">Open source</th>
<th align="center">Manuf.</th>
<th align="center">Ref.</th>
</tr>
</thead>
<tbody>
<tr>
<td>1.<break/><break/>2.<break/>3.</td>
<td>Edge impulse</td>
<td>Wio terminal<break/><break/>Arduino Nicla vision<break/>Arduino Nano 33 BLE sense</td>
<td>C&#x002B;&#x002B;</td>
<td>No</td>
<td>Edgeimpulse</td>
<td>[<xref ref-type="bibr" rid="ref-87">87</xref>,<xref ref-type="bibr" rid="ref-88">88</xref>], [<xref ref-type="bibr" rid="ref-92">92</xref>]<break/>[<xref ref-type="bibr" rid="ref-94">94</xref>]<break/>[<xref ref-type="bibr" rid="ref-106">106</xref>]<break/>[<xref ref-type="bibr" rid="ref-90">90</xref>]&#x00B8; [<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
</tr>
<tr>
<td>4.</td>
<td>Tensor Flow</td>
<td>Arduino Nano 33 BLE sense<break/>Wio terminal<break/>STM32F746 discovery kit<break/>Espress ESP-EYE<break/>Not mentioned<break/>Raspberry Pi 4</td>
<td>Python, C&#x002B;&#x002B; Java, javascript, swift, Go, TensorFlowLite</td>
<td>Yes</td>
<td>Google</td>
<td>[<xref ref-type="bibr" rid="ref-88">88</xref>,<xref ref-type="bibr" rid="ref-91">91</xref>,<xref ref-type="bibr" rid="ref-92">92</xref>,<xref ref-type="bibr" rid="ref-107">107</xref>,<xref ref-type="bibr" rid="ref-114">114</xref>]<break/>[<xref ref-type="bibr" rid="ref-89">89</xref>]<break/>[<xref ref-type="bibr" rid="ref-91">91</xref>]<break/>[<xref ref-type="bibr" rid="ref-91">91</xref>]<break/>[<xref ref-type="bibr" rid="ref-93">93</xref>]<break/>[<xref ref-type="bibr" rid="ref-108">108</xref>]</td>
</tr>
<tr>
<td>5.</td>
<td>TensorFlow lite</td>
<td>ESP32 development board<break/>Arduino Uno R3<break/>Arduino Nano 33 BLE sense<break/>Raspberry Pi 3 Model B<break/><break/>Raspberry Pi Pico<break/>Raspberry Pi 4<break/>nRF52832<break/>Not mentioned</td>
<td>Python, C&#x002B;&#x002B; Java, Swift, Objective-C, JavaScript</td>
<td>Yes</td>
<td>Google</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>,<xref ref-type="bibr" rid="ref-100">100</xref>,<xref ref-type="bibr" rid="ref-101">101</xref>,<break/><xref ref-type="bibr" rid="ref-105">105</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>]<break/>[<xref ref-type="bibr" rid="ref-94">94</xref>]<break/>[<xref ref-type="bibr" rid="ref-95">95</xref>&#x2013;<xref ref-type="bibr" rid="ref-98">98</xref>,<xref ref-type="bibr" rid="ref-102">102</xref>,<break/><xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-107">107</xref>],<break/>[<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>]<break/>[<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>]<break/>[<xref ref-type="bibr" rid="ref-108">108</xref>]<break/>[<xref ref-type="bibr" rid="ref-113">113</xref>]<break/>[<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Summary of sensors used in selected research works</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Board</th>
<th>Sensors</th>
<th>Ref.</th>
</tr>
</thead>
<tbody>
<tr>
<td>Wio terminal</td>
<td>Monitoring of temperature, humidity, moisture, light levels, and water levels</td>
<td>[<xref ref-type="bibr" rid="ref-87">87</xref>]</td>
</tr>
<tr>
<td/>
<td>Monitoring of temperature, humidity, moisture, water level, additional sensors such as motion sensors, microphone, and infrared emitter</td>
<td>[<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
</tr>
<tr>
<td>Arduino Nano 33 BLE sense</td>
<td>Motion sensors, environmental sensors, pressure sensors, microphones, and light sensors</td>
<td>[<xref ref-type="bibr" rid="ref-88">88</xref>]</td>
</tr>
<tr>
<td/>
<td>Camera</td>
<td>[<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
</tr>
<tr>
<td/>
<td>Motion sensors, camera module, LEDs, and buzzer</td>
<td>[<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
</tr>
<tr>
<td/>
<td>Motion sensors (accelerometer, gyroscope, magnetometer)</td>
<td>[<xref ref-type="bibr" rid="ref-95">95</xref>]</td>
</tr>
<tr>
<td/>
<td>Photoplethysmogram (ppg) sensors, ECG sensors</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
</tr>
<tr>
<td/>
<td>Fuel sensors, speed sensors, engine sensors, and environmental sensors</td>
<td>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td/>
<td>On-board diagnostics ii (obd-ii) port</td>
<td>[<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
</tr>
<tr>
<td/>
<td>3-axis accelerometer sensor</td>
<td>[<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
</tr>
<tr>
<td/>
<td>3-axis accelerometer for vibration measurement</td>
<td></td>
</tr>
<tr>
<td/>
<td>Infrared (ir) camera, temperature sensors, imu</td>
<td>[<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
</tr>
<tr>
<td>ESP32</td>
<td>GPS module, water level detection, and temperature sensors</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
</tr>
<tr>
<td/>
<td>Photoplethysmogram (ppg) sensors, ECG sensors</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
</tr>
<tr>
<td/>
<td>Fuel sensors, speed sensors, engine sensors, and environmental sensors</td>
<td>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td/>
<td>On-board diagnostics reader, embedded sensors in vehicles (speed sensors, fuel sensors, engine sensors, environmental sensors)</td>
<td>[<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
</tr>
<tr>
<td/>
<td>Network traffic sensors</td>
<td>[<xref ref-type="bibr" rid="ref-100">100</xref>]</td>
</tr>
<tr>
<td/>
<td>Monitoring speed, engine load, throttle position, GPS</td>
<td>[<xref ref-type="bibr" rid="ref-101">101</xref>]</td>
</tr>
<tr>
<td/>
<td>solar irradiance sensors, weather sensors</td>
<td>[<xref ref-type="bibr" rid="ref-109">109</xref>]</td>
</tr>
<tr>
<td>Raspberry Pi 3 Model B</td>
<td>Photoplethysmogram (ppg) sensors, ECG sensors<break/>Fuel sensors, speed sensors, engine sensors, and environmental sensors</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]<break/>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td/>
<td>Software defined radio (SDR)</td>
<td>[<xref ref-type="bibr" rid="ref-112">112</xref>]</td>
</tr>
<tr>
<td>Raspberry Pi Pico</td>
<td>Photoplethysmogram (ppg) sensors, ECG sensors<break/>Fuel sensors, speed sensors, engine sensors, and environmental sensors</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]<break/>[<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
</tr>
<tr>
<td>Arduino Nicla</td>
<td>High precision and quartz accelerometers, low precision accelerometers</td>
<td>[<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
</tr>
<tr>
<td/>
<td>acceleration sensors for collecting time-series vibration signals</td>
<td>[<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
</tr>
<tr>
<td>Arduino Uno R3</td>
<td>Monitoring of pulse rate and oxygen levels, sensors for room humidity and temperature, gas sensors for the detection of toxic gases, sound sensors</td>
<td>[<xref ref-type="bibr" rid="ref-94">94</xref>]</td>
</tr>
<tr>
<td>Siemens</td>
<td>3-axis accelerometer for vibration measurement.</td>
<td>[<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
<tr>
<td>Teensy 4.1</td>
<td>3-axis accelerometer, optitrack motion-capture system</td>
<td>[<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>Crazyflie</td>
<td>3-axis accelerometer, optitrack motion-capture system</td>
<td>[<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>nRF52832</td>
<td>Accelerometer for monitoring tool usage, temperature, and humidity sensors</td>
<td>[<xref ref-type="bibr" rid="ref-113">113</xref>]</td>
</tr>
<tr>
<td>&#x2013;</td>
<td>IMU sensors, heart rate, blood pressure, pulse rate, and temperature</td>
<td>[<xref ref-type="bibr" rid="ref-93">93</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the following, we briefly discussed all the selected research studies, organized into six categories. Consequently, <xref ref-type="sec" rid="s4_1">Sections 4.1</xref>&#x2013;<xref ref-type="sec" rid="s4_6">4.6</xref> provide a critical summary of selected research studies in smart agriculture, healthcare, vehicles, industry, energy, and security, respectively.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Smart Agriculture or Farming</title>
<p>It is a high-priority sector since it creates economic opportunities and generates most of the world&#x2019;s food. To satisfy the food demand, the major portion of agricultural tasks are required to be automated. While IoT communication technologies in smart agriculture have been reviewed previously [<xref ref-type="bibr" rid="ref-86">86</xref>], the use of intelligent IoT systems in the agriculture sector is a relatively new idea. This subsection provides practical examples (applications) of smart farming, where decision-making operations are performed using edge resources and TinyML techniques and tools.</p>
<p>One of the preliminary TinyML-based work in the agriculture sector is presented in [<xref ref-type="bibr" rid="ref-87">87</xref>]. The objective is to automate the watering of plants by monitoring environmental conditions and soil moisture, as shown in <xref ref-type="table" rid="table-13">Table 13</xref>. The system is developed around a Wio terminal board, as shown in <xref ref-type="table" rid="table-14">Table 14</xref>. The Edge Impulse (EI) is used for dataset generation, training, inference, and testing purposes, as shown in <xref ref-type="table" rid="table-15">Table 15</xref>. The employed sensors are for sensing the temperature, humidity, soil moisture, level of light, and level of water, as shown in <xref ref-type="table" rid="table-16">Table 16</xref>. While the work in [<xref ref-type="bibr" rid="ref-87">87</xref>] only elaborates on the initial idea, the complete IIoT systems for smart agriculture have been presented in [<xref ref-type="bibr" rid="ref-88">88</xref>,<xref ref-type="bibr" rid="ref-89">89</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>,<xref ref-type="bibr" rid="ref-112">112</xref>].</p>

<p>The IIoT system in [<xref ref-type="bibr" rid="ref-88">88</xref>] employs a customized CNN model to identify maize leaf disease. An important feature of this work is to evaluate the performance of TensorFlow and Edge Impulse frameworks. It has been observed that Edge Impulse is more user-friendly for data collection, labeling, and model deployment. Furthermore, lower memory footprint and power consumption make it suitable for deployment on low-powered edge devices. On the other hand, TensorFlow offers greater customization and control over the model architecture and training process. It has higher accuracy but requires more memory and computational resources. Another IIoT system in smart farming is presented in [<xref ref-type="bibr" rid="ref-89">89</xref>]. The system presents a power-aware and delay-aware TinyML model for monitoring and managing soil quality in agriculture. The model integrates dynamic voltage and frequency scaling (DVFS) as well as sleep/wake strategies based on genetic algorithms, energy harvesting, and task partitioning to optimize energy consumption and reduce delay. It employs CNN, RNNs, and Q-learning models. Like the work in [<xref ref-type="bibr" rid="ref-87">87</xref>], the Wio Terminal is used as the development board, while Tensor Flow is employed to deploy the model on the target development board.</p>
<p>A CNN-based approach for maize leaf disease detection and classification is presented in [<xref ref-type="bibr" rid="ref-90">90</xref>]. The customized CNN model, deployed on Arduino Nano 33 BLE Sense using the Edge Impulse platform, extracts important visual patterns from maize leaves, enhancing disease identification capabilities. The work presented in [<xref ref-type="bibr" rid="ref-91">91</xref>] aims to identify plant diseases using a customized CNN model. It has employed the TensorFlow framework to deploy the model on different development boards. In addition to the aforementioned applications in smart agriculture [<xref ref-type="bibr" rid="ref-87">87</xref>&#x2013;<xref ref-type="bibr" rid="ref-91">91</xref>], the other applications include but are not limited to real-time embedded prediction weather system and an end-to-end strategy for enhancing the security of the food supply chain.</p>
<p>To summarize the key findings of TinyML deployment in smart agriculture, it can be stated that TinyML-based agriculture systems offer significant results in analyzing specific conditions of individual farms, providing customized recommendations and actions based on real-time data. It has allowed practitioners to make decisions and perform tasks without constant human intervention, which is especially useful for large-scale farming. Moreover, by processing data locally, IIoT-based systems in smart farming are reducing the need for expensive cloud services. It not only lowers operational costs but also minimizes reliance on high-speed internet. Since data is processed on the device, the amount of data that needs to be transmitted is significantly reduced, conserving bandwidth and making the system more efficient.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Healthcare and Environmental Sector</title>
<p>An intelligent IoT-based healthcare decision support system is becoming paramount in today&#x2019;s healthcare domain. With the increasing prevalence of chronic diseases and the aging population, continuously and remotely monitoring patients&#x2019; health becomes crucial. IoT devices collect real-time data on vital signs, medication adherence, and lifestyle habits, enabling healthcare providers to make informed decisions swiftly. Consequently, it enhances patient outcomes by allowing early detection of potential health issues and timely interventions. Moreover, it reduces the burden on healthcare facilities by minimizing hospital visits and enabling efficient resource management.</p>
<p>One of the earlier target problems in healthcare is to process images and identify all those cases where a person is not wearing a face mask, as shown in [<xref ref-type="bibr" rid="ref-92">92</xref>]. Using the Edge Impulse platform, a low-cost solution using compressed CNN and transfer learning based on the MobileNetV1 architecture is deployed on the Arduino Nano 33 BLE Sense board. Another practical example of TinyML deployment in healthcare is classifying human activities from wearable sensors using an on-device deep learning inference mechanism, as shown in [<xref ref-type="bibr" rid="ref-93">93</xref>]. A lightweight CNN is trained offline on a stand-alone computer using TensorFlow. The specific microcontroller is not explicitly mentioned, but it is described as having minimal RAM (320 KB of SRAM) and operating at an 80 MHz clock frequency.</p>
<p>The classification of human activities in [<xref ref-type="bibr" rid="ref-93">93</xref>] is further extended to the analysis of human health parameters, as shown in [<xref ref-type="bibr" rid="ref-94">94</xref>]. The framework is centered on Raspberry Pi (fog layer) using TensorFlowLite; it reduces latency and ensures real-time data processing and analysis, which is crucial for time-sensitive healthcare applications. Another human activity recognition system is presented in [<xref ref-type="bibr" rid="ref-95">95</xref>]. The motion data (including acceleration, angular velocity, and magnetic field) is captured at a sampling frequency of 110 Hz. The compressed deep learning models (CNN and LSTM) with pruning and quantization techniques are deployed using TensorFlow Lite. Finally, a real-time blood pressure estimation problem is targetted in [<xref ref-type="bibr" rid="ref-96">96</xref>] using CNN types (AlexNet, LeNet, SqueezeNet, ResNet, and MobileNet) on various edge devices, including Raspberry Pi, ESP32, Raspberry Pi Pico, and Arduino Nano. The objective is to discuss the trade-offs between model size, accuracy, and inference time.</p>
<p>In addition to the aforementioned practical examples extracted from selected research works, there are other IoT healthcare applications, such as the prediction of chronic obstructive pulmonary disease, real-time activity tracking for elderly people and their nurses, a hand gesture recognition approach, etc. These advantages make TinyML a valuable technology in the healthcare sector, improving patient outcomes and operational efficiency.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Vehicles or Automotive Sector</title>
<p>Smart vehicles, also known as intelligent or connected vehicles, are transforming the automotive industry by integrating embedded technology, communication technology, and artificial intelligence technology to enhance safety, efficiency, and user experience [<xref ref-type="bibr" rid="ref-117">117</xref>]. The use of intelligent IoT systems has enabled the monitoring of vehicle components in real-time. For example, predicting maintenance proactively can reduce unexpected breakdowns and improve vehicle reliability. It, in turn, reduces the cost and warranty claims. Similarly, IIoT systems in the smart vehicles sector assist in collision avoidance and adaptive cruise control. It enables critical functions to operate without a constant internet connection, ensuring that safety features remain active in all conditions. Another problem in this sector is analyzing driver behavior and detecting signs of drowsiness or distraction, alerting the driver, or taking corrective actions to prevent accidents. It can also monitor the cabin environment (temperature, air quality) and adjust settings automatically to enhance passenger comfort. On-device processing ensures minimal delay (low latency) in executing critical functions, enhancing safety and performance. The low-cost hardware makes advanced features accessible in a broader range of vehicles and can be easily scaled across different vehicle models and types.</p>
<p>Scalable real-time processing of vehicular data streams on edge devices is presented in [<xref ref-type="bibr" rid="ref-97">97</xref>] by combining AutoCloud and TEDA (Typicality and Eccentricity Data Analytics). Real-time processing provides immediate insights and predictions, enhancing decision-making and operational efficiency. Four development boards (Raspberry Pi 3 Model B, ESP32 Wrover IE, Raspberry Pi Pico, and Arduino Nano 33 BLE) have been used to demonstrate the proposed algorithm&#x2019;s feasibility. Another application of smart vehicles is outlier detection and correction in vehicular data streams, as shown in [<xref ref-type="bibr" rid="ref-98">98</xref>]. The system employs the TEDA-RLS algorithm to increase data quality and reliability. The algorithm detects outliers using TEDA and corrects them using RLS filters. C&#x002B;&#x002B; is used to implement the algorithm on the ESP-32 microcontroller. In addition to real-time processing of vehicular data streams, intrusion detection systems are becoming increasingly important in the smart vehicles sector. An example of such a system can be found in [<xref ref-type="bibr" rid="ref-99">99</xref>], where a CNN-based approach is deployed on an nRF52840 microcontroller using TensorFlow Lite. A TinyML-based approach, using Multi-Layer Perceptron (MLP) and RF algorithms, is presented in [<xref ref-type="bibr" rid="ref-100">100</xref>] to enhance cybersecurity within the context of Electric Vehicle Charging Infrastructures.</p>
<p>The issues of real-time traffic management and driver behavior analysis in intelligent transportation Systems are addressed in [<xref ref-type="bibr" rid="ref-101">101</xref>] by presenting a multi-layered, stream-oriented data processing methodology for edge computing environments to detect and classify driver behavior patterns. The approach integrates the TEDA framework and an incremental clustering algorithm. The process starts by collecting data from vehicular physical sensors (e.g., speed, engine load, throttle position, RPM) and constructs a radar chart to represent multidimensional sensor readings as a polygon, with the area within the polygon reflecting the vehicle&#x2019;s resource utilization. Subsequently, it identifies and mitigates outliers in the data stream. Moreover, it implements a dynamic window to adapt to changes in data distribution and informs the incremental clustering algorithm about potential shifts. Finally, the AutoCloud algorithm is used for incremental clustering and does not require retaining datasets in memory. The TensorFlow Lite is used to deploy the machine learning models on the ESP32 microcontroller board. The customized AutoCloud algorithm was integrated into embedded systems using C&#x002B;&#x002B; and deployed on the hardware using the TensorFlow Lite library through Arduino IDE.</p>
<p>To summarize, the TinyML has been primarily used for predictive maintenance (to monitor vehicle components in real-time), driver assistance systems (to provide real-time alerts for lane departure, collision warnings, and driver drowsiness detection), and in-vehicle monitoring (monitoring the health and performance of various vehicle systems, such as the engine, brakes, and battery, to ensure optimal performance and safety). Moreover, the TEDA algorithm enhances the capabilities of ML and DL models in the automotive sector by providing robust, real-time data analytics, which is essential for maintaining vehicle safety, performance, and reliability.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Industrial and Robotics Sector</title>
<p>It is becoming increasingly important to analyze data from machinery in real time to predict failures before they occur. Similarly, monitoring production lines and energy usage is critical to ensure higher quality with reduced cost. In this context, TinyML enables robots to process sensory data in real-time for recognition and classification.</p>
<p>One of the TinyML-based framework pioneers, TinyOL, is presented in [<xref ref-type="bibr" rid="ref-102">102</xref>] where anomaly detection using Autoencoders on Arduino Nano 33 BLE Sense board using the TensorFlowLite framework. It processes streaming data, updates running mean and variance, scales input, makes predictions, and updates weights using online gradient descent algorithms. Consequently, an unsupervised autoencoder is transformed into a supervised anomaly classification model. The encoder&#x2019;s output and reconstruction error are used as features for classifying different anomaly patterns. The system&#x2019;s performance is evaluated in terms of fine-tuning (an ability to adapt existing neural network on MCUs to new data) and multi-anomaly classification (an ability to replace the final layer of an autoencoder with TinyOL, enabling post-training in an online mode).</p>
<p>The work in [<xref ref-type="bibr" rid="ref-103">103</xref>] leverages the W3C Web of Things (WoT) to semantically express IoT device capabilities through thing descriptions (TD). The TD describes IoT devices&#x2019; metadata and interactions in a standardized format. It introduces semantic models for on-device applications, specifically for neural networks (NN) and complex event processing (CEP) rules. Subsequently, the enriched semantic knowledge is hosted to discover and interoperate edge devices and applications across decentralized networks using a knowledge graph (KG). The case study of the framework is the implementation of a conveyor belt to monitor operational processes and detect irregularities. An Arduino board connected to a camera uses a CNN to detect the presence of workers near a workstation. The CNN processes the image data to determine whether a worker is present.</p>
<p>A powerful tool for controlling dynamic robotic systems with complex constraints is presented in [<xref ref-type="bibr" rid="ref-104">104</xref>], where a high-speed model-predictive control solver (named TinyMPC) is implanted using the Alternating Direction Method of Multipliers (ADMM) algorithm. The ADMM is used to solve convex optimization problems. The demonstrated applications are high-speed trajectory tracking, dynamic obstacle avoidance, and recovery from extreme attitudes. Moreover, the Teensy 4.1 Development Board is used to benchmark TinyMPC against randomly generated trajectory tracking problems. On the other hand, Crazyflie 2.1 Quadrotor demonstrates TinyMPC&#x2019;s performance in real-time dynamic control tasks, including figure-eight trajectory tracking, recovery from extreme initial attitudes, and dynamic obstacle avoidance.</p>
<p>The work in [<xref ref-type="bibr" rid="ref-105">105</xref>] presents a transfer learning framework combined with TinyML-powered CNN architecture for vibration-based fault diagnosis of different machines. Various time-domain features are extracted from vibration signals for fault diagnosis. Features include mean, median, variance, standard deviation, skew, kurtosis, crest, impulse, and shape factors. While the conventional transfer learning strategy is to retain dense layers while freezing convolutional layers, the work in [<xref ref-type="bibr" rid="ref-19">19</xref>] retains convolutional layers while freezing dense layers. To achieve memory efficiency, it retains only the biases of hidden layers. For edge implementation, the online training is conducted on a Raspberry Pi single-board computer, while the edge inference is performed on an ESP32 microcontroller board using TensorFlow Lite.</p>
<p>Another application of intelligent IoT systems in industry and robotics is ensuring the structural integrity of steel constructions through climbing inspection robots, as presented in [<xref ref-type="bibr" rid="ref-106">106</xref>]. The work introduces a real-time bolt-defect detection system using TinyML and a magnetic climbing inspection robot. The magnetic climbing robot has 3D-printed wheels embedded with permanent magnets for secure adhesion to metallic surfaces. The system employs the Faster Objects, More Objects (FOMO) algorithm optimized for edge computing on microcontrollers. It captures images, processes them using the FOMO model, and streams annotated images to the user via the real-time streaming protocol (RTSP). The FOMO model is simplified for multi-object classification and optimized for microcontrollers. Images of bolts in various conditions (normal, loose, missing) were collected and used to train the model.</p>
<p>TinyML is revolutionizing the industrial and robotics sectors through various innovative applications. The TinyOL framework enhances anomaly detection, and W3C WoT standardizes IoT device capabilities, enabling efficient monitoring and worker detection. Similarly, TinyMPC, a high-speed control solver, improves dynamic robotic tasks like trajectory tracking and obstacle avoidance. Lastly, climbing inspection robots utilize TinyML and the FOMO algorithm for real-time bolt-defect detection, demonstrating effective multi-object classification on microcontrollers. These advancements highlight TinyML&#x2019;s potential to improve industrial and robotic systems significantly.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Energy Sector</title>
<p>Integrating digital technologies is important for managing energy resources efficiently and sustainably, especially when updating old systems. It involves using sensors and devices to monitor and control energy accurately. The growing reliance on renewable energy requires efficient maintenance of large-scale photovoltaic (PV) solar plants. The faults in PV panels, such as hot spots, can significantly reduce efficiency, and therefore, early detection of faults is critical through some monitoring mechanisms.</p>
<p>The work in [<xref ref-type="bibr" rid="ref-107">107</xref>] presents a framework to classify defects in PV modules. The framework captures infrared (IR) images using a low-cost IR camera and processes these images in real-time. The images are pre-processed to reduce noise and then fed into the Tiny CNN model for classification. The model is optimized and integrated into a low-cost, low-power microcontroller for real-time fault diagnosis. The trained TensorFlow TinyCNN model is converted to a TensorFlow Lite version to optimize it for deployment on a microcontroller. The Arduino-integrated development environment is used for writing and uploading the C&#x002B;&#x002B; code to microcontrollers, while the OpenCV-Python library preprocesses IR images. Another microcontroller (Arduino Nano 33 IoT) is used to communicate with the Tiny CNN microcontroller and post the results online. This setup enables remote monitoring and real-time visualization of the PV module status on a dedicated webpage.</p>
<p>The work in [<xref ref-type="bibr" rid="ref-108">108</xref>] employs thermographic images of a PV array for fault diagnosis using two deep convolutional neural network (DCNN) models. A binary classifier model architecture detects whether a PV module is faulty. Moreover, a multiclass classifier is also used to diagnose the specific type of fault in a PV module. The models are trained using TensorFlow and Keras libraries. The trained models are converted to TensorFlow Lite format to reduce their size and make them suitable for edge-devices deployment. The optimized models are embedded into a Raspberry Pi 4 microprocessor. Python scripts run the models, send notifications via SMS and email, and display results on an LCD. The work in [<xref ref-type="bibr" rid="ref-109">109</xref>] started by making a solar farm dataset, including power generation and weather-related information. Then, this data is preprocessed using min-max scaling. Various machine learning models are trained and evaluated for their performance in predicting solar energy yield, focusing on tuning hyperparameters to optimize accuracy. Additionally, the study explores the deployment of these models on resource-constrained edge devices using TinyML to enable real-time, low-cost forecasting.</p>
<p>The objective of the work in [<xref ref-type="bibr" rid="ref-110">110</xref>] is to develop a privacy-preserving architectural framework for hybrid energy management systems (HEMS) that leverages TinyML models for short-term energy consumption prediction on mobile devices. This approach aims to enhance user privacy, ensure efficient energy management, and provide real-time, on-device processing capabilities. It employs LSTM neural networks to predict short-term energy consumption in hybrid energy management systems. These models are converted to CoreML and TensorFlow Lite formats to enable deployment on mobile devices. TensorFlow Lite shows minimal performance degradation and thus is more suitable for real-time, on-device processing. The hardware used in the study includes mobile devices, specifically the iPhone 12 Pro, which serves as an edge device for processing and storing data locally. Finally, the work in [<xref ref-type="bibr" rid="ref-111">111</xref>] presents a framework to upgrade the energy distribution panel of an old manufacturing plant by installing sensor devices. These sensors enable remote monitoring and decentralize predictive analysis. The analysis uses 15-min energy demand forecast models based on two-layer LSTM networks.</p>
<p>The key findings from the aforementioned works highlight the innovative use of TinyML in various applications. A framework for classifying PV module defects uses IR images processed by Tiny CNN models for real-time fault diagnosis. Thermographic images of PV arrays are used with DCNN models for fault detection and classification. A privacy-preserving framework for hybrid energy management systems uses LSTM models on mobile devices to predict short-term energy consumption, ensuring efficient energy management. Lastly, sensor devices installed in an old manufacturing plant enable remote monitoring and predictive analysis using LSTM networks, improving energy management. These findings demonstrate TinyML&#x2019;s potential in enhancing real-time monitoring, fault diagnosis, and energy management across various sectors.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Consumer Electronics and Safety in Smart Cities</title>
<p>All the above categories (smart agriculture, smart healthcare, smart vehicles, smart industrial processes, and smart energy management) have significantly enhanced our living standards. All these smart processes have given birth to &#x201C;smart cities&#x201D;. The concept of smart cities also includes consumer electronics in smart homes and the security of all the automated processes. This section will review some state-of-the-art consumer electronics in smart homes and the security of some automated processes in smart homes. Consumer electronics in smart homes include devices that can be controlled remotely via smartphones or voice assistants. Security in smart cities includes surveillance systems, emergency response systems, advanced access control systems, cybersecurity, etc. Existing review articles, such as [<xref ref-type="bibr" rid="ref-118">118</xref>], review the technical advancements in consumer electronics and smart cities. However, they purely discuss smart cities in terms of IoT, machine learning, cloud computing, and edge computing. This section reviews some state-of-the-art IIoT in consumer electronics and security.</p>
<p>The work in [<xref ref-type="bibr" rid="ref-112">112</xref>] enhances the detection of jamming attacks in IoT wireless networks using a CNN-based approach deployed on edge devices. The model is trained to classify two types of jamming attacks (constant and periodic). Received Signal Strength (RSS) data is the primary metric for detecting and classifying jamming attacks. The data is collected for different scenarios, including normal channel conditions and two types of jamming attacks. Finally, the trained model is deployed using TensorFlow Lite on edge devices like the Raspberry Pi. The deployed model performs real-time inference on RSS data, detecting and classifying jamming attacks. Another CNN-based approach is presented in [<xref ref-type="bibr" rid="ref-113">113</xref>], where the objective is to classify the usage of handheld power tools. The model was trained and validated using a dataset of over 280 min of three-axis accelerations during different activities (tool transportation, no-load, metal drilling, and wood drilling). The trained CNN model is converted to TensorFlow Lite format using post-training quantization (to convert the 32-bit floating-point weights to 8-bit integers), reducing the model size significantly. The converted model is deployed on the NRF52832 board.</p>
<p>The work in [<xref ref-type="bibr" rid="ref-114">114</xref>] presents a framework (TinyAP) for detecting and mitigating attacks at the access point level in a Wi-Fi network. The objective is to ensure the safety and privacy of smart home devices. It employs DNN and LSTM, trained on a general-purpose computer, and then converted into a format suitable for deployment on microcontrollers. While the work in [<xref ref-type="bibr" rid="ref-114">114</xref>] focuses on the security of smart homes, the security of medical IoT systems is discussed in [<xref ref-type="bibr" rid="ref-115">115</xref>], enabling real-time health monitoring and intrusion detection. It addresses privacy and security concerns using edge processing for data encryption and TinyML for real-time vital sign analysis and emergency alerts. Moreover, Naive Bayes and SVM are used for real-time analysis of vital signs. However, the article does not mention the type of hardware used for monitoring and processing real-time vital signs (e.g., blood pressure and heart rate). The framework&#x2019;s encryption and intrusion detection capabilities help maintain data integrity and privacy.</p>
<p>Finally, the work in [<xref ref-type="bibr" rid="ref-116">116</xref>] presents an energy-harvesting resource allocation algorithm for UAV-assisted TinyML consumer electronics in low-power IoT networks. Optimizing energy harvesting and resource allocation in IoT networks involves handling interference, resource conflicts, and real-time decision-making. Therefore, Naive Bayes, SVM, RF, and KNN are efficient for classification and regression tasks in resource-constrained environments. However, the article did not provide the corresponding hardware and software details for the model deployment.</p>
<p>To summarize, the key applications in smart home security and consumer electronics are the detection of jamming attacks in IoT wireless networks, the usage of handheld power tools, detecting and mitigating attacks at the access point level in Wi-Fi networks, the security of medical IoT systems and energy-harvesting resource allocation algorithm for UAV-assisted TinyML consumer electronics in low-power IoT networks, optimizing energy use and resource allocation. These findings demonstrate the potential of TinyML and edge computing to enhance the security and efficiency of smart cities.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Responses to Formulated Research Question</title>
<p>This section formulates responses to the target research questions (identified in the Introduction Section). The responses to research questions are based on the results of <xref ref-type="sec" rid="s3">Sections 3</xref> and <xref ref-type="sec" rid="s4">4</xref>.</p>
<p><bold>Research Question 1:</bold> What are the state-of-the-art model conversion techniques used in EdgeML, and how do they impact the performance, efficiency, and deployment of machine learning models on resource-constrained devices?</p>
<p><bold>Answer:</bold> This work has classified state-of-the-art model conversion techniques in EdgeML into four categories: Pruning, quantization, low-rank factorization, and knowledge distillation. Their impact on performance, efficiency, and model deployment has been discussed in <xref ref-type="sec" rid="s3">Section 3</xref>. The pruning methods have been compared in <xref ref-type="table" rid="table-5">Table 5</xref>. Results on quantization methods and quantization-aware training have been summarized in <xref ref-type="table" rid="table-6">Tables 6</xref> and <xref ref-type="table" rid="table-7">7</xref>, respectively. Similarly, model conversion methods based on low-rank factorization and knowledge distillation methods have been analyzed in <xref ref-type="table" rid="table-8">Tables 8</xref> and <xref ref-type="table" rid="table-9">9</xref>, respectively.</p>

<p><bold>Research Question 2:</bold> What are the current state-of-the-art inference mechanisms used in EdgeML, and how do they compare in performance and efficiency?</p>
<p><bold>Answer:</bold> Edge devices can perform inference in two ways once a model is deployed: either entirely on the device (on-device inference) or in collaboration with other edge devices or servers (distributed inference). <xref ref-type="sec" rid="s3_2_1">Sections 3.2.1</xref> and <xref ref-type="sec" rid="s3_2_2">3.2.2</xref> detail these methods, and <xref ref-type="table" rid="table-10">Table 10</xref> summarizes the different inference mechanisms.</p>

<p><bold>Research Question 3:</bold> How do different learning strategies impact the performance and efficiency of ML models deployed in edge computing environments?</p>
<p><bold>Answer:</bold> The results of our study on different learning strategies are categorized into three main areas: (a) online learning strategies, (b) continual or life-long learning, and (c) federated learning. <xref ref-type="sec" rid="s3_3_1">Sections 3.3.1</xref>&#x2013;<xref ref-type="sec" rid="s3_3_3">3.3.3</xref> present the findings for each area, respectively. Additionally, <xref ref-type="table" rid="table-11">Table 11</xref> summarizes the results for continual learning, while <xref ref-type="table" rid="table-12">Table 12</xref> summarizes federated learning. A comprehensive analysis of selected research areas in these sections and tables highlights how various learning strategies impact the performance and efficiency of ML models deployed in edge computing environments.</p>

<p><bold>Research Question 4:</bold> What are the key challenges and problems and the associated ML/DL models in deploying TinyML on resource-constrained devices in various sectors?</p>
<p><bold>Answer:</bold> <xref ref-type="sec" rid="s4">Section 4</xref> presents the results of deploying TinyML models on resource-constrained devices organized in different sectors to reflect TinyML&#x2019;s diverse applications. Consequently, we have identified six key sectors where TinyML-based model deployment has shown promising results such that <xref ref-type="sec" rid="s4_1">Sections 4.1</xref>&#x2013;<xref ref-type="sec" rid="s4_6">4.6</xref> provide critical analysis of the selected research studies in smart agriculture, healthcare, vehicles, industry, energy, and security sectors, respectively. Moreover, <xref ref-type="table" rid="table-13">Table 13</xref> lists various applications and the associated ML/DL models used in each sector.</p>

<p><bold>Research Question 5:</bold> What are the latest advancements in hardware, software frameworks, and sensors designed explicitly for TinyML frameworks?</p>
<p><bold>Answer:</bold> The synthesized results for six different TinyML sectors are detailed in <xref ref-type="table" rid="table-14">Tables 14</xref>&#x2013;<xref ref-type="table" rid="table-16">16</xref>. <xref ref-type="table" rid="table-14">Table 14</xref> summarizes 15 hardware development boards from the selected research works, highlighting their features and organizing the articles by hardware development. <xref ref-type="table" rid="table-15">Table 15</xref> categorizes the selected research works by three major software frameworks. Finally, <xref ref-type="table" rid="table-16">Table 16</xref> details the sensors used in all the research papers selected.</p>

</sec>
<sec id="s6">
<label>6</label>
<title>Discussion and Limitations of the Literature Review</title>
<p><xref ref-type="sec" rid="s5">Section 5</xref> has answered the target research questions (listed in the Introduction) by utilizing the results in <xref ref-type="sec" rid="s3">Sections 3</xref> and <xref ref-type="sec" rid="s4">4</xref>. This section further discusses some important aspects of model design and deployment. Mainly, it includes automating quantization and pruning policies for efficient model conversion (<xref ref-type="sec" rid="s6_1">Section 6.1</xref>), an in-depth discussion on fault resilience in distributed edge machine learning models (<xref ref-type="sec" rid="s6_2">Section 6.2</xref>), the influence of hardware constraints on edge devices (<xref ref-type="sec" rid="s6_3">Section 6.3</xref>), ethical concerns of deployment of Edge ML (<xref ref-type="sec" rid="s6_4">Section 6.4</xref>) and limitations of this literature review (<xref ref-type="sec" rid="s6_5">Section 6.5</xref>).</p>
<sec id="s6_1">
<label>6.1</label>
<title>Automating Quantization &#x0026; Pruning Policies for Efficient Model Conversion</title>
<p><xref ref-type="sec" rid="s3">Section 3</xref> reveals that converting models for edge devices involves transforming large, complex models into smaller, efficient versions suitable for resource-constrained hardware. Common techniques include quantization, pruning, low-rank factorization, and knowledge distillation. Quantization, the most widely used technique, can be applied post-training or during training (quantization-aware training). Post-training quantization is easier to implement but may reduce accuracy, while quantization-aware training maintains better accuracy. Pruning, the second most common technique, offers better accuracy but requires more memory and computation. Low-rank factorization and knowledge distillation are less common due to their complexity and longer training times.</p>
<p>Hernandez et al. [<xref ref-type="bibr" rid="ref-119">119</xref>] have provided an interesting theoretical analysis of the generalization of the quantization approaches and studied an algorithm-independent generalization error bound. However, many quantization-aware training algorithms show efficacy through practical applications with lesser theoretical proof of the performance bound of asymptotic convergence analysis. Both quantization-aware training and post-training quantization have advantages and disadvantages. Post-training quantization is easy to implement and does not usually require retraining. So, in situations where we have a trained model and we don&#x2019;t have computational resources and/or time for training the model, post-training quantization is a good option. A quantization policy during the training can maintain higher model performance in quantization-aware training. Moreover, a model can be designed according to the hardware requirements. It will require additional time and computational resources.</p>
<p>In addition to the results in <xref ref-type="sec" rid="s3">Section 3</xref>, automating quantization and pruning policies according to hardware constraints is crucial in model conversion. In [<xref ref-type="bibr" rid="ref-120">120</xref>], the quantization policy is defined as an optimization problem aimed at minimizing cross-entropy loss while adhering to a hardware budget constraint. This constraint limits the number of bits assigned to the model&#x2019;s weight tensors within the hardware budget. On the MNIST dataset, the optimized average bit size of the weights is 1.46, with no accuracy loss, achieving a compression rate of over 2000 on LeNet. Similarly, hardware-aware automated quantization in [<xref ref-type="bibr" rid="ref-121">121</xref>] uses reinforcement learning to optimize policies based on latency and energy feedback. The ResPrune method [<xref ref-type="bibr" rid="ref-122">122</xref>] is another technique that uses stochastic optimization to prune filters, saving over 50% of FLOPS on various datasets. An evolutionary pruning model [<xref ref-type="bibr" rid="ref-123">123</xref>] optimizes pruned models for better classification results. Chen et al. [<xref ref-type="bibr" rid="ref-124">124</xref>] have proposed a two-stage framework for optimizing compression ratios, achieving 3 to 36 times compression with 2 to 3-bit quantization. Guo et al. [<xref ref-type="bibr" rid="ref-125">125</xref>] have introduced a multi-agent reinforcement learning framework for automatic pruning, achieving 30% to 50% compression. Albanese et al. [<xref ref-type="bibr" rid="ref-126">126</xref>] used pruned and quantized MobileNetV2 and SqueezeNet models to detect artifacts in plastic components with high accuracy and inference rates.</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Fault Resilience in Distributed Edge Machine Learning Models</title>
<p>It has been observed from <xref ref-type="sec" rid="s3">Section 3</xref> that inference can be performed entirely on the device (on-device inference) or in collaboration with other devices or servers (distributed inference). For this purpose, EdgeML incorporates various learning strategies to optimize performance and resource usage, including (a) online learning, (b) continual learning, and (c) federated learning. While distributed edge machine learning solutions using model conversion techniques are highly effective, these frameworks are also susceptible to various sources of errors and faults [<xref ref-type="bibr" rid="ref-127">127</xref>]. Therefore, the purpose of fault-resilient edge machine learning models is to maintain the performance and functionality of the model despite various types of errors and faults.</p>
<p>There are many sources of faults and errors in these frameworks [<xref ref-type="bibr" rid="ref-128">128</xref>], such as data-related issues, which include noisy data, incomplete data, and drift in the data distribution used to train the models. Similarly, ensuring redundancy in the inference of edge machine learning models is crucial for maintaining reliability and performance. It implies that if one edge device fails, others can take over its tasks without significant disruption. Communication channels are another critical point where errors can be introduced in edge machine learning systems. During data transfer between edge devices or between edge devices and servers, errors such as packet loss, corruption, or delays can occur. It can lead to incomplete or inaccurate data being processed.</p>
<p>Identifying and diagnosing faults early in the framework can prevent incorrect inferences or system failures. Making the system fault-tolerant involves error correction codes, inherent model design, and decentralized or distributed management [<xref ref-type="bibr" rid="ref-129">129</xref>]. Decentralized management distributes decision-making across multiple nodes, reducing the impact of individual node failures [<xref ref-type="bibr" rid="ref-127">127</xref>]. Deep learning models are programmed and deployed on the hardware, assuming that input data is genuine and error-free, there is no error or bug in the program, and the hardware is provided as described without any pre-deployment or post-deployment fault. In reality, the input data may be corrupt or noisy due to failure or malfunction of sensors or adversarial attacks [<xref ref-type="bibr" rid="ref-130">130</xref>]. Programs may have bugs or data-oriented wrong calculations, and hardware may face failures due to environmental or structural effects. The fault resilience of the deep learning models should be analyzed against these types of faults, reflecting the models&#x2019; reliability under such circumstances. A systematic review of designing a framework for fault injection can be found in [<xref ref-type="bibr" rid="ref-131">131</xref>]. In the following, we discuss some research works focusing on fault injection.</p>
<p>Narayanan et al. [<xref ref-type="bibr" rid="ref-132">132</xref>] have proposed high-level fault injection frameworks for TensorFlow-based applications. Similarly, Laster et al. [<xref ref-type="bibr" rid="ref-133">133</xref>] have studied the effect of transient hardware faults on the misclassification of DNN. Syed et al. [<xref ref-type="bibr" rid="ref-134">134</xref>] investigated the impact of faults on the deep neural network at different quantization levels. They have found that good quantization of the model increases the resiliency of the DNN. Ruospo et al. [<xref ref-type="bibr" rid="ref-135">135</xref>] have investigated the effect of fixed and floating point quantization of CNN on reliability. An open-source fault injection framework, darknet, tests the CNN resilience against the faults. They have tested LeNet and YOLO architectures by injecting faults at different layers. They have concluded that fixed point data quantization provides a better tradeoff between memory footprint reduction and resilience.</p>
<p>Liu et al. [<xref ref-type="bibr" rid="ref-136">136</xref>] have proposed a distribution-based error detector to improve the bit error resilience of DNN. They have used memory errors and register fault injectors. The results on LeNet and AlexNet showed that DED improved the error resilience of the models. The effect of the two pruning methods, namely magnitude-based pruning and filter-based structured pruning, on the reliability of the DNN deployed on FPGA is studied in [<xref ref-type="bibr" rid="ref-137">137</xref>]. A hardware injection tool is used for reliability evaluations.</p>
<p>Furthermore, the effect of weight quantization is also investigated. The classification accuracy of VGG16 on the CIFAR-10 dataset is studied for different quantization levels and pruning methods. They have found that 8-bit quantization is a better option, and reliability does not increase much as we increase the bit-width to be larger than 8. Filter-based structured pruning method is less reliable than magnitude-based pruning. DNN with higher pruning rates is more robust to weight errors but less reliable to errors on the configuration bits.</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Influence of Hardware Constraints on Edge Devices</title>
<p>The specific hardware constraints of edge devices significantly influence the design and deployment of models in EdgeML and TinyML.</p>
<p>Memory Constraints: Edge devices often have limited memory, necessitating model optimization techniques to reduce the size of machine learning models. Techniques such as pruning, quantization, and knowledge distillation are commonly used to compress models without significantly compromising accuracy. For instance, pruning removes unnecessary parameters, while quantization reduces the precision of the model weights, both of which help in fitting the model within the limited memory available on edge devices.</p>
<p>Computational Capabilities: The computational power of edge devices varies widely. Devices with limited computational capabilities require computationally efficient models. It often involves designing lightweight models or using specialized architectures like MobileNets or SqueezeNet, optimized for low-power and low-latency inference. Additionally, techniques such as model partitioning can distribute the computational load across multiple devices or offload some tasks to more powerful servers, balancing the computational requirements.</p>
<p>Energy Consumption: Energy efficiency is crucial for edge devices, especially those that rely on battery power. Models deployed on these devices must be optimized to minimize energy consumption. It can be achieved through techniques like event-driven computing, where the device only activates when necessary, and by using energy-efficient algorithms that reduce the number of computations required. For example, event-driven computing can significantly extend battery life by avoiding constant data processing.</p>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Ethical Concerns of Deployment of Edge ML</title>
<p>When machine learning models are deployed on resource-constrained devices such as smartphones, smartwatches, and other IoT devices, handling the privacy issue of sensitive personal data is a significant concern that needs to be addressed. Robust security infrastructure is required to optimize the robustness of the IoT devices against data breaches. This issue has been discussed in many ways, including with the newly developed blockchain technology. Still, incorporating robust security measures in resource-constrained devices is difficult due to limited power availability. The learning of the models with limited access to the training data from a localized environment may create a bias in the models. Such a problem can be handled by sharing the models among the edge devices and placing the central model on the cloud or local servers. Many decisions are made by edge machine learning models in sensitive environments, such as maintaining security or providing healthcare in smart cities. In such cases, transparency and accountability of automated decision-making are more significant ethical concerns.</p>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Limitations of the Literature Review</title>
<p>Despite strictly following the guidelines and adhering to our review protocol, there are certain limitations: (1) While we used appropriate search terms and thoroughly scanned the results, some terms returned thousands of results that could not be exhaustively reviewed. Additionally, some research was rejected based on titles, which may not accurately reflect the content. Therefore, we do not claim our research is exhaustive. (2) We utilized four renowned scientific databases: IEEE, ELSEVIER, ACM, and SPRINGER, which provide many journal and conference publications. However, other databases also contain significant research work. Consequently, there is a chance we missed relevant recent studies from other sources. Nonetheless, we believe the ultimate findings of this literature review are not significantly affected, as the selected databases provide high-quality, up-to-date research literature.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Challenges and Future Directions</title>
<p><xref ref-type="sec" rid="s3">Sections 3</xref> and <xref ref-type="sec" rid="s4">4</xref> provide results into two major categories: (1) model conversion, inference, and learning in EdgeML and (2) model deployment on resource-constrained devices with TinyML. Based on these results, <xref ref-type="sec" rid="s5">Sections 5</xref> and <xref ref-type="sec" rid="s6">6</xref> have formulated responses to identified research questions and associated discussion. Consequently, this section identifies probable issues due to the rapid increase in smart devices, which are projected to reach 29 billion by 2030 [<xref ref-type="bibr" rid="ref-138">138</xref>]. These devices generate vast amounts of data, which can be utilized to build numerous IIoT systems. However, deploying these models on edge devices has several challenges due to their computational and energy constraints. Here are some key challenges and future research directions:</p>
<p><bold>Resource Optimization for Dynamic and Heterogeneous Edge Environments:</bold> The edge environment is becoming increasingly dynamic, heterogeneous, and diverse. Edge devices now operate in various environments, from remote, communication-constrained areas to highly demanding and dangerous situations such as firefighting, battlefields, and harsh weather conditions [<xref ref-type="bibr" rid="ref-139">139</xref>]. Adaptive solutions are needed to dynamically allocate resources based on decision-making requirements, current workload, and communication status among edge devices or between edge devices and edge or cloud servers [<xref ref-type="bibr" rid="ref-140">140</xref>]. Therefore, effective resource management is crucial for end-to-end machine learning applications deployed on autonomous devices. For battery-operated edge devices, extending operational life is also a key concern. Consequently, efficient and innovative computational algorithms are becoming a priority. Developing multi-objective optimal policies for sharing and balancing resources among collaborating edge devices is another important future research direction [<xref ref-type="bibr" rid="ref-141">141</xref>].</p>
<p>Additionally, optimizing the placement of edge nodes and servers is essential to enhance the quality of collaborative tasks for mobile edge devices [<xref ref-type="bibr" rid="ref-142">142</xref>]. We need a suitable prediction mechanism to predict the workload patterns and allocate the resources in the dynamic environment. Resource scheduling requires a robust mechanism in the presence of communication instability and variation of available resources. Reinforcement learning can be a good future direction in solving this problem. In the heterogeneous environment, scalability issues are a significant concern for resource management strategies, which can be handled by solving dynamic scheduling issues.</p>
<p><bold>Resource-Aware Learning and Sustainability:</bold> Although models can be compressed to fit on-edge devices, inference, and continual learning (as discussed in <xref ref-type="sec" rid="s3">Section 3</xref> of this article) still demand additional resources. It implies that maintaining learned knowledge and acquiring new information from dynamic environments is challenging. Analyzing data, selecting important information for learning, and discarding irrelevant data require extra computational resources and energy. When edge devices collaborate or communicate with servers, designing an optimal policy for inference and learning is always challenging [<xref ref-type="bibr" rid="ref-143">143</xref>]. Therefore, balancing acquiring new knowledge with making the model available for inference necessitates task-based scheduling procedures. Moreover, ensuring timeliness and accuracy through specialized model design will be crucial for future edge ML applications [<xref ref-type="bibr" rid="ref-144">144</xref>]. Consequently, resource sustainability will be a significant research topic, as resource-aware federated learning uses neural architecture search to generate deployable models on heterogeneous devices based on their resources. Yu et al. [<xref ref-type="bibr" rid="ref-145">145</xref>] proposed on-demand customized model deployment considering the available resources on edge devices. Extensive research is also required in this direction. Selecting the data relevant to the learning already trained is a primary step in efficient learning without wasting computation resources on learning redundant data. Ge et al. [<xref ref-type="bibr" rid="ref-146">146</xref>] proposed an adaptive personalized FL when the data distribution among the users is homogenous. Hence, the scheme identifies users with homogenous patterns and engages them in collaborative FL using one-shot screening based on the learning loss without communicating the original data. Due to privacy issues related to personal data, it is not safe to transmit this data to the edge server. Hence, specific properties of the dataset on which the machine learning model is trained should be kept and used to predict the novelty of the new information.</p>
<p><bold>New Learning Paradigms and Algorithms:</bold> While this literature review has discussed various learning paradigms, including on-device, federated, continual, or life-long, and split model learning, adapting to changing data distributions remains a significant challenge and requires stable learning procedures. The review highlights that substantial work has been done on quantizing and pruning pre-trained models. However, there is still a need for extensive research in quantization-aware and prune-aware training of models. Specialized learning algorithms for quantized models based on multi-objective loss functions could be more beneficial for future model designs [<xref ref-type="bibr" rid="ref-147">147</xref>]. Due to its limited computational resources, TinyML faces additional challenges in incorporating these learning strategies effectively. Addressing these challenges will be crucial for advancing the capabilities and applications of TinyML [<xref ref-type="bibr" rid="ref-148">148</xref>].</p>
<p><bold>Hardware-Oriented Model Design:</bold> As EdgeML and TinyML technologies evolve, optimizing hardware to meet the specific needs of various applications and algorithms becomes increasingly important [<xref ref-type="bibr" rid="ref-149">149</xref>]. Custom hardware can be designed to accelerate particular machine learning algorithms, improving performance and efficiency. Tailoring hardware to the needs of specific applications can significantly reduce power consumption. By designing hardware optimized for specific tasks, it is possible to reduce latency, enabling real-time data processing and decision-making. However, creating hardware optimized for particular applications can limit its scalability across different use cases, requiring multiple designs for various applications. Moreover, developing custom hardware can be expensive and time-consuming, which may not be feasible for all organizations. In the hardware-oriented model design, hardware constraints can be identified, and a multi-objective optimization algorithm can be used to achieve a tradeoff between model complexity and hardware design.</p>
<p><bold>Better Interoperability and Scalability:</bold> <xref ref-type="sec" rid="s4">Section 4</xref> of this LITERATURE REVIEW has highlighted that the EdgeML and TinyML paradigms are applied to diverse applications (such as healthcare, automotive, energy, vehicles, industry, and energy) in a dynamic and varied environment. The diversity of edge devices, each with unique constraints, adds complexity to EdgeML and TinyML framework [<xref ref-type="bibr" rid="ref-11">11</xref>]. In collaborative inference and learning, these edge devices must work together to complete tasks despite their hardware and software variability. This heterogeneity necessitates standard protocols for integrating information and learning. Recent research [<xref ref-type="bibr" rid="ref-150">150</xref>&#x2013;<xref ref-type="bibr" rid="ref-152">152</xref>] has addressed some interoperability issues, but much more work is needed in this area. Scalability is another significant concern in deploying edge computing solutions. Computation offloading to edge servers or other edge devices can improve performance, save energy, and reduce computation time. However, the capacity of edge servers to handle requests from numerous edge devices can constrain performance.</p>
<p>Additionally, the heterogeneity of edge machine learning architectures can lead to scalability challenges. Interoperability can be improved by ensuring the compatibility of the communication protocols among the devices. Hence, the standardization of communication protocols for heterogeneous devices is desirable for future practical applications. Making smart multi-protocol gateways can be a future solution to this problem.</p>
<p><bold>New Computational Algorithms Aiming for Edge Devices:</bold> New computational algorithms are required to efficiently deploy and compute machine learning models and are well-suited for limited-resource edge devices. Hyperdimensional computing (HDC) is a new learning paradigm suitable for lesser computation and energy requirements [<xref ref-type="bibr" rid="ref-153">153</xref>]. Hence, efficient machine learning solutions can be developed in the high-dimensional space with lower precision parameters. Moreover, implementing the HDC on diverse hardware platforms is challenging and needs further research and development [<xref ref-type="bibr" rid="ref-154">154</xref>].</p>
<p><bold>New Edge AI Architectures Combining Cloud, Fog, and Edge Layers:</bold> New edge AI (artificial intelligence) architectures are needed to accommodate the diverse range of edge devices and servers, optimizing synchronization and communication within limited bandwidth constraints. Typically, EdgeML architectures are multi-tiered, comprising multiple computing layers, from edge devices to fog servers and cloud servers. Efficient resource utilization is challenging due to the varying levels of granularity across these layers [<xref ref-type="bibr" rid="ref-155">155</xref>]. Multi-tenant edge AI allows multiple applications to run concurrently, optimizing resource use. However, this architecture faces key challenges in meeting performance requirements such as latency, accuracy, and energy consumption [<xref ref-type="bibr" rid="ref-156">156</xref>].</p>
<p><bold>Data Privacy and Access Security:</bold> Edge ML is decentralized in computing and contains various edge devices and servers. Therefore, the vulnerability of edge devices is a big issue. Due to their distributed nature and operation in an untrusted environment, edge devices are prone to various cyber-attacks. Hence, protecting the sensitive data collected by edge devices and stored locally is a significant challenge. Securing the transmission of the data over multiple communication channels is also a challenging task. Moreover, maintaining a dynamic edge ML environment where edge devices may leave or join a dynamic trust model between devices is an issue to tackle [<xref ref-type="bibr" rid="ref-157">157</xref>]. Conventional cryptographical solutions are no longer safe after the emergence of quantum computing. So, in edge machine learning, creating quantum-safe protocols will be a future challenge [<xref ref-type="bibr" rid="ref-158">158</xref>].</p>
<p><bold>Operational Efficiency:</bold> While edge devices collect data for inference and learning, misleading data may corrupt the learning by edge devices. So, avoiding the adversarial data and recognizing the legitimate data for learning and inference is also challenging [<xref ref-type="bibr" rid="ref-159">159</xref>]. Adversarial robustness in a machine learning model can be achieved by defining new adversarial learning strategies that can incorporate adversarial examples in the learning process. In this way, the trained models can recognize and resist adversarial attacks on the continual learning process. Model hardening is another helpful research direction that includes data sanitization, adversarial training, and differential privacy to mitigate the threats. A good survey in this direction has recently been published [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p><bold>Edge ML on Quantum Edge Devices:</bold> Integrating EdgeML and TinyML with quantum edge devices is another promising future research direction that aims to utilize the advantages of quantum computing, such as enhanced processing power and speed, to optimize further machine learning models deployed on edge devices [<xref ref-type="bibr" rid="ref-160">160</xref>]. In addition to the ability of quantum computing to solve complex problems more efficiently than classical computing, post-quantum cryptography algorithms offer advanced security features that could protect data processed on edge devices [<xref ref-type="bibr" rid="ref-161">161</xref>]. However, it may require developing quantum algorithms tailored for EdgeML and TinyML applications. Developing specific quantum algorithms for EdgeML and TinyML can unlock new possibilities for real-time data processing and analytics, making these technologies even more powerful and versatile.</p>
<p><bold>Photonic Neural Networks on the Edge Devices:</bold> Photonic neural networks use light for computations and leverage important characteristics of light in terms of low energy consumption, parallel implementation, and faster processing speed. Hence, photonic deep neural networks can be a good choice regarding efficiency and performance [<xref ref-type="bibr" rid="ref-162">162</xref>]. However, coherent optical processing is challenging in photonic neural networks [<xref ref-type="bibr" rid="ref-163">163</xref>]. Moreover, deploying the photonic neural network on resource-constrained devices still requires further research in this area, with a major focus on miniaturization, cost efficiency, and integration with other components on edge devices.</p>
<p><bold>Compliance with Regulations and Standards:</bold> As EdgeML and TinyML technologies become more integrated into various applications, they must meet regulatory requirements and industry standards for widespread adoption and safe deployment. Edge devices often handle sensitive data, so compliance with data protection regulations like GDPR or CCPA is essential [<xref ref-type="bibr" rid="ref-164">164</xref>]. Similarly, meeting safety standards is vital for applications in critical sectors such as healthcare or automotive. In addition to security and safety standards, interoperability is also important. Furthermore, developing standards that ensure different devices and systems can work together seamlessly is important for the scalability of EdgeML and TinyML solutions [<xref ref-type="bibr" rid="ref-165">165</xref>]. Finally, compliance with environmental regulations regarding energy consumption and electronic waste is becoming increasingly important as the number of deployed devices grows.</p>
</sec>
<sec id="s8">
<label>8</label>
<title>Conclusion</title>
<p>This literature review has highlighted significant advancements in model conversion, inference mechanisms, and learning strategies in EdgeML, emphasizing their impact on the performance, efficiency, and deployment of machine learning models on resource-constrained devices with TinyML. It identifies and discusses various techniques, such as pruning, quantization, knowledge distillation, and low-rank factorization, along with their respective advantages and challenges. The document also explores different inference mechanisms, including on-device and distributed inference, and various learning strategies like continual learning and federated learning. Furthermore, it delves into deploying TinyML models across different sectors, including smart agriculture, healthcare, automotive, industry, energy, and security. It provides insights into the hardware development boards, software frameworks, and sensors used in these applications, showcasing the versatility and potential of TinyML in real-world scenarios. Despite the significant progress, the review acknowledges several challenges and future research directions. These include resource optimization, new learning paradigms, interoperability, scalability, security, and the integration of quantum computing with EdgeML and TinyML. Addressing these challenges will be crucial for the continued advancement and widespread adoption of these technologies.</p>
<p>In conclusion, this review is a valuable resource for researchers, practitioners, and developers, offering a detailed overview of the current landscape and future directions in EdgeML and TinyML. Practitioners can leverage techniques like pruning, quantization, and knowledge distillation to reduce model size and improve inference speed, making deploying complex models on resource-constrained devices feasible. Similarly, developers can choose between on-device and distributed inference based on the specific application requirements, balancing latency, computational load, and energy consumption. Moreover, continual and federated learning can help maintain model accuracy over time without frequent retraining, particularly useful in dynamic environments. The insights into hardware development boards, software frameworks, and sensors can guide practitioners in selecting the right tools and components for deploying TinyML models in various sectors such as smart agriculture, healthcare, automotive, industry, energy, and security. By incorporating these practical implications, practitioners and developers can effectively apply the advancements discussed in the review to their projects, driving innovation and improving the deployment of machine learning models on resource-constrained devices.</p>
</sec>
</body>
<back>
<ack>
<p>The authors acknowledge the Computers, Materials &#x0026; Continua for their support for the paper.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm their contribution to the paper as follows: study conception and design: Muhammad Arif and Muhammad Rashid; data collection: Muhammad Arif; analysis and interpretation of results: Muhammad Arif and Muhammad Rashid; draft manuscript preparation: Muhammad Arif and Muhammad Rashid. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>No dataset is used in the paper.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmed</surname> <given-names>SF</given-names></string-name>, <string-name><surname>Alam</surname> <given-names>MSB</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rozbu</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Ishtiak</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rafa</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning modelling techniques: current progress, applications, advantages, and challenges</article-title>. <source>Artif Intell Rev</source>. <year>2023</year>;<volume>56</volume>(<issue>11</issue>):<fpage>13521</fpage>&#x2013;<lpage>617</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-023-10466-8</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Oliveira</surname> <given-names>F</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>DG</given-names></string-name>, <string-name><surname>Assis</surname> <given-names>F</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Internet of intelligent things: a convergence of embedded systems, edge computing and machine learning</article-title>. <source>Internet Things</source>. <year>2024</year>;<volume>26</volume>(<issue>9</issue>):<fpage>101153</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2024.101153</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lesch</surname> <given-names>V</given-names></string-name>, <string-name><surname>Z&#x00FC;fle</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bauer</surname> <given-names>A</given-names></string-name>, <string-name><surname>Iffl&#x00E4;nder</surname> <given-names>L</given-names></string-name>, <string-name><surname>Krupitzer</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kounev</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A literature review of IoT and CPS&#x2014;what they are, and what they are not</article-title>. <source>J Syst Softw</source>. <year>2023</year>;<volume>200</volume>(<issue>3</issue>):<fpage>111631</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jss.2023.111631</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Parast</surname> <given-names>FK</given-names></string-name>, <string-name><surname>Sindhav</surname> <given-names>C</given-names></string-name>, <string-name><surname>Nikam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yekta</surname> <given-names>HI</given-names></string-name>, <string-name><surname>Kent</surname> <given-names>KB</given-names></string-name>, <string-name><surname>Hakak</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Cloud computing security: a survey of service-based models</article-title>. <source>Comput Secur</source>. <year>2022</year>;<volume>114</volume>(<issue>1</issue>):<fpage>102580</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2021.102580</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gill</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>Edge AI: a survey</article-title>. <source>Internet Things Cyber-Phys Syst</source>. <year>2023</year>;<volume>3</volume>(<issue>5</issue>):<fpage>71</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iotcps.2023.02.004</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duan</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Combining federated learning and edge computing toward ubiquitous intelligence in 6G network: challenges, recent advances, and future directions</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2023</year>;<volume>25</volume>(<issue>9</issue>):<fpage>2892</fpage>&#x2013;<lpage>950</lpage>. doi:<pub-id pub-id-type="doi">10.1109/COMST.2023.3316615</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ba&#x015F;ar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Diggavi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Eldar</surname> <given-names>YC</given-names></string-name>, <string-name><surname>Letaief</surname> <given-names>KB</given-names></string-name>, <string-name><surname>Poor</surname> <given-names>HV</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Communication-efficient distributed learning: an overview</article-title>. <source>IEEE J Sel Areas Commun</source>. <year>2023</year>;<volume>41</volume>(<issue>4</issue>):<fpage>851</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSAC.2023.3242710</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>An</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Edge computing on IoT for machine signal processing and fault diagnosis: a review</article-title>. <source>IEEE Internet Things J</source>. <year>2023</year>;<volume>10</volume>(<issue>13</issue>):<fpage>11093</fpage>&#x2013;<lpage>116</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2023.3239944</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Immonen</surname> <given-names>R</given-names></string-name>, <string-name><surname>H&#x00E4;m&#x00E4;l&#x00E4;inen</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Tiny machine learning for resource-constrained microcontrollers</article-title>. <source>J Sens</source>. <year>2022</year>;<volume>2022</volume>(<issue>1</issue>):<fpage>7437023</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2022/7437023</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tsoukas</surname> <given-names>V</given-names></string-name>, <string-name><surname>Gkogkidis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Boumpa</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kakarountas</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A review on the emerging technology of TinyML</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>10</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3661820</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Meuser</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lov&#x00E9;n</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bhuyan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Patil</surname> <given-names>SG</given-names></string-name>, <string-name><surname>Dustdar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aral</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Revisiting edge AI: opportunities and challenges</article-title>. <source>IEEE Internet Comput</source>. <year>2024</year>;<volume>28</volume>(<issue>4</issue>):<fpage>49</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MIC.2024.3383758</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hua</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Edge computing with artificial intelligence: a machine learning perspective</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3555802</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hoffpauir</surname> <given-names>K</given-names></string-name>, <string-name><surname>Simmons</surname> <given-names>J</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pittala</surname> <given-names>R</given-names></string-name>, <string-name><surname>Briggs</surname> <given-names>I</given-names></string-name>, <string-name><surname>Makani</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on edge intelligence and lightweight machine learning support for future applications and services</article-title>. <source>ACM J Data Inform Qual</source>. <year>2023</year>;<volume>15</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3581759</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhalgaonkar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Munot</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Model compression of deep neural network architectures for visual pattern recognition: current status and future directions</article-title>. <source>Comput Elect Eng</source>. <year>2024</year>;<volume>116</volume>(<issue>3</issue>):<fpage>109180</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2024.109180</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grzesik</surname> <given-names>P</given-names></string-name>, <string-name><surname>Mrozek</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Combining machine learning and edge computing: opportunities, challenges, platforms, frameworks, and use cases</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>3</issue>):<fpage>640</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13030640</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Capogrosso</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cunico</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>DS</given-names></string-name>, <string-name><surname>Fummi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cristani</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A machine learning-oriented survey on tiny machine learning</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>1</issue>):<fpage>23406</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3365349</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Paracha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Arshad</surname> <given-names>J</given-names></string-name>, <string-name><surname>Farah</surname> <given-names>MB</given-names></string-name>, <string-name><surname>Ismail</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Machine learning security and privacy: a review of threats and countermeasures</article-title>. <source>EURASIP J Inf Secur</source>. <year>2024</year>;<volume>2024</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1186/s13635-024-00158-3</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Maybank</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Knowledge distillation: a survey</article-title>. <source>Int J Comput Vis</source>. <year>2021</year>;<volume>129</volume>(<issue>6</issue>):<fpage>1789</fpage>&#x2013;<lpage>819</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-021-01453-z</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Barijough</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Gerstlauer</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Deepthings: distributed adaptive deep learning inference on resource-constrained iot edge clusters</article-title>. <source>IEEE Trans Comput Aided Des Integr Circuits Syst</source>. <year>2018</year>;<volume>37</volume>(<issue>11</issue>):<fpage>2348</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCAD.2018.2858384</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kone&#x010D;n&#x00FD;</surname> <given-names>J</given-names></string-name>, <string-name><surname>McMahan</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Ramage</surname> <given-names>D</given-names></string-name>, <string-name><surname>Richt&#x00E1;rik</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Federated optimization: distributed machine learning for on-device intelligence</article-title>. <comment>arXiv:161002527</comment>. <comment>2016</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning</article-title>. <source>J Syst Archit</source>. <year>2022</year>;<volume>124</volume>(<issue>10</issue>):<fpage>102431</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.sysarc.2022.102431</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Asymptotic soft filter pruning for deep convolutional neural networks</article-title>. <source>IEEE Trans Cybern</source>. <year>2019</year>;<volume>50</volume>(<issue>8</issue>):<fpage>3594</fpage>&#x2013;<lpage>604</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCYB.2019.2933477</pub-id>; <pub-id pub-id-type="pmid">31478883</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shaikh</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bilal</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Pruning filters with L1-norm and capped L1-norm for CNN compression</article-title>. <source>Appl Intell</source>. <year>2021</year>;<volume>51</volume>(<issue>2</issue>):<fpage>1152</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-020-01894-y</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Han</surname> <given-names>C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>EasiEdge: a novel global deep neural networks pruning method for efficient edge computing</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>8</volume>(<issue>3</issue>):<fpage>1259</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2020.3034925</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mondal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Das</surname> <given-names>B</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>SD</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>P</given-names></string-name>, <string-name><surname>Lall</surname> <given-names>B</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>SD</given-names></string-name></person-group>. <article-title>Adaptive CNN filter pruning using global importance metric</article-title>. <source>Comput Vis Image Underst</source>. <year>2022</year>;<volume>222</volume>(<issue>12</issue>):<fpage>103511</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cviu.2022.103511</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sarvani</surname> <given-names>C</given-names></string-name>, <string-name><surname>Dubey</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Ghorai</surname> <given-names>M</given-names></string-name></person-group>. <article-title>UFKT: unimportant filters knowledge transfer for CNN pruning</article-title>. <source>Neurocomputing</source>. <year>2022</year>;<volume>514</volume>(<issue>3</issue>):<fpage>101</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2022.09.150</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Enhancing CNN efficiency through mutual information-based filter pruning</article-title>. <source>Digit Signal Process</source>. <year>2024</year>;<volume>151</volume>(<issue>1</issue>):<fpage>104547</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dsp.2024.104547</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name></person-group>. <article-title>FPWT: filter pruning via wavelet transform for CNNs</article-title>. <source>Neural Netw</source>. <year>2024</year>;<volume>179</volume>(<issue>2</issue>):<fpage>106577</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2024.106577</pub-id>; <pub-id pub-id-type="pmid">39098265</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chung</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tsang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Asadipour</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Multi-objective evolutionary architectural pruning of deep convolutional neural networks with weights inheritance</article-title>. <source>Inf Sci</source>. <year>2024</year>;<volume>685</volume>(<issue>8</issue>):<fpage>121265</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ins.2024.121265</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kolf</surname> <given-names>JN</given-names></string-name>, <string-name><surname>Elliesen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Damer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Boutros</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Towards extreme face and periocular recognition model compression with mixed-precision quantization</article-title>. <source>Eng Appl Artif Intell</source>. <year>2024</year>;<volume>137</volume>(<issue>5</issue>):<fpage>109114</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.109114</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Minimal loss DNN model compression with vectorized weight quantization</article-title>. <source>IEEE Trans Comput</source>. <year>2020</year>;<volume>70</volume>(<issue>5</issue>):<fpage>696</fpage>&#x2013;<lpage>710</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TC.2020.2995593</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H-T</given-names></string-name></person-group>. <article-title>Deep network quantization via error compensation</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2021</year>;<volume>33</volume>(<issue>9</issue>):<fpage>4960</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2021.3064293</pub-id>; <pub-id pub-id-type="pmid">33852390</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Smart-DNN&#x002B;: a memory-efficient neural networks compression framework for the model inference</article-title>. <source>ACM Trans Archit Code Optim</source>. <year>2023</year>;<volume>20</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3617688</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>R</given-names></string-name></person-group>. <article-title>MBQuant: a novel multi-branch topology method for arbitrary bit-width network quantization</article-title>. <source>Pattern Recognit</source>. <year>2025</year>;<volume>158</volume>(<issue>7</issue>):<fpage>111061</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2024.111061</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>AC</given-names></string-name></person-group>. <article-title>MedQ: lossless ultra-low-bit neural network quantization for medical image segmentation</article-title>. <source>Med Image Anal</source>. <year>2021</year>;<volume>73</volume>(<issue>1</issue>):<fpage>102200</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.media.2021.102200</pub-id>; <pub-id pub-id-type="pmid">34416578</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shamim</surname> <given-names>MZM</given-names></string-name></person-group>. <article-title>Hardware deployable edge-AI solution for prescreening of oral tongue lesions using TinyML on embedded devices</article-title>. <source>IEEE Embedd Syst Lett</source>. <year>2022</year>;<volume>14</volume>(<issue>4</issue>):<fpage>183</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/LES.2022.3160281</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Park</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A practical wearable fall detection system based on tiny convolutional neural networks</article-title>. <source>Biomed Signal Process Control</source>. <year>2023</year>;<volume>86</volume>(<issue>4</issue>):<fpage>105325</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2023.105325</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>An ultra-low power tinyml system for real-time visual processing at edge</article-title>. <source>IEEE Trans Circuits Syst II: Express Briefs</source>. <year>2023</year>;<volume>70</volume>(<issue>7</issue>):<fpage>2640</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCSII.2023.3239044</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thonglek</surname> <given-names>K</given-names></string-name>, <string-name><surname>Takahashi</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ichikawa</surname> <given-names>K</given-names></string-name>, <string-name><surname>Nakasan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Nakada</surname> <given-names>H</given-names></string-name>, <string-name><surname>Takano</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Automated quantization and retraining for neural network models without labeled data</article-title>. <source>IEEE Access</source>. <year>2022</year>;<volume>10</volume>:<fpage>73818</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2022.3190627</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>AC</given-names></string-name></person-group>. <article-title>EfficientQ: an efficient and accurate post-training neural network quantization method for medical image segmentation</article-title>. <source>Med Image Anal</source>. <year>2024</year>;<volume>97</volume>(<issue>1</issue>):<fpage>103277</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.media.2024.103277</pub-id>; <pub-id pub-id-type="pmid">39094461</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>J</given-names></string-name></person-group>. <article-title>MPQ-YOLO: ultra low mixed-precision quantization of YOLO for edge devices deployment</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>574</volume>(<issue>7</issue>):<fpage>127210</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2023.127210</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Pei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>G</given-names></string-name></person-group>. <article-title>GXNOR-Net: training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework</article-title>. <source>Neural Netw</source>. <year>2018</year>;<volume>100</volume>:<fpage>49</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2018.01.010</pub-id>; <pub-id pub-id-type="pmid">29471195</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Enderich</surname> <given-names>L</given-names></string-name>, <string-name><surname>Timm</surname> <given-names>F</given-names></string-name>, <string-name><surname>Burgard</surname> <given-names>W</given-names></string-name></person-group>. <article-title>SYMOG: learning symmetric mixture of Gaussian modes for improved fixed-point quantization</article-title>. <source>Neurocomputing</source>. <year>2020</year>;<volume>416</volume>:<fpage>310</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2019.11.114</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Quantization through search: a novel scheme to quantize convolutional neural networks in finite weight space</article-title>. In: <conf-name>28th Asia and South Pacific Design Automation Conference (ASP-DAC)</conf-name>; <year>2023</year>; <publisher-loc>Tokyo, Japan</publisher-loc>. p. <fpage>378</fpage>&#x2013;<lpage>83</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Hessian-based mixed-precision quantization with transition aware training for neural networks</article-title>. <source>Neural Netw</source>. <year>2024</year>;<volume>182</volume>(<issue>1&#x2013;4</issue>):<fpage>106910</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2024.106910</pub-id>; <pub-id pub-id-type="pmid">39579751</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sharma</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kidambi</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Mukhopadhyay</surname> <given-names>S</given-names></string-name></person-group>. <article-title>HamQ: hamming weight-based energy aware quantization for analog compute-in-memory accelerator in intelligent sensors</article-title>. <source>IEEE Sens J</source>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2024.3382479</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jung</surname> <given-names>S</given-names></string-name>, <string-name><surname>Son</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Son</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>J-J</given-names></string-name>, <string-name><surname>Kwak</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning to quantize deep networks by optimizing quantization intervals with task loss</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2019</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>4345</fpage>&#x2013;<lpage>54</lpage>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>So</surname> <given-names>HC</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep convolutional neural network compression via coupled tensor decomposition</article-title>. <source>IEEE J Sel Top Signal Process</source>. <year>2020</year>;<volume>15</volume>(<issue>3</issue>):<fpage>603</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSTSP.2020.3038227</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Swaminathan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Garg</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kannan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Andres</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Sparse low rank factorization for deep neural network compression</article-title>. <source>Neurocomputing</source>. <year>2020</year>;<volume>398</volume>(<issue>11</issue>):<fpage>185</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2020.02.035</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nekooei</surname> <given-names>A</given-names></string-name>, <string-name><surname>Safari</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Compression of deep neural networks based on quantized tensor decomposition to implement on reconfigurable hardware platforms</article-title>. <source>Neural Netw</source>. <year>2022</year>;<volume>150</volume>:<fpage>350</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2022.02.024</pub-id>; <pub-id pub-id-type="pmid">35344706</pub-id></mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Joint matrix decomposition for deep convolutional neural networks compression</article-title>. <source>Neurocomputing</source>. <year>2023</year>;<volume>516</volume>(<issue>3</issue>):<fpage>11</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2022.10.021</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hsiao</surname> <given-names>T-Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>Y-C</given-names></string-name>, <string-name><surname>Chou</surname> <given-names>H-H</given-names></string-name>, <string-name><surname>Chiu</surname> <given-names>C-T</given-names></string-name></person-group>. <article-title>Filter-based deep-compression with global average pooling for convolutional networks</article-title>. <source>J Syst Archit</source>. <year>2019</year>;<volume>95</volume>(<issue>3</issue>):<fpage>9</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.sysarc.2019.02.008</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hussain</surname> <given-names>I</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A knowledge distillation based deep learning framework for cropped images detection in spatial domain</article-title>. <source>Signal Process Image Commun</source>. <year>2024</year>;<volume>124</volume>(<issue>24</issue>):<fpage>117117</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.image.2024.117117</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dai</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A light-weight skeleton human action recognition model with knowledge distillation for edge intelligent surveillance applications</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>151</volume>(<issue>10</issue>):<fpage>111166</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2023.111166</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Amjad</surname> <given-names>K</given-names></string-name>, <string-name><surname>Asif</surname> <given-names>S</given-names></string-name>, <string-name><surname>Waheed</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A novel lightweight deep learning framework with knowledge distillation for efficient diabetic foot ulcer detection</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>167</volume>(<issue>1</issue>):<fpage>112296</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112296</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>He</surname> <given-names>P</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A novel model compression method based on joint distillation for deepfake video detection</article-title>. <source>J King Saud Univ Comput Inf Sci</source>. <year>2023</year>;<volume>35</volume>(<issue>9</issue>):<fpage>101792</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jksuci.2023.101792</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ham</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J-H</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Difficulty level-based knowledge distillation</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>606</volume>:<fpage>128375</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2024.128375</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Hybrid knowledge distillation from intermediate layers for efficient single image super-resolution</article-title>. <source>Neurocomputing</source>. <year>2023</year>;<volume>554</volume>(<issue>7</issue>):<fpage>126592</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2023.126592</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Niyaz</surname> <given-names>U</given-names></string-name>, <string-name><surname>Sambyal</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Bathula</surname> <given-names>DR</given-names></string-name></person-group>. <article-title>Leveraging different learning styles for improved knowledge distillation in biomedical imaging</article-title>. <source>Comput Biol Med</source>. <year>2024</year>;<volume>168</volume>(<issue>12</issue>):<fpage>107764</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.107764</pub-id>; <pub-id pub-id-type="pmid">38056210</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Multidimensional knowledge distillation for multimodal scene classification of remote sensing images</article-title>. <source>Digit Signal Process</source>. <year>2024</year>;<volume>157</volume>:<fpage>104876</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dsp.2024.104876</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nooruddin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Karray</surname> <given-names>F</given-names></string-name>, <string-name><surname>Muhammad</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A multi-resolution fusion approach for human activity recognition from video data in tiny edge devices</article-title>. <source>Inf Fusion</source>. <year>2023</year>;<volume>100</volume>(<issue>3</issue>):<fpage>101953</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2023.101953</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Pushp</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>DeepWear: adaptive local offloading for on-wearable deep learning</article-title>. <source>IEEE Trans Mob Comput</source>. <year>2019</year>;<volume>19</volume>(<issue>2</issue>):<fpage>314</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TMC.2019.2893250</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hsu</surname> <given-names>C-H</given-names></string-name></person-group>. <article-title>DNNOff: offloading DNN-based intelligent IoT applications in mobile edge computing</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2021</year>;<volume>18</volume>(<issue>4</issue>):<fpage>2820</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2021.3075464</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gauttam</surname> <given-names>H</given-names></string-name>, <string-name><surname>Pattanaik</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Bhadauria</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nain</surname> <given-names>G</given-names></string-name>, <string-name><surname>Prakash</surname> <given-names>PB</given-names></string-name></person-group>. <article-title>An efficient DNN splitting scheme for edge-AI enabled smart manufacturing</article-title>. <source>J Ind Inf Integr</source>. <year>2023</year>;<volume>34</volume>(<issue>14</issue>):<fpage>100481</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jii.2023.100481</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wolter</surname> <given-names>K</given-names></string-name></person-group>. <article-title>DDPQN: an efficient DNN offloading strategy in local-edge-cloud collaborative environments</article-title>. <source>IEEE Trans Serv Comput</source>. <year>2021</year>;<volume>15</volume>(<issue>2</issue>):<fpage>640</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TSC.2021.3116597</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mwase</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Da Xu</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Self-aware collaborative edge inference with embedded devices for IIoT</article-title>. <source>Future Gener Comput Syst</source>. <year>2025</year>;<volume>163</volume>:<fpage>107535</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2024.107535</pub-id>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Palena</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cerquitelli</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chiasserini</surname> <given-names>CF</given-names></string-name></person-group>. <article-title>Edge-device collaborative computing for multi-view classification</article-title>. <source>Comput Netw</source>. <year>2024</year>;<volume>254</volume>(<issue>8</issue>):<fpage>110823</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comnet.2024.110823</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Pimentel</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Stefanov</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Model and system robustness in distributed CNN inference at the edge</article-title>. <source>Integration</source>. <year>2025</year>;<volume>100</volume>(<issue>1</issue>):<fpage>102299</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.vlsi.2024.102299</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Hamila</surname> <given-names>R</given-names></string-name>, <string-name><surname>Erbad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gabbouj</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Distributed inference in resource-constrained iot for real-time video surveillance</article-title>. <source>IEEE Syst J</source>. <year>2022</year>;<volume>17</volume>(<issue>1</issue>):<fpage>1512</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSYST.2022.3198711</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>D</given-names></string-name>, <string-name><surname>Castiglione</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>BB</given-names></string-name>, <string-name><surname>Karuppiah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>PCNN CEC: efficient and privacy-preserving convolutional neural network inference based on cloud-edge-client collaboration</article-title>. <source>IEEE Trans Netw Sci Eng</source>. <year>2022</year>;<volume>10</volume>(<issue>5</issue>):<fpage>2906</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNSE.2022.3177755</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>C-H</given-names></string-name>, <string-name><surname>Jha</surname> <given-names>NK</given-names></string-name></person-group>. <article-title>DOCTOR: a multi-disease detection continual learning framework based on wearable medical sensors</article-title>. <source>ACM Trans Embed Comput Syst</source>. <year>2024</year>;<volume>23</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3679050</pub-id>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pellegrini</surname> <given-names>L</given-names></string-name>, <string-name><surname>Graffieti</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lomonaco</surname> <given-names>V</given-names></string-name>, <string-name><surname>Maltoni</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Latent replay for real-time continual learning</article-title>. In: <conf-name>2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</conf-name>; <year>2020</year>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>10203</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ravaglia</surname> <given-names>L</given-names></string-name>, <string-name><surname>Rusci</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nadalini</surname> <given-names>D</given-names></string-name>, <string-name><surname>Capotondi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Conti</surname> <given-names>F</given-names></string-name>, <string-name><surname>Benini</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A tinyml platform for on-device continual learning with quantized latent replays</article-title>. <source>IEEE J Emerg Sel Top Circuits Syst</source>. <year>2021</year>;<volume>11</volume>(<issue>4</issue>):<fpage>789</fpage>&#x2013;<lpage>802</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JETCAS.2021.3121554</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deutel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hannig</surname> <given-names>F</given-names></string-name>, <string-name><surname>Mutschler</surname> <given-names>C</given-names></string-name>, <string-name><surname>Teich</surname> <given-names>J</given-names></string-name></person-group>. <article-title>On-device training of fully quantized deep neural networks on Cortex-M microcontrollers</article-title>. <source>IEEE Trans Comput Aided Des Integr Circuits Syst</source>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.1109/TCAD.2024.3484354</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xia</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>LEAF: an adaptation framework against noisy data on edge through ultra low-cost training</article-title>. In: <conf-name>Proceedings of the 61st ACM/IEEE Design Automation Conference</conf-name>; <year>2024</year>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Novoa-Paradela</surname> <given-names>D</given-names></string-name>, <string-name><surname>Fontenla-Romero</surname> <given-names>O</given-names></string-name>, <string-name><surname>Guijarro-Berdi&#x00F1;as</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Fast deep autoencoder for federated learning</article-title>. <source>Pattern Recognit</source>. <year>2023</year>;<volume>143</volume>(<issue>2</issue>):<fpage>109805</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2023.109805</pub-id>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bosch</surname> <given-names>J</given-names></string-name>, <string-name><surname>Olsson</surname> <given-names>HH</given-names></string-name></person-group>. <article-title>Enabling efficient and low-effort decentralized federated learning with the EdgeFL framework</article-title>. <source>Inf Softw Tech</source>. <year>2025</year>;<volume>178</volume>(<issue>3</issue>):<fpage>107600</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.infsof.2024.107600</pub-id>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Efficient knowledge management for heterogeneous federated continual learning on resource-constrained edge devices</article-title>. <source>Future Gener Comput Syst</source>. <year>2024</year>;<volume>156</volume>:<fpage>16</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2024.02.018</pub-id>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hamalainen</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Importance-aware data selection and resource allocation for hierarchical federated edge learning</article-title>. <source>Future Gener Comput Syst</source>. <year>2024</year>;<volume>154</volume>(<issue>6</issue>):<fpage>35</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2023.12.014</pub-id>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>FedDA: resource-adaptive federated learning with dual-alignment aggregation optimization for heterogeneous edge devices</article-title>. <source>Future Gener Comput Syst</source>. <year>2025</year>;<volume>163</volume>(<issue>3</issue>):<fpage>107551</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2024.107551</pub-id>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Federated learning framework based on trimmed mean aggregation rules</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>270</volume>(<issue>11</issue>):<fpage>126354</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.126354</pub-id>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>AC-C</given-names></string-name></person-group>. <article-title>FedCM: federated learning with client-level momentum</article-title>. <comment>arXiv:210610874</comment>. <comment>2021</comment>.</mixed-citation></ref>
<ref id="ref-83"><label>[83]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thorgeirsson</surname> <given-names>AT</given-names></string-name>, <string-name><surname>Gauterin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Probabilistic predictions with federated learning</article-title>. <source>Entropy</source>. <year>2020</year>;<volume>23</volume>(<issue>1</issue>):<fpage>41</fpage>. doi:<pub-id pub-id-type="doi">10.3390/e23010041</pub-id>; <pub-id pub-id-type="pmid">33396677</pub-id></mixed-citation></ref>
<ref id="ref-84"><label>[84]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Dynamic aggregation for heterogeneous quantization in federated learning</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2021</year>;<volume>20</volume>(<issue>10</issue>):<fpage>6804</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TWC.2021.3076613</pub-id>.</mixed-citation></ref>
<ref id="ref-85"><label>[85]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Akhtarshenas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vahedifar</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ayoobi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Maham</surname> <given-names>B</given-names></string-name>, <string-name><surname>Alizadeh</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ebrahimi</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Federated learning: a cutting-edge survey of the latest advancements and applications</article-title>. <source>Comput Commun</source>. <year>2024</year>;<volume>228</volume>:<fpage>107964</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comcom.2024.107964</pub-id>.</mixed-citation></ref>
<ref id="ref-86"><label>[86]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karunathilake</surname> <given-names>E</given-names></string-name>, <string-name><surname>Le</surname> <given-names>AT</given-names></string-name>, <string-name><surname>Heo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>YS</given-names></string-name>, <string-name><surname>Mansoor</surname> <given-names>S</given-names></string-name></person-group>. <article-title>The path to smart farming: innovations and opportunities in precision agriculture</article-title>. <source>Agriculture</source>. <year>2023</year>;<volume>13</volume>(<issue>8</issue>):<fpage>1593</fpage>. doi:<pub-id pub-id-type="doi">10.3390/agriculture13081593</pub-id>.</mixed-citation></ref>
<ref id="ref-87"><label>[87]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tsoukas</surname> <given-names>V</given-names></string-name>, <string-name><surname>Gkogkidis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kakarountas</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A tinyml-based system for smart agriculture</article-title>. In: <conf-name>Proceedings of the 26th Pan-Hellenic Conference on Informatics</conf-name>; <year>2022</year>; <publisher-loc>Athens, Greece</publisher-loc>. p. <fpage>207</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-88"><label>[88]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gookyi</surname> <given-names>DAN</given-names></string-name>, <string-name><surname>Wulnye</surname> <given-names>FA</given-names></string-name>, <string-name><surname>Arthur</surname> <given-names>EAE</given-names></string-name>, <string-name><surname>Ahiadormey</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Agyemang</surname> <given-names>JO</given-names></string-name>, <string-name><surname>Agyekum</surname> <given-names>KO-BO</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>TinyML for smart agriculture: comparative analysis of TinyML platforms and practical deployment for maize leaf disease identification</article-title>. <source>Smart Agri Technol</source>. <year>2024</year>;<volume>8</volume>(<issue>2</issue>):<fpage>100490</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.atech.2024.100490</pub-id>.</mixed-citation></ref>
<ref id="ref-89"><label>[89]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhattacharya</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pandey</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deploying an energy efficient, secure &#x0026; high-speed sidechain-based TinyML model for soil quality monitoring and management in agriculture</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>242</volume>(<issue>5</issue>):<fpage>122735</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.122735</pub-id>.</mixed-citation></ref>
<ref id="ref-90"><label>[90]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wulnye</surname> <given-names>FA</given-names></string-name>, <string-name><surname>Arthur</surname> <given-names>EAE</given-names></string-name>, <string-name><surname>Gookyi</surname> <given-names>DAN</given-names></string-name>, <string-name><surname>Asiedu</surname> <given-names>DKP</given-names></string-name>, <string-name><surname>Wilson</surname> <given-names>M</given-names></string-name>, <string-name><surname>Agyemang</surname> <given-names>JO</given-names></string-name></person-group>. <article-title>TinyML implementation on microcontrollers: the case of maize leaf disease identification</article-title>. In: <conf-name>2024 Conference on Information Communications Technology and Society (ICTAS)</conf-name>; <year>2024</year>; <publisher-loc>Durban, South Africa</publisher-loc>. p. <fpage>180</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-91"><label>[91]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dockendorf</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mitra</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mohanty</surname> <given-names>SP</given-names></string-name>, <string-name><surname>Kougianos</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Lite-Agro: exploring light-duty computing platforms for IoAT-Edge AI in plant disease identification.In</article-title>. In: <conf-name>IFIP International Internet of Things Conference</conf-name>; <year>2023</year>; <publisher-loc>Denton, TX, USA</publisher-loc>. p. <fpage>371</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-92"><label>[92]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Azevedo</surname> <given-names>MB</given-names></string-name>, <string-name><surname>de Medeiros</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Medeiros</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>I</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>DG</given-names></string-name></person-group>. <article-title>Detecting face masks through embedded machine learning algorithms: a transfer learning approach for affordable microcontrollers</article-title>. <source>Mach Learn Appl</source>. <year>2023</year>;<volume>14</volume>(<issue>10</issue>):<fpage>100498</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.mlwa.2023.100498</pub-id>.</mixed-citation></ref>
<ref id="ref-93"><label>[93]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Saha</surname> <given-names>B</given-names></string-name>, <string-name><surname>Samanta</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ghosh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>RB</given-names></string-name></person-group>. <article-title>BandX: an intelligent IoT-band for human activity recognition based on TinyML</article-title>. In: <conf-name>Proceedings of the 24th International Conference on Distributed Computing and Networking</conf-name>; <year>2023</year>; <publisher-loc>Kharagpur, India</publisher-loc>. p. <fpage>284</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-94"><label>[94]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arthi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Krishnaveni</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Optimized tiny machine learning and explainable AI for trustable and energy-efficient fog-enabled healthcare decision support system</article-title>. <source>Int J Comput Intell Syst</source>. <year>2024</year>;<volume>17</volume>(<issue>1</issue>):<fpage>229</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s44196-024-00631-4</pub-id>.</mixed-citation></ref>
<ref id="ref-95"><label>[95]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gaud</surname> <given-names>N</given-names></string-name>, <string-name><surname>Rathore</surname> <given-names>M</given-names></string-name>, <string-name><surname>Suman</surname> <given-names>U</given-names></string-name></person-group>. <article-title>MHCNLS-HAR: multi-headed CNN-LSTM based human activity recognition leveraging a novel wearable edge device for elderly health care</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>21</issue>):<fpage>35394</fpage>&#x2013;<lpage>405</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2024.3450499</pub-id>.</mixed-citation></ref>
<ref id="ref-96"><label>[96]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bayes</surname> <given-names>S</given-names></string-name>, <string-name><surname>Abotaleb</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>The case for TinyML in healthcare: CNNs for real-time on-edge blood pressure estimation</article-title>. In: <conf-name>Proceedings of the 38th ACM/SIGAPP Symposium on Applied Computing</conf-name>; <year>2023</year>; <publisher-loc>Tallinn, Estonia</publisher-loc>. p. <fpage>629</fpage>&#x2013;<lpage>38</lpage>.</mixed-citation></ref>
<ref id="ref-97"><label>[97]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Andrade</surname> <given-names>P</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>I</given-names></string-name>, <string-name><surname>Diniz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Flores</surname> <given-names>T</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>DG</given-names></string-name>, <string-name><surname>Soares</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Online processing of vehicular data on the edge through an unsupervised TinyML regression technique</article-title>. <source>ACM Trans Embed Comput Syst</source>. <year>2024</year>;<volume>23</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3591356</pub-id>.</mixed-citation></ref>
<ref id="ref-98"><label>[98]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Andrade</surname> <given-names>P</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>M</given-names></string-name>, <string-name><surname>Medeiros</surname> <given-names>M</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>DG</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>I</given-names></string-name></person-group>. <article-title>TEDA-RLS: a TinyML incremental learning approach for outlier detection and correction</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>22</issue>):<fpage>38165</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2024.3458917</pub-id>.</mixed-citation></ref>
<ref id="ref-99"><label>[99]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Im</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>S</given-names></string-name></person-group>. <article-title>TinyML-based intrusion detection system for in-vehicle network using convolutional neural network on embedded devices</article-title>. <source>IEEE Embedd Syst Lett</source>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.1109/LES.2024.3475470</pub-id>.</mixed-citation></ref>
<ref id="ref-100"><label>[100]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saini</surname> <given-names>M</given-names></string-name>, <string-name><surname>Adebayo</surname> <given-names>SO</given-names></string-name>, <string-name><surname>Arora</surname> <given-names>V</given-names></string-name></person-group>. <article-title>IoT-Fog-based framework to prevent vehicle-road accidents caused by self-visual distracted drivers</article-title>. <source>Multimed Tools Appl</source>. <year>2024</year>;<volume>83</volume>(<issue>42</issue>):<fpage>90133</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-024-19050-w</pub-id>.</mixed-citation></ref>
<ref id="ref-101"><label>[101]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Medeiros</surname> <given-names>M</given-names></string-name>, <string-name><surname>Flores</surname> <given-names>T</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>M</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>I</given-names></string-name></person-group>. <article-title>A multi-layered methodology for driver behavior analysis using TinyML and edge computing</article-title>. In: <conf-name>2024 IEEE International Conference on Evolving and Adaptive Intelligent Systems (EAIS)</conf-name>; <year>2024</year>; <publisher-loc>Madrid, Spain</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-102"><label>[102]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>H</given-names></string-name>, <string-name><surname>Anicic</surname> <given-names>D</given-names></string-name>, <string-name><surname>Runkler</surname> <given-names>TA</given-names></string-name></person-group>. <article-title>Tinyol: TinyML with online-learning on microcontrollers</article-title>. In: <conf-name>2021 International Joint Conference on Neural Networks (IJCNN)</conf-name>; <year>2021</year>; <publisher-loc>Shenzhen, China</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-103"><label>[103]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>H</given-names></string-name>, <string-name><surname>Anicic</surname> <given-names>D</given-names></string-name>, <string-name><surname>Runkler</surname> <given-names>TA</given-names></string-name></person-group>. <article-title>Towards semantic management of on-device applications in industrial IoT</article-title>. <source>ACM Trans Internet Technol</source>. <year>2022</year>;<volume>22</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3510820</pub-id>.</mixed-citation></ref>
<ref id="ref-104"><label>[104]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Schoedel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alavilli</surname> <given-names>A</given-names></string-name>, <string-name><surname>Plancher</surname> <given-names>B</given-names></string-name>, <string-name><surname>Manchester</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>TinyMPC: model-predictive control on resource-constrained microcontrollers</article-title>. In: <conf-name>2024 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>; <year>2024</year>; <publisher-loc>Yokohama, Japan</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-105"><label>[105]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asutkar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chalke</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shivgan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Tallur</surname> <given-names>S</given-names></string-name></person-group>. <article-title>TinyML-enabled edge implementation of transfer learning framework for domain generalization in machine fault diagnosis</article-title>. <source>Expert Syst Appl</source>. <year>2023</year>;<volume>213</volume>(<issue>1</issue>):<fpage>119016</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2022.119016</pub-id>.</mixed-citation></ref>
<ref id="ref-106"><label>[106]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>T-H</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>C-T</given-names></string-name>, <string-name><surname>Putranto</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Tiny machine learning empowers climbing inspection robots for real-time multiobject bolt-defect detection</article-title>. <source>Eng Appl Artif Intell</source>. <year>2024</year>;<volume>133</volume>(<issue>12</issue>):<fpage>108618</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.108618</pub-id>.</mixed-citation></ref>
<ref id="ref-107"><label>[107]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ksira</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mellit</surname> <given-names>A</given-names></string-name>, <string-name><surname>Blasuttigh</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pavan</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>A novel embedded system for real-time fault diagnosis of photovoltaic modules</article-title>. <source>IEEE J Photovolt</source>. <year>2024</year>;<volume>14</volume>(<issue>12</issue>):<fpage>354</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JPHOTOV.2024.3359462</pub-id>.</mixed-citation></ref>
<ref id="ref-108"><label>[108]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mellit</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An embedded solution for fault detection and diagnosis of photovoltaic modules using thermographic images and deep convolutional neural networks</article-title>. <source>Eng Appl Artif Intell</source>. <year>2022</year>;<volume>116</volume>(<issue>5668</issue>):<fpage>105459</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2022.105459</pub-id>.</mixed-citation></ref>
<ref id="ref-109"><label>[109]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hayajneh</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Alasali</surname> <given-names>F</given-names></string-name>, <string-name><surname>Salama</surname> <given-names>A</given-names></string-name>, <string-name><surname>Holderbaum</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Intelligent solar forecasts: modern machine learning models &#x0026; TinyML role for improved solar energy yield predictions</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>10846</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3354703</pub-id>.</mixed-citation></ref>
<ref id="ref-110"><label>[110]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Boiko</surname> <given-names>O</given-names></string-name>, <string-name><surname>Komin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shendryk</surname> <given-names>V</given-names></string-name>, <string-name><surname>Malekian</surname> <given-names>R</given-names></string-name>, <string-name><surname>Davidsson</surname> <given-names>P</given-names></string-name></person-group>. <article-title>TinyML on mobile devices for hybrid energy management systems</article-title>. In: <conf-name>2024 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing &#x0026; Communications (GreenCom) and IEEE Cyber, Physical &#x0026; Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics</conf-name>; <year>2024</year>; <publisher-loc>Copenhagen, Denmark</publisher-loc>. p. <fpage>200</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-111"><label>[111]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fernandes</surname> <given-names>R</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gomes</surname> <given-names>R</given-names></string-name>, <string-name><surname>Vila&#x00E7;a</surname> <given-names>N</given-names></string-name></person-group>. <article-title>SmartLVEnergy: an AIoT framework for energy management through distributed processing and sensor-actuator integration in legacy low-voltage systems</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>13</issue>):<fpage>20726</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2024.3403484</pub-id>.</mixed-citation></ref>
<ref id="ref-112"><label>[112]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hussain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abughanam</surname> <given-names>N</given-names></string-name>, <string-name><surname>Qadir</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Jamming detection in IoT wireless networks: an edge-AI based approach</article-title>. In: <conf-name>Proceedings of the 12th International Conference on the Internet of Things</conf-name>; <year>2023</year>; <publisher-loc>Delft, Netherlands</publisher-loc>. p. <fpage>57</fpage>&#x2013;<lpage>64</lpage>.</mixed-citation></ref>
<ref id="ref-113"><label>[113]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Giordano</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baumann</surname> <given-names>N</given-names></string-name>, <string-name><surname>Crabolu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fischer</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bellusci</surname> <given-names>G</given-names></string-name>, <string-name><surname>Magno</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Design and performance evaluation of an ultralow-power smart IoT device with embedded TinyML for asset activity monitoring</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2022</year>;<volume>71</volume>:<fpage>2510711</fpage>. doi:<pub-id pub-id-type="doi">10.1109/TIM.2022.3165816</pub-id>.</mixed-citation></ref>
<ref id="ref-114"><label>[114]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Agrawal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Maiti</surname> <given-names>RR</given-names></string-name></person-group>. <article-title>TinyAP: an intelligent access point to combat Wi-Fi Attacks using TinyML</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>12</volume>(<issue>2</issue>):<fpage>2135</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3467328</pub-id>.</mixed-citation></ref>
<ref id="ref-115"><label>[115]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Saranya</surname> <given-names>T</given-names></string-name>, <string-name><surname>Jeyamala</surname> <given-names>D</given-names></string-name>, <string-name><surname>Sellamuthu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A secure framework for MIoT: TinyML-powered emergency alerts and intrusion detection for secure real-time monitoring</article-title>. In: <conf-name>2024 8th International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC)</conf-name>; <year>2024</year>; <publisher-loc>Kirtipur, Nepal</publisher-loc>. p. <fpage>13</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-116"><label>[116]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chakraborty</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Alharbi</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An energy harvesting algorithm for UAV-assisted TinyML consumer electronic in low-power IoT networks</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2024</year>;<volume>70</volume>(<issue>4</issue>):<fpage>7346</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCE.2024.3419784</pub-id>.</mixed-citation></ref>
<ref id="ref-117"><label>[117]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmad</surname> <given-names>U</given-names></string-name>, <string-name><surname>Han</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jolfaei</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jabbar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ibrar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Erbad</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive survey and tutorial on smart vehicles: emerging technologies, security issues, and solutions using machine learning</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2024</year>;<volume>25</volume>(<issue>11</issue>):<fpage>15314</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2024.3419988</pub-id>.</mixed-citation></ref>
<ref id="ref-118"><label>[118]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huda</surname> <given-names>NU</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>I</given-names></string-name>, <string-name><surname>Adnan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>M</given-names></string-name>, <string-name><surname>Naeem</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Experts and intelligent systems for smart homes&#x2019; transformation to sustainable smart cities: a comprehensive review</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>238</volume>(<issue>9</issue>):<fpage>122380</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.122380</pub-id>.</mixed-citation></ref>
<ref id="ref-119"><label>[119]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hernandez</surname> <given-names>C</given-names></string-name>, <string-name><surname>Taslimi</surname> <given-names>B</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>HY</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Pardalos</surname> <given-names>PM</given-names></string-name></person-group>. <article-title>Training generalizable quantized deep neural nets</article-title>. <source>Expert Syst Appl</source>. <year>2023</year>;<volume>213</volume>(<issue>2</issue>):<fpage>118736</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2022.118736</pub-id>.</mixed-citation></ref>
<ref id="ref-120"><label>[120]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gui</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Automatic neural network compression by sparsity-quantization joint learning: a constrained optimization-based approach</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2020</year>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>2175</fpage>&#x2013;<lpage>85</lpage>.</mixed-citation></ref>
<ref id="ref-121"><label>[121]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>SHAQ</given-names></string-name></person-group>. <article-title>Hardware-aware automated quantization with mixed precision</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2019</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>8612</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-122"><label>[122]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jayasimhan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pabitha</surname> <given-names>P</given-names></string-name></person-group>. <article-title>ResPrune: an energy-efficient restorative filter pruning method using stochastic optimization for accelerating CNN</article-title>. <source>Pattern Recognit</source>. <year>2024</year>;<volume>155</volume>:<fpage>110671</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2024.110671</pub-id>.</mixed-citation></ref>
<ref id="ref-123"><label>[123]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Poyatos</surname> <given-names>J</given-names></string-name>, <string-name><surname>Molina</surname> <given-names>D</given-names></string-name>, <string-name><surname>Martinez</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Del Ser</surname> <given-names>J</given-names></string-name>, <string-name><surname>Herrera</surname> <given-names>F</given-names></string-name></person-group>. <article-title>EvoPruneDeepTL: an evolutionary pruning model for transfer learning based deep neural networks</article-title>. <source>Neural Netw</source>. <year>2023</year>;<volume>158</volume>(<issue>3</issue>):<fpage>59</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2022.10.011</pub-id>; <pub-id pub-id-type="pmid">36442374</pub-id></mixed-citation></ref>
<ref id="ref-124"><label>[124]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Towards automatic model compression via a unified two-stage framework</article-title>. <source>Pattern Recognit</source>. <year>2023</year>;<volume>140</volume>(<issue>1</issue>):<fpage>109527</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2023.109527</pub-id>.</mixed-citation></ref>
<ref id="ref-125"><label>[125]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>C-M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ARLP: automatic multi-agent transformer reinforcement learning pruner for one-shot neural network pruning</article-title>. <source>Knowl Based Syst</source>. <year>2024</year>;<volume>300</volume>:<fpage>112122</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2024.112122</pub-id>.</mixed-citation></ref>
<ref id="ref-126"><label>[126]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Albanese</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nardello</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fiacco</surname> <given-names>G</given-names></string-name>, <string-name><surname>Brunelli</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Tiny machine learning for high accuracy product quality inspection</article-title>. <source>IEEE Sens J</source>. <year>2022</year>;<volume>23</volume>(<issue>2</issue>):<fpage>1575</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2022.3225227</pub-id>.</mixed-citation></ref>
<ref id="ref-127"><label>[127]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Multicenter hierarchical federated learning with fault-tolerance mechanisms for resilient edge computing networks</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2024</year>;<volume>36</volume>(<issue>1</issue>):<fpage>47</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2024.3362974</pub-id>; <pub-id pub-id-type="pmid">38546988</pub-id></mixed-citation></ref>
<ref id="ref-128"><label>[128]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brandic</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Sustainable and trustworthy edge machine learning</article-title>. <source>IEEE Internet Comput</source>. <year>2021</year>;<volume>25</volume>(<issue>5</issue>):<fpage>5</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MIC.2021.3104383</pub-id>.</mixed-citation></ref>
<ref id="ref-129"><label>[129]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Tiwari</surname> <given-names>RG</given-names></string-name>, <string-name><surname>Haroon</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tripathi</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>V</given-names></string-name></person-group>. <chapter-title>A system model of fault tolerance technique in distributed system and scalable system using machine learning</chapter-title>. In: <source>Software-defined network frameworks</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>CRC Press</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>16</lpage>.</mixed-citation></ref>
<ref id="ref-130"><label>[130]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Binh</surname> <given-names>HTT</given-names></string-name>, <string-name><surname>PoisonGAN</surname> <given-names>YS</given-names></string-name></person-group>. <article-title>Generative poisoning attacks against federated learning in edge computing systems</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>8</volume>(<issue>5</issue>):<fpage>3310</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2020.3023126</pub-id>.</mixed-citation></ref>
<ref id="ref-131"><label>[131]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bolchini</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cassano</surname> <given-names>L</given-names></string-name>, <string-name><surname>Miele</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Resilience of deep learning applications: a systematic literature review of analysis and hardening techniques</article-title>. <source>Comput Sci Rev</source>. <year>2024</year>;<volume>54</volume>(<issue>521</issue>):<fpage>100682</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cosrev.2024.100682</pub-id>.</mixed-citation></ref>
<ref id="ref-132"><label>[132]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Narayanan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pattabiraman</surname> <given-names>K</given-names></string-name>, <string-name><surname>Debardeleben</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Fault injection for TensorFlow applications</article-title>. <source>IEEE Trans Dependable Secure Comput</source>. <year>2022</year>;<volume>20</volume>(<issue>4</issue>):<fpage>2677</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TDSC.2022.3175930</pub-id>.</mixed-citation></ref>
<ref id="ref-133"><label>[133]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Laskar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Characterizing deep learning neural network failures between algorithmic inaccuracy and transient hardware faults</article-title>. In: <conf-name>2022 IEEE 27th Pacific Rim International Symposium on Dependable Computing (PRDC)</conf-name>; <year>2022</year>; <publisher-loc>Beijing, China</publisher-loc>. p. <fpage>54</fpage>&#x2013;<lpage>67</lpage>.</mixed-citation></ref>
<ref id="ref-134"><label>[134]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Syed</surname> <given-names>RT</given-names></string-name>, <string-name><surname>Ulbricht</surname> <given-names>M</given-names></string-name>, <string-name><surname>Piotrowski</surname> <given-names>K</given-names></string-name>, <string-name><surname>Krstic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Fault resilience analysis of quantized deep neural networks</article-title>. In: <conf-name>2021 IEEE 32nd International Conference on Microelectronics (MIEL)</conf-name>; <year>2021</year>; <publisher-loc>Nis, Serbia</publisher-loc>. p. <fpage>275</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-135"><label>[135]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ruospo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sanchez</surname> <given-names>E</given-names></string-name>, <string-name><surname>Traiola</surname> <given-names>M</given-names></string-name>, <string-name><surname>O&#x2019;connor</surname> <given-names>I</given-names></string-name>, <string-name><surname>Bosio</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Investigating data representation for efficient and reliable convolutional neural networks</article-title>. <source>Microprocess Microsyst</source>. <year>2021</year>;<volume>86</volume>(<issue>7</issue>):<fpage>104318</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.micpro.2021.104318</pub-id>.</mixed-citation></ref>
<ref id="ref-136"><label>[136]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>An efficient structure to improve the reliability of deep neural networks on ARMs</article-title>. <source>Microelectron Reliab</source>. <year>2022</year>;<volume>136</volume>(<issue>2</issue>):<fpage>114729</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.microrel.2022.114729</pub-id>.</mixed-citation></ref>
<ref id="ref-137"><label>[137]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Reliability evaluation of FPGA based pruned neural networks</article-title>. <source>Microelectron Reliab</source>. <year>2022</year>;<volume>130</volume>(<issue>8</issue>):<fpage>114498</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.microrel.2022.114498</pub-id>.</mixed-citation></ref>
<ref id="ref-138"><label>[138]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ficco</surname> <given-names>M</given-names></string-name>, <string-name><surname>Guerriero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Milite</surname> <given-names>E</given-names></string-name>, <string-name><surname>Palmieri</surname> <given-names>F</given-names></string-name>, <string-name><surname>Pietrantuono</surname> <given-names>R</given-names></string-name>, <string-name><surname>Russo</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Federated learning for IoT devices: enhancing TinyML with on-board training</article-title>. <source>Inf Fusion</source>. <year>2024</year>;<volume>104</volume>(<issue>3</issue>):<fpage>102189</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2023.102189</pub-id>.</mixed-citation></ref>
<ref id="ref-139"><label>[139]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gupta</surname> <given-names>BB</given-names></string-name>, <string-name><surname>Quamara</surname> <given-names>M</given-names></string-name></person-group>. <article-title>An overview of Internet of Things (IoT): architectural aspects, challenges, and protocols</article-title>. <source>Concurr Comput</source>. <year>2020</year>;<volume>32</volume>(<issue>21</issue>):<fpage>e4946</fpage>. doi:<pub-id pub-id-type="doi">10.1002/cpe.4946</pub-id>.</mixed-citation></ref>
<ref id="ref-140"><label>[140]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>MT</given-names></string-name>, <string-name><surname>Truong</surname> <given-names>HL</given-names></string-name></person-group>. <article-title>On optimizing resources for real-time end-to-end machine learning in heterogeneous edges</article-title>. <source>Softw Pract Exp</source>. <year>2025</year>;<volume>55</volume>(<issue>3</issue>):<fpage>541</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1002/spe.3383</pub-id>.</mixed-citation></ref>
<ref id="ref-141"><label>[141]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Mo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Multi-objective resource allocation in mobile edge computing using PAES for Internet of Things</article-title>. <source>Wirel Netw</source>. <year>2024</year>;<volume>30</volume>(<issue>5</issue>):<fpage>3533</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11276-020-02409-w</pub-id>.</mixed-citation></ref>
<ref id="ref-142"><label>[142]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asghari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Azgomi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Darvishmofarahi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Multi-objective edge server placement using the whale optimization algorithm and game theory</article-title>. <source>Soft Comput</source>. <year>2023</year>;<volume>27</volume>(<issue>21</issue>):<fpage>16143</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00500-023-07995-3</pub-id>.</mixed-citation></ref>
<ref id="ref-143"><label>[143]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Moustakas</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tziouvaras</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kolomvatsos</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Data and resource aware incremental ML training in support of pervasive applications</article-title>. <source>Computing</source>. <year>2024</year>;<volume>106</volume>(<issue>11</issue>):<fpage>3727</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00607-024-01338-2</pub-id>.</mixed-citation></ref>
<ref id="ref-144"><label>[144]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yoosefi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kargahi</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Resource-aware in-edge distributed real-time deep learning</article-title>. <source>Internet of Things</source>. <year>2024</year>;<volume>27</volume>(<issue>8</issue>):<fpage>101263</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2024.101263</pub-id>.</mixed-citation></ref>
<ref id="ref-145"><label>[145]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mu&#x00F1;oz</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Jannesari</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Resource-aware heterogeneous federated learning with specialized local models</article-title>. In: <conf-name>European Conference on Parallel Processing</conf-name>; <year>2024</year>; <publisher-loc>Madrid, Spain</publisher-loc>. p. <fpage>389</fpage>&#x2013;<lpage>403</lpage>.</mixed-citation></ref>
<ref id="ref-146"><label>[146]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Adaptive personalized federated learning with one-shot screening</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>11</volume>(<issue>9</issue>):<fpage>15375</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2023.3346900</pub-id>.</mixed-citation></ref>
<ref id="ref-147"><label>[147]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Nimmagadda</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Model optimization techniques for edge devices</chapter-title>. In: <source>Model optimization methods for efficient and edge AI: federated learning architectures, frameworks and applications</source>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>Wiley-IEEE Press</publisher-name>; <year>2025</year>. p. <fpage>57</fpage>&#x2013;<lpage>85</lpage>.</mixed-citation></ref>
<ref id="ref-148"><label>[148]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rajapakse</surname> <given-names>V</given-names></string-name>, <string-name><surname>Karunanayake</surname> <given-names>I</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Intelligence at the extreme edge: a survey on reformable tinyml</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>13s</issue>):<fpage>1</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3583683</pub-id>.</mixed-citation></ref>
<ref id="ref-149"><label>[149]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>W</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Review of neural network model acceleration techniques based on FPGA platforms</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>610</volume>(<issue>7</issue>):<fpage>128511</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2024.128511</pub-id>.</mixed-citation></ref>
<ref id="ref-150"><label>[150]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Subramaniam</surname> <given-names>EVD</given-names></string-name>, <string-name><surname>Srinivasan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Qaisar</surname> <given-names>SM</given-names></string-name>, <string-name><surname>P&#x0142;awiak</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Interoperable IoMT approach for remote diagnosis with privacy-preservation perspective in edge systems</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>17</issue>):<fpage>7474</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23177474</pub-id>; <pub-id pub-id-type="pmid">37687933</pub-id></mixed-citation></ref>
<ref id="ref-151"><label>[151]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pliatsios</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kotis</surname> <given-names>K</given-names></string-name>, <string-name><surname>Goumopoulos</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A systematic review on semantic interoperability in the IoE-enabled smart cities</article-title>. <source>Internet Things</source>. <year>2023</year>;<volume>22</volume>(<issue>1</issue>):<fpage>100754</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2023.100754</pub-id>.</mixed-citation></ref>
<ref id="ref-152"><label>[152]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nilsson</surname> <given-names>J</given-names></string-name>, <string-name><surname>Javed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Albertsson</surname> <given-names>K</given-names></string-name>, <string-name><surname>Delsing</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liwicki</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sandin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>AI concepts for system of systems dynamic interoperability</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>9</issue>):<fpage>2921</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24092921</pub-id>; <pub-id pub-id-type="pmid">38733028</pub-id></mixed-citation></ref>
<ref id="ref-153"><label>[153]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>C-Y</given-names></string-name>, <string-name><surname>Chuang</surname> <given-names>Y-C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>C-T</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>A-Y</given-names></string-name></person-group>. <article-title>Recent progress and development of hyperdimensional computing (HDC) for edge intelligence</article-title>. <source>IEEE J Emerg Sel Top Circuits Syst</source>. <year>2023</year>;<volume>13</volume>(<issue>1</issue>):<fpage>119</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JETCAS.2023.3242767</pub-id>.</mixed-citation></ref>
<ref id="ref-154"><label>[154]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hassan</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bettayeb</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mohammad</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Advancing hardware implementation of hyperdimensional computing for edge intelligence</article-title>. In: <conf-name>2024 IEEE 6th International Conference on AI Circuits and Systems (AICAS)</conf-name>; <year>2024</year>; <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>. p. <fpage>169</fpage>&#x2013;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="ref-155"><label>[155]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multi-granularity fusion resource allocation algorithm based on dual-attention deep reinforcement learning and lifelong learning architecture in heterogeneous IIoT</article-title>. <source>Inf Fusion</source>. <year>2023</year>;<volume>99</volume>:<fpage>101871</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2023.101871</pub-id>.</mixed-citation></ref>
<ref id="ref-156"><label>[156]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mir</surname> <given-names>NF</given-names></string-name></person-group>. <article-title>AI-assisted edge computing for multi-tenant management of edge devices in 6G networks</article-title>. In: <conf-name>2023 2nd International Conference on 6G Networking (6GNet)</conf-name>; <year>2023</year>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>3</lpage>.</mixed-citation></ref>
<ref id="ref-157"><label>[157]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jedidi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Dynamic trust security approach for edge computing-based mobile IoT devices using artificial intelligence</article-title>. <source>Eng Res Express</source>. <year>2024</year>;<volume>6</volume>(<issue>2</issue>):<fpage>25211</fpage>. doi:<pub-id pub-id-type="doi">10.1088/2631-8695/ad43b5</pub-id>.</mixed-citation></ref>
<ref id="ref-158"><label>[158]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Puri</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Challenges and opportunities in implementing quantum-safe key distribution in IoT devices</article-title>. In: <conf-name>2024 3rd International Conference for Innovation in Technology (INOCON)</conf-name>; <year>2024</year>; <publisher-loc>Bangalore, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-159"><label>[159]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dharani</surname> <given-names>D</given-names></string-name>, <string-name><surname>Anitha Kumari</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A smart surveillance system utilizing modified federated machine learning: gossip-verifiable and quantum-safe approach</article-title>. <source>Concurr Comput</source>. <year>2024</year>;<volume>36</volume>(<issue>24</issue>):<fpage>e8238</fpage>. doi:<pub-id pub-id-type="doi">10.1002/cpe.8238</pub-id>.</mixed-citation></ref>
<ref id="ref-160"><label>[160]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ansere</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Gyamfi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>V</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dobre</surname> <given-names>OA</given-names></string-name>, <string-name><surname>Duong</surname> <given-names>TQ</given-names></string-name></person-group>. <article-title>Quantum deep reinforcement learning for dynamic resource allocation in mobile edge computing-based IoT systems</article-title>. <source>IEEE Trans Wirel Commun</source>. <year>2023</year>;<volume>23</volume>(<issue>6</issue>):<fpage>6221</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TWC.2023.3330868</pub-id>.</mixed-citation></ref>
<ref id="ref-161"><label>[161]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karakaya</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ulu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A survey on post-quantum based approaches for edge computing security</article-title>. <source>Wiley Interdiscip Rev Comput Stat</source>. <year>2024</year>;<volume>16</volume>(<issue>1</issue>):<fpage>e1644</fpage>. doi:<pub-id pub-id-type="doi">10.1002/wics.1644</pub-id>.</mixed-citation></ref>
<ref id="ref-162"><label>[162]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khonina</surname> <given-names>SN</given-names></string-name>, <string-name><surname>Kazanskiy</surname> <given-names>NL</given-names></string-name>, <string-name><surname>Skidanov</surname> <given-names>RV</given-names></string-name>, <string-name><surname>Butt</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Exploring types of photonic neural networks for imaging and computing&#x2014;a review</article-title>. <source>Nanomaterials</source>. <year>2024</year>;<volume>14</volume>(<issue>8</issue>):<fpage>697</fpage>. doi:<pub-id pub-id-type="doi">10.3390/nano14080697</pub-id>; <pub-id pub-id-type="pmid">38668191</pub-id></mixed-citation></ref>
<ref id="ref-163"><label>[163]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bandyopadhyay</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sludds</surname> <given-names>A</given-names></string-name>, <string-name><surname>Krastanov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hamerly</surname> <given-names>R</given-names></string-name>, <string-name><surname>Harris</surname> <given-names>N</given-names></string-name>, <string-name><surname>Bunandar</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Single-chip photonic deep neural network with forward-only training</article-title>. <source>Nat Photonics</source>. <year>2024</year>;<volume>18</volume>(<issue>12</issue>):<fpage>1335</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41566-024-01567-z</pub-id>.</mixed-citation></ref>
<ref id="ref-164"><label>[164]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rani</surname> <given-names>F</given-names></string-name>, <string-name><surname>Chollet</surname> <given-names>N</given-names></string-name>, <string-name><surname>Vogt</surname> <given-names>L</given-names></string-name>, <string-name><surname>Urbas</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Industrial edge MLOps: overview and challenges</article-title>. <source>Comput Aided Chem Eng</source>. <year>2024</year>;<volume>53</volume>(<issue>11</issue>):<fpage>3019</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1016/B978-0-443-28824-1.50504-4</pub-id>.</mixed-citation></ref>
<ref id="ref-165"><label>[165]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Shabir</surname> <given-names>MY</given-names></string-name>, <string-name><surname>Torta</surname> <given-names>G</given-names></string-name>, <string-name><surname>Basso</surname> <given-names>A</given-names></string-name>, <string-name><surname>Damiani</surname> <given-names>F</given-names></string-name></person-group>. <chapter-title>Toward secure TinyML on a standardized AI architecture</chapter-title>. In: <source>Device-edge-cloud continuum: paradigms, architectures and applications</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>121</fpage>&#x2013;<lpage>39</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>






















