<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">76413</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.076413</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Adaptive Intrusion Detection Framework for IoT: Balancing Accuracy and Computational Efficiency</article-title>
<alt-title alt-title-type="left-running-head">An Adaptive Intrusion Detection Framework for IoT: Balancing Accuracy and Computational Efficiency</alt-title>
<alt-title alt-title-type="right-running-head">An Adaptive Intrusion Detection Framework for IoT: Balancing Accuracy and Computational Efficiency</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Alsulami</surname><given-names>Abdulaziz A.</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>aaalsulami10@kau.edu.sa</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Alturki</surname><given-names>Badraddin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Tayeb</surname><given-names>Ahmad J.</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Alsemmeari</surname><given-names>Rayan A.</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Alsini</surname><given-names>Raed</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Information Systems, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Information Technology, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Abdulaziz A. Alsulami. Email: <email>aaalsulami10@kau.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>9</day><month>4</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>3</issue>
<elocation-id>48</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>11</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_76413.pdf"></self-uri>
<abstract>
<p>Intrusion Detection Systems (IDS) play a critical role in protecting networked environments from cyberattacks. They have become increasingly important in smart environments such as the Internet of Things (IoT) systems. However, IDS for IoT networks face critical challenges due to hardware constraints, including limited computational resources and storage capacity, which lead to high feature dimensionality, prediction uncertainty, and increased processing cost. These factors make many conventional detection approaches unsuitable for real-time IoT deployment. To address these challenges, this paper proposes an adaptive intrusion detection framework that intelligently balances detection accuracy and computational efficiency. The proposed framework integrates mutual information (MI) feature selection model, deep contextual embeddings, and an adaptive decision mechanism. The MI model identifies and retains the most informative features, which reduces dimensionality while maintaining high detection accuracy. The adaptive decision dynamically selects between multiple inference paths to ensure that additional computation is needed only when the uncertainty level is high. Experimental evaluations on benchmark IoT datasets namely RT-IoT-2022, CIC-IoT-2023 and CIC-IoMT-2024 show that the proposed framework achieves F1-score of 99.92%, 96.66%, and 99.84%, respectively, with an average inference time of approximately 0.105 ms per sample. These results demonstrate that the framework effectively adapts inference complexity to data uncertainty, which provides an intelligent, interpretable and efficient solution for real-world IoT intrusion detection.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Adaptive IDS</kwd>
<kwd>cyberattack detection</kwd>
<kwd>computational cost</kwd>
<kwd>IoT security</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Deanship of Scientific Research (DSR) at King Abdulaziz University</funding-source>
<award-id>753-611-2025</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The Internet of Things (IoT) is an emerging technology that is already gaining acceptance and has been used in a variety of application areas including wearables, smart homes, healthcare services, agriculture, logistics, smart cities, and automotive [<xref ref-type="bibr" rid="ref-1">1</xref>]. Because it supports efficient communication of information across real-world objects and accomplishes this without the need of user intervention. IoT is a ubiquitous, self-organized network that has the ability to significantly change our way of interaction with others and our environment [<xref ref-type="bibr" rid="ref-2">2</xref>]. The ability to remotely monitor and manage the environment has been made feasible by the connectivity of IoT devices, which improves the ability to make decisions and thus improves the life of individuals and communities [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>There are around 20 billion connected devices around the world, and by the end of 2030, there will be around 31 billion [<xref ref-type="bibr" rid="ref-4">4</xref>]. Additionally, by 2030, the IoT is projected to generate 684.1 billion dollars in income yearly worldwide [<xref ref-type="bibr" rid="ref-5">5</xref>]. Individuals, businesses and community sectors could benefit from the widespread implementation of IoT in several asymmetrical environments and application domains. While IoT devices are becoming more accessible, they often lack strong security features [<xref ref-type="bibr" rid="ref-6">6</xref>]. The existing constraints on IoT devices cause them to be a target for attackers, which leads to a range of attacks such as man-in-the-middle attacks, and network scanning, spear phishing, malware infiltration, eavesdropping, selective forwarding, keystroke logging, SQL injections and tampering, physical damage [<xref ref-type="bibr" rid="ref-7">7</xref>]. A Denial of Service (DoS) attack can be of the most harmful attack type to IoT devices as it restricts genuine users from reaching services. A cyberattack of this kind can cause considerable damage to IoT services and smart ecosystem applications that operate within a network of IoT devices. As a result, protecting IoT systems has become a growing challenge [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>In the context of IoT, intrusion is defined as any attack, malicious activity, or unauthorized access that threatens the availability, integrity, or confidentiality of connected devices or data across consumer and enterprise networks [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. Because IoT systems are diverse, resource-constrained and reliant on lightweight standards, they offer a broader attack surface since they are widely deployed in critical sectors such as production, energy, healthcare, and transportation [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>Conventional security systems were initially created for static, standard, and high power systems, but they are often not effective in the constantly evolving attack environment of IoT [<xref ref-type="bibr" rid="ref-13">13</xref>]. The IoT infrastructures have constraints including processing power, memory and real-time responsiveness that differentiate them against traditional networks. They are therefore particularly vulnerable to adaptive cyber threats, which provide a risk and have the ability to compromise system security and interrupt with important operations of the company. Furthermore, the growing number of autonomous and sensitive IoT systems requires the use of creative, intelligent, and context-aware security mechanisms [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>The security of IoT networks and systems requires intrusion detection systems (IDSs) and intrusion prevention systems (IPSs). These systems are necessary to protect data, prevent cyberattacks, and respond to security regulations [<xref ref-type="bibr" rid="ref-16">16</xref>]. Utilizing a variety of techniques, including anomaly detection, behavioral analysis, and investigating existing attack patterns, IDSs identify potential security breaches and immediately alert users of any threats to ensure rapid responses of security concerns [<xref ref-type="bibr" rid="ref-17">17</xref>]. Since that many IoT devices lack adequate security features, IPSs helps in network protection by examining data packets and preventing known threats. Security of IoT devices can also be ensured by IPSs, which address vulnerabilities, encrypt communication for added safety, and act as an essential defensive mechanism in the complicated IoT system [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>Because IoT environments are dynamic, distributed and diverse, intrusion detection is a challenging but essential task [<xref ref-type="bibr" rid="ref-18">18</xref>]. Since these platforms comprise a variety of devices with different processing capabilities, protocols of communication and data formats, which makes performing cybersecurity analysis a challenging task in such contexts. Furthermore, the large amount of real-time data and the absence of annotated attack instances particularly for undetected attacks, which limits the ability to identify threats in an adaptive way [<xref ref-type="bibr" rid="ref-19">19</xref>]. Class imbalance, noise and low-latency computing requirements make it more difficult to use traditional deep learning (DL) techniques that emphasizes the need for more specialized, flexible and resource-effective methods [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>Sensor data reliability is a critical aspect in IoT security systems because IDS performance is heavily dependent on the quality and accuracy of data generated by heterogeneous sensors such as vibration, temperature, humidity and motion sensors. Faulty, noisy, drifting or spoofed sensor readings can introduce false patterns into the network, which affects both anomaly detection accuracy and model stability. Recent studies like [<xref ref-type="bibr" rid="ref-21">21</xref>] highlight that unreliable sensors significantly degrade the trustworthiness of IoT operations and can even mimic intrusion signatures. Therefore, incorporating reliability assessment or fault tolerant preprocessing techniques is essential for robust intrusion detection.</p>
<p>Existing machine learning (ML) and DL based intrusion detection systems are resilient and provide high accuracy in classification in conventional Internet environments. However, existing IDS methods can consume more resources and can be considered as unsuitable for IoT [<xref ref-type="bibr" rid="ref-22">22</xref>]. Furthermore, the latest developments in ML based IDS require more resources and computing power which can help improve accuracy. As a result, these methods need to be resource effective to be appropriate for IoT devices with limited resources [<xref ref-type="bibr" rid="ref-23">23</xref>]. Due to all the challenges that are above mentioned, there is a need to have an adaptive and cost effective IDS in IoT environments.</p>
<p>Our study aims to address the challenges of IoT environments characterized by limited computational resources, high feature dimensionality, and heterogeneous architectures. To overcome these limitations, we propose an adaptive intrusion detection framework that integrates mutual information (MI) model for feature selection, calibrated random forest classifiers, and deep contextual embeddings from the efficiently learning an encoder that classifies token replacements accurately (ELECTRA) transformer. The framework employs a Linear Upper Confidence Bound (LinUCB) algorithm as an adaptive decision mechanism that dynamically selects between multiple inference paths, effectively balancing detection accuracy and computational efficiency.</p>
<p>Our main contribution in this work is summarized as follows:
<list list-type="bullet">
<list-item>
<p>We propose an adaptive IDS framework that dynamically balances detection accuracy and computational efficiency for IoT environments, supported by a mutual-information (MI) feature selection model that reduces dimensionality by removing redundant and less-informative features while preserving important content.</p></list-item>
<list-item>
<p>We implement a parallel pipeline architecture using random forest classifiers: a Base Model trained on all features and a Reduced Model trained on the top MI-selected features. This setup enables adaptive decision-making during inference.</p></list-item>
<list-item>
<p>We integrate a pretrained ELECTRA language model to generate deep contextual embeddings, which are combined with a logistic regression Teacher Probe to refine predictions under high uncertainty and enhance the understanding of complex IoT traffic patterns.</p></list-item>
<list-item>
<p>We use a contextual LinUCB bandit algorithm to select among four inference actions namely Reduced Model, Base Model, large language model (LLM) Refinement and Escalation, which optimizes the trade-off between accuracy and computational cost. The framework is evaluated on three IoT datasets&#x2014;RT-IoT-2022, CIC-IoT-2023, and CIC-IoMT-2024&#x2014;achieving F1-scores of 99.94%, 94%, and 98%, respectively, with significant reduction in computational overhead.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews the related work. <xref ref-type="sec" rid="s3">Section 3</xref> describes the proposed method, including the system architecture and feature extraction. <xref ref-type="sec" rid="s4">Section 4</xref> presents the experimental results, compares them with existing benchmark models, and discusses the findings. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper and outlines directions for future research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>In this section, we discuss an overview of recent related works that design advanced IDS models for IoT setting. Then, we highlight the used datasets, learning methods, feature selection techniques, processes, and findings. The discussion emphasizes the way that the previous research has attempted to improve detection accuracy, computational effectiveness, and flexibility of integrating the DL and ML methods. <xref ref-type="table" rid="table-1">Table 1</xref> discusses recent IDS methods in IoT environments including datasets utilized, approaches and techniques used, the accuracy achieved, the computational cost and research gap.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of related IDS approaches in IoT.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ref.</th>
<th>Year</th>
<th>Methods</th>
<th>Datasets</th>
<th>Accuracy</th>
<th>Computational Cost</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>2024</td>
<td>MLP&#x002B; FS (IG, GR, CFS, Pearson, SU)</td>
<td>RT-IoT-2022</td>
<td>96.4%</td>
<td>Runtime reduced by 66.4% via feature selection</td>
<td>Uses static set of 16 features and lacks an adaptive decision layer</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>2023</td>
<td>FS &#x002B; PCA &#x002B; ANN/DNN/ TabNet</td>
<td>RT-IoT-2022</td>
<td>99.7%</td>
<td>Lightweight computation via reduced features from 83 to 5</td>
<td>Uses static dimensionality reduction (83 to 5 features)</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>2025</td>
<td>CSMCR &#x002B; RegNet &#x002B; FBNet</td>
<td>IoTID20, N-BaIoT, RT-IoT-2022, Bot-IoT</td>
<td>91%&#x2013;100%</td>
<td>Training time reduced by 53% via CSMCR</td>
<td>High computational costs without a cost-aware mechanism</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>2024</td>
<td>Lightweight Host IDS &#x002B; XGBoost &#x002B; Entropy</td>
<td>CIC-IoT-2022, Bot-IoT</td>
<td>99.97%</td>
<td>99.7% reduction in processing time; 86.4% in memory</td>
<td>Primarily host-based</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>2024</td>
<td>CNN&#x002B;LSTM Stacked &#x002B; Random Forest Meta-Classifier</td>
<td>CIC-IoT-2022 &#x002B; NF datasets</td>
<td>91%&#x2013;99.99%</td>
<td>Average inference time of 1.5 ms</td>
<td>High inference latency</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>2024</td>
<td>CNN-LSTM &#x002B; Random Forest &#x002B; Transfer Learning &#x002B; SHAP/LIME</td>
<td>CIC-IoT-2022, CIC-IoT-2023, Edge-IIoT</td>
<td>98.2%</td>
<td>Handled big data by Spark optimization and reduced the size of the network by data preprocessing techniques</td>
<td>Lacks an intelligent bandit-based controller for real-time resource optimization</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>2025</td>
<td>Incremental Learning &#x002B; Blockchain &#x002B; Encryption</td>
<td>CIC-IoT-2023, CIC-DDoS2019, TON-IoT2020</td>
<td>98%&#x2013;99.89%</td>
<td>Low encryption and decryption overhead 0.0001 s&#x2013;0.0003 s and transaction time 2&#x2013;3 s</td>
<td>Blockchain integration introduces high latency</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>2025</td>
<td>RF, DT, KNN, SVM, AdaBoost</td>
<td>CIC-IoT-2023</td>
<td>99.125%</td>
<td>Optimized for resource efficiency via feature selection</td>
<td>Evaluation is limited to a single dataset</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>2025</td>
<td>Correlation &#x002B; SFS &#x002B; Cascaded LSTM &#x002B; NB</td>
<td>CIC-DDoS2019, CIC-IoT-2023, CIC-IoV-2024</td>
<td>99.7%&#x2013;99.9%</td>
<td>Fastest intrusion detection time is 0.041 s</td>
<td>Employs a static hybrid classification model</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>2025</td>
<td>RF, DT, XGB, FNN</td>
<td>CICIoT-2023, CICIoMT-2024</td>
<td>99.85%</td>
<td>Resource efficiency by optimizing the dataset and hyperparameter tuning</td>
<td>Lacks an intelligent decision layer</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>2023</td>
<td>XGBoost &#x002B; SHAP &#x002B; Voting Fusion</td>
<td>General IoT threat data</td>
<td>97%</td>
<td>Bandwidth reduced by 73% and detection latency below 200 ms</td>
<td>Relies on post-hoc SHAP values for interpretability</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>2023</td>
<td>UNet&#x002B;&#x002B; &#x002B; LSTM for IoMT</td>
<td>CICIoMT-2024</td>
<td>87.96%&#x2013;99.92%</td>
<td>The inference speed is 166.12 <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="normal">&#x0B5;</mml:mi></mml:mrow></mml:math></inline-formula>s per sample and the model size is 8.04 MB</td>
<td>Uses a complex UNet&#x002B;&#x002B;/LSTM stack that is computationally expensive and difficult to reduce</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The authors in [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed an anomaly IDS architecture that uses a Multi-Layer Perceptron (MLP) classifier on the reduced set of features and a combined feature selection strategy to find the optimal features for the detection of IoT attacks. They used information gain, gain ratio, correlation-based feature subset selection, pearson&#x2019;s correlation analysis, symmetric uncertainty and MLP as a classifier. They evaluated their model using the RT-IoT-2022 dataset. The MLP achieved an accuracy rate of 93.5% with the original data with 83 features and 96.4% with all FS methods with 16 features. The study in [<xref ref-type="bibr" rid="ref-25">25</xref>] improved intrusion detection by introducing a semi supervised learning technique that uses a self-training loop and uncertainty filtering based on entropy. They have used XGBoost, random forest, extremely randomized trees (XRT), gradient boosting classifier (GBC), decision tree (DT), and entropy-based dynamic threshold filtering with self-training. In their evaluation, they used three datasets namely RT-IoT-2022, CICIoT-2023, and CICIoMT-2024 datasets. They achieved the highest overall accuracy of 100% with DT using the RT-IoT-2022 dataset, and the highest overall accuracy of 93% with XGBoost and XRT using the CICIoT2023 dataset, and achieved accuracy of 98% with random forest and XGB using CICIoMT2024. The researchers in [<xref ref-type="bibr" rid="ref-26">26</xref>] introduced an approach to maximize feature dimensionality for real-time intrusion detection in IoT contexts by integrating five feature selection techniques including gain ratio, correlation based feature subset selection, Pearson analysis, information gain, and symmetrical uncertainty with PCA and classifiers used such artificial neural networks (ANNs), deep neural networks (DNNs) and TabNet. They evaluated their approach using the RT-IoT-2022 dataset. They have achieved the highest accuracy rate of 99.7% by using Pearson&#x2013;PCA with ANN that are applied on the RT-IoT-2022 dataset. According to the comparison table provided by the authors that their work outperformed other related works in the field. The work in [<xref ref-type="bibr" rid="ref-27">27</xref>] presented a method named cosine similarity based majority class reduction (CSMCR). It reduces duplicated majority class instances while maintains the integrity of the dataset. Despite traditional methods like synthetic minority over sampling technique (SMOTE) or random undersampling, they used CSMCR to ensure the retained majority samples remain diverse by evaluating similarity across features. This reduces information loss and avoids unnecessary data duplication. In addition, they created a hybrid DL model that improved feature extraction and classification performance by integrating the RegNet and FBNet architectures. In the evaluation process, they used four datasets including IoTID20, N-BaIoT, RT-IoT-2022, UNSW Bot-IoT. Their model achieved accuracy rates of 96.18% on the IoTID20 dataset, 100% on N-BaIoT, 98.40% on RT-IoT-2022, and 91.10% on UNSW Bot-IoT.</p>
<p>The article in [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a lightweight host-based intrusion detection system that can characterize communication activities with a variety of entropies. They used methods namely XGBoost, host-based intrusion detection systems, data aggregation and several entropies. They evaluated their methods using the Bot-IoT Dataset and the CIC IoT 2022 Dataset. The accuracy rating achieved by their approaches was 99.97%. According to the findings, the proposed approach can reduce the processing time by 99.7% (2916 ms) and memory usage by 86.4% (633 MiB). The study in [<xref ref-type="bibr" rid="ref-29">29</xref>] proposed a stacking ensemble IDS model called (StaEn-IDS). A CNN and a LSTM model are integrated with a Deep Intrusion Detection System (DIDS). To increase accuracy, the random forest algorithm is used as a meta-classifier throughout the stacking procedure. Their StaEn-IDS was evaluated toward other models in the literature using a range of datasets including NF-ToN-IoT, NF-BoT IoT, NF-UNSW NB15, NF-CSE CIC IDS 2018 and NF-UQ NIDS. They achieved accuracies of 91.58% on NF-ToN-IoT, 95.82% on NF-BoT IoT, 99.49% on NF-UNSW NB15, 98.44% on NF-CSE CIC IDS 2018, 97.72% on NF-UQ NIDS and approximately 99.99% on CIC-IoT IDS 2022. The results show promising accuracy among all benchmark datasets. The authors in [<xref ref-type="bibr" rid="ref-30">30</xref>] present an enhanced IDS for IoT security that uses transfer learning and multimodal big data representation. They have used LIME and SHAP, Deep IDS (DIDS), CNN and LSTM as base learners. In addition, they utilized random forest as a meta-classifier and for Feature Selection they used Correlation and MI. The CIC-IoT 2022, CIC-IoT 2023 and Edge-IIoT are three benchmark IoT-based datasets that are utilized to evaluate the proposed approach. The proposed CNN-LSTM method achieves an accuracy of 98.2% on CIC-IoT dataset 2022.</p>
<p>The researchers in [<xref ref-type="bibr" rid="ref-31">31</xref>] used blockchain, hybrid encryption and incremental learning in three steps to create an integrated framework that secures and enhances IDS performance in the IoT. First, a model was created using incremental learning including SGD Classifier to identify cyberattacks. In the second step, the model has the ability to retrain itself on new data. The information has signed and encrypted in digital manner. The encrypted data is stored using a blockchain model. The incremental model is retrained using new data taken from the blockchain in the third step. They used CIC-IoT-2023 dataset to train their model. After retraining, the model evaluated with data that can be considered as unseen data for the model namely CIC IDS2019 and TON IoT2020 datasets. Their method achieved accuracy rates of 99.7% on CIC-DDoS2019 dataset, 98.27% on TON-IoT2020 dataset and 99.89% on CIC-IoT-2023 dataset. The research in [<xref ref-type="bibr" rid="ref-32">32</xref>] introduces an IDS method that improves cyber threat recognition and mitigation in smart cities by using ML. They created and evaluated a number of models including random forest, decision tree, k-nearest neighbors (KNN), support vector machine (SVM), and adaptive boosting (AdaBoost) by utilizing the CIC-IoT-2023 Dataset. Their proposed IDS system focuses on real-time detection of threats and ensures to obtain a high accuracy and minimal latency. Their solution performs well in cyberthreat detection and prevention due to robust data preparation and precise model training. The results of this study demonstrate the power of using AI techniques by presenting significant improvements in privacy and security of smart city in IoT architectures. They achieved an accuracy rate of 99.125% with the KNN model in binary classification. The authors in [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed a hybrid DL and ML approach that helps to identify DDoS and spoofing threats, lower false alarms and then put the required security in action. The first step of the proposed approach is to provide a feature selection that uses a correlation coefficient and a sequential feature selector. The second step involves a combination of DL neural networks with a cascaded LSTM and a Naive Bayes classifier. In the third stage, ports are restricted and network security measures are enhanced. The three datasets namely CIC-DDoS2019, CIC-IoT2023 and CIC IoV2024 were used in training and evaluating the performance of the proposed method and were balanced to provide useful outcomes. The method achieved an accuracy rates of 99.91% on CIC-DDoS 2019, 99.88% on CIC-IoT 2023 and 99.77% CIC IoV 2024 dataset. The test data was validated with cross-validation technique to ensure there was no overfitting.</p>
<p>The authors in [<xref ref-type="bibr" rid="ref-34">34</xref>] compared the performance of ML models that were trained on two datasets namely CICIoT2023 and CICIoMT2024, and they emphasized the importance of using domain specific data. Their results show that when models trained on single dataset are evaluated on another, the F1-score decreased significantly by around 66.87%. They evaluated in CICIoMT2024 and suggested baseline optimization methods such as appropriate train,validation and test splits, uniform windowing, modifying temporal dependencies for time series data and balance the imbalance dataset. For intrusion detection, they used three tree based ML classification methods: extreme gradient boosting (XGB), random forest (RF) and decision support (DT). They extended the scope of the research and emphasized the performance differences that a DL model could have in their situation, they also included a Feedforward Neural Network (FNN). Compared to other methods in the study, they achieved an accuracy rate of 99.85% and considered this a notable gain in IDS performance. The study in [<xref ref-type="bibr" rid="ref-35">35</xref>] proposed a high-performance cybersecurity framework that is interpretable and uses a precisely tuned XGBoost classifier to identify malicious attacks with high prediction accuracy. Their study compares the proposed approach with a baseline of regular logistic regression. They used SHAP (SHapley Additive exPlanations) to find important features that influenced predictions and examined the security cost trade-off while building ML systems for threat detection. They introduced a fusion technique based on max voting to effectively combine the advantages of the models. XGBoost outperforms Logistic Regression in terms of accuracy by achieving 97% and recall 100%, the results showed that the fusion model provided a more balanced performance with enhanced precision 98% and fewer false negatives. The researchers in [<xref ref-type="bibr" rid="ref-36">36</xref>] presented a DL method for analyzing Internet of Medical Things (IoMT) traffic and detecting malicious activity. They enhanced the prediction accuracy of current neural network models by using the CIC IoMT dataset. LSTM models and UNet&#x002B;&#x002B; are combined in their method to efficiently extract network traffic characteristics. The proposed approach outperformed conventional algorithms, according to experimental data, they achieved an accuracy rate of 87.96% in attack classification and 99.92% in anomaly detection.</p>
<p>Most existing studies on IDS for the IoT focused on improving performance by optimizing classifiers through techniques such as undersampling including CSMCR and dimensionality reduction such as PCA. However, these approaches have not fully explored the integration of deep contextual embeddings or adaptive decision mechanisms, such as bandit-based controllers. The reviewed papers employ ensemble tree models such as random forest, XGBoost or DL architectures like ANN, DNN, RegNet, FBNet and CNN without leveraging deep contextual embeddings such as ELECTRA for feature representation or utilizing adaptive selection strategies including LinUCB for dynamic model control. While existing studies achieve high accuracy, they typically rely on static computational costs, lack semantic context from deep embeddings and provide no adaptive paths for uncertain cases.</p>
<p>However, our study introduces an adaptive intrusion detection framework that works in IoT environments. The framework dynamically adjusts its behavior based on contextual factors such as computational power, classification accuracy, and security requirements. The adaptability behavior of our framework continuously optimizes the performance in real time by using deep contextual embeddings for feature representation and a bandit-based controller for classifier selection. This helps maintain high detection accuracy under the computational constraints of IoT devices.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>In this section, we present the research methodology. We propose an adaptive intrusion detection framework for IoT environments that integrates feature learning, deep embeddings, and intelligent decision control. Within this framework, multiple models run in parallel, allowing the system to choose dynamically among four possible actions: Reduced Model, Base Model, LLM Refinement, or Escalation based on current conditions. This helps to balance high detection accuracy with computational efficiency. The framework uses an MI-based feature selection model with dynamic cumulative fraction policies to keep the most informative features. This reduces redundancy while preserving essential information. Calibrated random forest classifiers operate in both base and reduced configurations. The ELECTRA embedding pipeline is combined with a logistic regression (Teacher Probe) to refine predictions in uncertain cases. During inference, a LinUCB algorithm dynamically chooses the best action to balance accuracy and computational cost.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Proposed Framework</title>
<p>The overall architecture of the proposed adaptive intrusion detection framework is illustrated in <xref ref-type="fig" rid="fig-1">Figs. 1</xref> and <xref ref-type="fig" rid="fig-2">2</xref>, which depict the training pipeline and the adaptive inference and decision-making process, respectively.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Architecture of the proposed IDS showing data preprocessing, MI-based feature selection, and training of the reduced and base models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-1.tif"/>
</fig><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Adaptive inference workflow using a LinUCB controller to select inference actions based on uncertainty and computational cost.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-2.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the framework begins with data preprocessing, where the dataset is divided into training, validation, and test subsets. Two parallel pipelines are then constructed for model training. In the first pipeline, MI with a cumulative fraction policy is applied to rank features according to their information contribution. Only features that collectively explain 90% of the total information are retained, and these features are used to train the Reduced Model, a lightweight random forest classifier designed to minimize computational cost during inference.</p>

<p>In parallel, the second pipeline trains the Base Model using the complete feature set. This model is implemented as a calibrated random forest to ensure reliable probability estimates, which are critical for downstream uncertainty assessment. In addition to these pipelines, both training and test samples are transformed into dense vector representations using the ELECTRA model [<xref ref-type="bibr" rid="ref-37">37</xref>]. These embeddings capture richer contextual relationships in the IoT traffic data and are used to train a logistic regression classifier, referred to as the Teacher Probe. The Teacher Probe serves as a refinement component that assists decision-making when model uncertainty is high.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the adaptive inference process during testing. Instead of relying on a single static classifier, the framework employs a LinUCB contextual bandit to dynamically select the most suitable inference action for each batch of test samples. This ensures a balance between detection accuracy and computational cost. For each batch, a context vector is constructed using batch-level statistics, such as the mean and standard deviation of the input features, and provided to the LinUCB controller.</p>
<p>Based on this context, the controller selects one of four possible actions: (i) the Reduced Model for efficient low-cost inference, (ii) the Base Model when higher reliability is required, (iii) LLM-based refinement using the Teacher Probe when uncertainty remains high, or (iv) escalation for manual review under extreme uncertainty. Each selected action produces a prediction output, which is evaluated using accuracy and F1-score metrics.</p>
<p>The selected action, prediction performance, and associated computational cost are then used to compute a reward signal, defined as performance minus action cost. This reward is used to update the LinUCB controller, completing the feedback loop. Through this iterative process, the controller learns to select inference actions that achieve an optimal trade-off between accuracy and efficiency under varying operational conditions.</p>
<p>The overall operation of the proposed framework is summarized in Algorithm 1, which presents the end-to-end training and adaptive inference procedure.</p>
<fig id="fig-12">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-12.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Proposed Hybrid IDS Components</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>SMOTE</title>
<p>Balancing the classes of an imbalance dataset is a crucial step to eliminate biased decisions caused by major classes over minority ones. In this work, SMOTE is applied only to the training subset in order to avoid information leakage and preserve the integrity of the validation and test distributions. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the process begins with loading the training set, then class distribution is examined to detect imbalance between majority and minority classes. The dataset is then separated into features <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>x</mml:mi></mml:math></inline-formula> and labels <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>y</mml:mi></mml:math></inline-formula>, and SMOTE is applied to generate synthetic samples for the minority classes. This step balances the dataset by ensuring that all classes are more evenly represented. The resampled features and labels are recombined into a new balanced dataset, and the class distribution is rechecked to confirm the effect of SMOTE. Finally, the balanced dataset is saved as a CSV file for use in subsequent model training. This process ensures that the learning algorithm does not become biased toward the majority class and can achieve more reliable and fair performance across all classes.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Data preprocessing workflow.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-3.tif"/>
</fig>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Mutual Information (MI) Feature Selection</title>
<p>MI between a feature <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> and the label <italic>Y</italic> is:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Y</mml:mi></mml:mrow></mml:munder><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> defines the standard MI feature selection, which measures how much knowing a feature <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> helps in predicting the class label <italic>Y</italic>. In simple terms, MI evaluates the dependency between a feature and the label. The term <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the joint probability of the feature value and the label occurring together. However, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represent their individual probabilities. Therefore, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> and <italic>Y</italic> are independent, then <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and the logarithmic term becomes zero. This means that the feature provides no useful information about the label. Larger MI values indicate that the feature carries more information about the label, making it more valuable for classification tasks. Features are ranked by <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and the top-<italic>K</italic> are selected.</p>
<p>The cumulative fraction policy (cumfrac) in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> selects the smallest number of top features <italic>K</italic> such that their combined MI accounts for at least 90% of the total MI across all features. In practice, this means we first sort features by how much information they share with the label. Then we keep adding features, one by one, until the fraction of total information they explain reaches the threshold <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula>. This ensures that most of the useful information is kept while removing redundant or less important features [<xref ref-type="bibr" rid="ref-38">38</xref>].
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>k</mml:mi><mml:mo>:</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></disp-formula>where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> are features sorted by decreasing MI score.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Base and Reduced Models</title>
<p>The framework employs two random forest classifiers, which are a base model and a reduced model. The base model uses the entire feature set with dimensionality <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>d</mml:mi></mml:math></inline-formula>, whereas the reduced model is restricted to only the top-<italic>K</italic> most informative features selected using MI. This allows the reduced model to achieve faster predictions with lower computational cost, while still maintaining competitive accuracy. The mathematical representations of the two models are given below:</p>
<p>The base model uses all <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>d</mml:mi></mml:math></inline-formula> features is given in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>RF</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></disp-formula></p>
<p>While the reduced model uses the top-<italic>K</italic> features is represented in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>reduced</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>RF</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>K</mml:mi><mml:mo>&#x226A;</mml:mo><mml:mi>d</mml:mi></mml:math></disp-formula></p>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> demonstrates the workflow of the base and reduced models. The base model directly uses all available input features, training a random forest followed by probability calibration without applying feature selection.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Workflow of base and reduced models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-4.tif"/>
</fig>
<p>In contrast, the reduced model applies MI feature ranking to the input features, retaining the top-<italic>K</italic> subset based on a cumulative fraction threshold of <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula>, which ensures that at least 90% of the overall information content is preserved. A random forest classifier is then trained on this reduced feature set. This dual setup, one model using all features and another using only the most informative subset, allows a clear comparison of their performance. This shows how the MI method helps remove redundancy, reduce dimensionality, and minimize noise.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>ELECTRA Embeddings and Teacher Probe</title>
<p>Although IoT features are numeric, transforming each feature vector into a tokenized text sequence allows ELECTRA to model dependencies between features that traditional tabular methods treat independently. The transformer captures contextual relationships, feature combinations, and non-linear interactions that often characterize malicious IoT traffic. Using a lightweight logistic regression probe on top of these embeddings provides an efficient refinement step when the base model is uncertain, enabling richer representation learning with minimal additional cost. As shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, each feature vector <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is first converted into a text string and then passed through the ELECTRA model, which produces a dense embedding <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>h</mml:mi></mml:msup></mml:math></inline-formula>. These embeddings capture richer semantic and structural patterns from the input features, which effectively maps raw numerical data into a representation space where complex relationships are easier to learn.</p>
<p>Features are transformed to text and embedded with ELECTRA:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ELECTRA</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mtext>Tokenizer</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>features</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>h</mml:mi></mml:msup></mml:math></disp-formula></p>
<p>After the embeddings are created the Teacher Probe is implemented as a logistic regression classifier as represented in <xref ref-type="disp-formula" rid="eqn-6">(6)</xref>. The probe is trained on embeddings derived from the training set, making it computationally efficient while still leveraging the representational power of ELECTRA. At test time, when the base random forest model is uncertain, the Teacher Probe predicts labels using the test embeddings. This combination allows the framework to integrate deep language model representations into a classical ML pipeline without requiring fine-tuning of the large transformer itself.</p>
<p>A teacher probe (logistic regression) is trained on embeddings:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>probe</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>LR</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_2_5">
<label>3.2.5</label>
<title>Entropy and Adaptive Thresholding</title>
<p>Entropy is used to to measure the uncertainty of a model or dataset [<xref ref-type="bibr" rid="ref-25">25</xref>]. Given a probability distribution <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> over <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>i</mml:mi></mml:math></inline-formula> classes, the entropy as shown in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> calculates how confused the model is. If the distribution is sharp for example, one class has probability close to 1, the entropy is low, indicating high confidence. While, if the probabilities are not sharp, the entropy is higher, meaning the model is uncertain.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the predicted probability for class <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>i</mml:mi></mml:math></inline-formula>.</p>
<p>To decide when to trust the base model and when to move to a more expensive option (such as the Teacher Probe), an adaptive threshold T is used as in <xref ref-type="disp-formula" rid="eqn-8">(8)</xref>. This threshold is computed using the <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>z</mml:mi></mml:math></inline-formula>-score of the entropy values across a batch. Specifically, the dynamic threshold is set to the mean entropy <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0B5;</mml:mi></mml:mrow><mml:mi>H</mml:mi></mml:msub></mml:math></inline-formula> plus a multiple <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula> of the standard deviation <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>H</mml:mi></mml:msub></mml:math></inline-formula>.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0B5;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0B5;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> are the mean and standard deviation of entropy values, and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula> is a scaling factor.</p>
</sec>
<sec id="s3_2_6">
<label>3.2.6</label>
<title>Action Definitions</title>
<p>The agent has four possible actions, each representing a different trade-off between accuracy and computational cost.</p>
<p><bold>Reduced Model:</bold> The first action is using random forest that is trained only on the top-<italic>K</italic> features selected using the MI cumfrac policy. By focusing on the most informative features, the reduced model is faster and simpler, with no added cost. However, it may sacrifice some accuracy compared to the full model. Reduced Features action can be represented as in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>red</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p><bold>Base Model:</bold> This is the second action as defined in <xref ref-type="disp-formula" rid="eqn-10">(10)</xref>, which uses the calibrated random forest model trained on all <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>d</mml:mi></mml:math></inline-formula> features. It provides strong predictive performance since no information is discarded. Importantly, this option is considered cost-free, making it a reliable baseline.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p><bold>LLM Refinement:</bold> In this third action, the Base Model initially performs the prediction, and its confidence is quantified using the entropy of the predicted probability distribution, denoted as <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. If the entropy is below the adaptive threshold <italic>T</italic> (indicating low uncertainty), the prediction from the Base Model is accepted directly. Conversely, if the entropy exceeds the threshold (indicating high uncertainty), the system activates the Teacher Probe, a logistic regression classifier trained on ELECTRA embeddings. This mechanism refines the decision for ambiguous samples by leveraging deep contextual representations. The decision rule for this action is formally expressed in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>probe</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><bold>Escalate to Analyst:</bold> The fourth action is used when predictions are highly uncertain. The top <italic>K</italic> samples with the highest entropy values are forwarded to an analyst. In practice, this means that the system does not assign a predicted label for these cases and instead marks them as unresolved pending manual verification. As shown in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>, if the sample <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>i</mml:mi></mml:math></inline-formula> is not in the Top-K uncertain samples <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2209;</mml:mo><mml:mtext>TopK</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the prediction is made by the base model <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. However, if sample <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>i</mml:mi></mml:math></inline-formula> is in the Top-K uncertain samples <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mtext>TopK</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the output is assigned an UNRESOLVED label to reflect escalation to the analyst.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2209;</mml:mo><mml:mrow><mml:mtext>TopK</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>UNRESOLVED</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mtext>TopK</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_2_7">
<label>3.2.7</label>
<title>Adaptive Decision Agent Workflow</title>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the workflow of the adaptive decision approach, which controls the dynamic action selection process during inference. The input batch, which comes from the test set, is passed to MI model to perform feature selection using a cumulative fraction policy to identify the most informative subset of features. This step ensures that redundant or less relevant features are excluded.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Adaptive decision controller.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-5.tif"/>
</fig>
<p>The selected features are then submitted to the Base Model to produce class predictions along with confidence scores. These scores are checked using an adaptive entropy method to decide how reliable each prediction is.</p>
<p>As the figure shows, the model dynamically selects one of four decision pathways based on the evaluated entropy level. For instance, in case of low-entropy, the Reduced Model path is used after selecting the top-<italic>K</italic> most informative features. It is worth noting that this option is computationally efficient and suitable when predictions are already confident. The system switches to the Base Model when the entropy level increased to operate on the complete feature set. This enhances reliability in uncertain cases. The LLM Refinement process is activated by the framework when the entropy rises. This uses deep contextual representations together with logistic regression and ELECTRA embeddings to enhance predictions.</p>
<p>Finally, the agent initiates the Escalate to Analyst pathway when the entropy stays over a certain threshold. For manual verification, it sends the top-<italic>K</italic> highly uncertain instances. The process concludes at the Accept Prediction module where the result of the action selected is used to make the final decision of the system. This adaptive control improves accuracy, interpretability and computational efficiency by intelligent allocation of computational resources based on the uncertainty of each input sample.</p>
<p>Overall, these four steps allow the model to effectively balance computational costs toward predictive performance. In order to maintain efficiency without reducing accuracy, the framework is encouraged by the cost structure to keep the more costly choices (Actions 3 and 4) to only highly uncertain samples.</p>
</sec>
<sec id="s3_2_8">
<label>3.2.8</label>
<title>Reward Function</title>
<p>The reward function is designed to balance prediction accuracy with the computational or operational penalty of each action. The reward calculates the difference between the performance score of the model and the penalty of the chosen action as defined in <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>. We use accuracy as the scoring function that measures the proportion of correctly classified samples. Actions that achieve high accuracy with minimal penalty will receive higher rewards. However, actions that are more expensive or less accurate will produce lower rewards. For instance, there is no penalty associated with the usage of the Reduced or Base models, therefore the accuracy is equivalent to the reward. By contrast, utilizing the Teacher Probe or forwarding the Top-K uncertain samples to the analyst incurs additional penalties, which reduces the reward even when accuracy is high. This design ensures that the system learns to prefer strategies that are not only correct but also efficient, reflecting a realistic trade-off between accuracy and resource usage.
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Score</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>Penalty</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_2_9">
<label>3.2.9</label>
<title>LinUCB Algorithm</title>
<p>Each action in the LinUCB algorithm keeps its own parameters, which are represented as a vector <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>b</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and a matrix <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>A</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula>. The matrix <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>A</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> captures how much information has been gathered about the action based on past feature contexts, while <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>b</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula> accumulates the rewards observed [<xref ref-type="bibr" rid="ref-39">39</xref>]. Using these values, the algorithm estimates the parameter vector <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi>b</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula>, which can be interpreted as the model&#x2019;s best guess about the relationship between the context and the expected reward for action <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>a</mml:mi></mml:math></inline-formula>. When deciding which action to take, LinUCB computes an upper confidence bound (UCB) score for each action as written in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>a</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msqrt><mml:msup><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:msubsup><mml:mi>A</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mi>x</mml:mi></mml:msqrt><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The first term, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>a</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:mi>x</mml:mi></mml:math></inline-formula>, represents the predicted reward for the action given the current context (features) <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>x</mml:mi></mml:math></inline-formula>. The second term adds an exploration bonus, scaled by the parameter <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>. This bonus is larger when the algorithm has less certainty about the action (i.e., when <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>A</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is large), encouraging exploration of under-tested actions.</p>
<p>The action with the highest score is then selected as shown in <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mi>a</mml:mi></mml:munder><mml:msub><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the estimated upper confidence bound for action <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>a</mml:mi></mml:math></inline-formula> given context <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>. This means the agent selects the action that maximizes the combination of the predicted reward and the uncertainty bonus.</p>
<p>Once the reward <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is observed, the algorithm updates its internal parameters for the chosen action as in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref>:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi>A</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>b</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>These updates allow LinUCB to continuously refine its confidence estimates, gradually learning which actions yield higher rewards under different contexts.</p>
</sec>
<sec id="s3_2_10">
<label>3.2.10</label>
<title>Datasets</title>
<p>In this research, we evaluated our proposed adaptive intrusion detection framework on three benchmark datasets including RT-IoT-2022 [<xref ref-type="bibr" rid="ref-40">40</xref>], CIC-IoT-2023 [<xref ref-type="bibr" rid="ref-41">41</xref>] and CIC-IoMT-2024 [<xref ref-type="bibr" rid="ref-42">42</xref>]. We chose these datasets because they represent modern network behavior, diverse device types and unseen attacks in IoT and IoMT environments.</p>
<p>The first dataset used is RT-IoT-2022 and contains attack and normal traffic from a real-time IoT setup. It was designed for intrusion and anomaly detection research and provides 83 flow-based features such as packet statistics, payload details, timing and TCP flags [<xref ref-type="bibr" rid="ref-40">40</xref>]. The dataset has nine attack types that are SYN flooding, Slowloris distributed denial of service (DDoS), ARP poisoning and Nmap scans as well as three types of benign traffic. This variety makes this dataset a strong test for classification methods, particularly given the limited resources common in IoT systems.</p>
<p>The second dataset is CIC-IoT-2023 that also supports research on intrusion and anomaly detection and was developed using a network of 105 real IoT devices [<xref ref-type="bibr" rid="ref-41">41</xref>]. It has 33 attack types from compromised devices that are grouped into seven categories namely DDoS, DoS, reconnaissance, web attacks, brute force, spoofing and Mirai botnet variants. The dataset provides 46 features describing network flows, packet statistics and timing. This range enables us to evaluate our detection methods on a variety of new IoT attacks and takes various devices and protocols into account.</p>
<p>The third dataset is CIC-IoMT-2024 that was collected from a simulated healthcare network with typical medical IoT devices and protocols [<xref ref-type="bibr" rid="ref-42">42</xref>]. The dataset includes both benign and malicious traffic such as resource loss attacks, illegal access, and probing. It provides 44 flow-based features on protocol use, packet statistics and timing. This setup lets us test our detection methods in healthcare where device variety and the need for constant service add extra challenges.</p>
<p>Overall, the use of the aforementioned datasets to evaluate our intrusion detection framework helps to demonstrate its effectiveness and adaptability to diverse IoT and IoMT environments. The variety of attack types, device behaviors and protocol characteristics in RT-IoT-2022, CIC-IoT-2023, and CIC-IoMT-2024 enables a thorough evaluation of the classification performance of our model under realistic conditions. <xref ref-type="table" rid="table-2">Table 2</xref> summarizes the key characteristics of the datasets used in this study, highlighting their differences in environment, feature dimensionality, attack diversity, and application domain.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of datasets used in this study.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Environment</th>
<th># Features</th>
<th>Attack Diversity</th>
<th>Traffic Representation</th>
<th>Application Domain</th>
</tr>
</thead>
<tbody>
<tr>
<td>RT-IoT-2022</td>
<td>Real-time IoT network</td>
<td>83</td>
<td>Multiple DoS, scanning, poisoning, and benign traffic types</td>
<td>Flow-based</td>
<td>General IoT security</td>
</tr>
<tr>
<td>CIC-IoT-2023</td>
<td>Large-scale IoT testbed</td>
<td>46</td>
<td>33 attacks grouped into 7 categories (DDoS, DoS, reconnaissance, web, brute force, spoofing, Mirai)</td>
<td>Flow-based</td>
<td>Heterogeneous IoT networks</td>
</tr>
<tr>
<td>CIC-IoMT-2024</td>
<td>Simulated healthcare network</td>
<td>44</td>
<td>Multiple resource abuse, illegal access, and probing attacks</td>
<td>Flow-based</td>
<td>Medical IoT (IoMT)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<p>In this section, we present and discuss the experimental results of our proposed adaptive intrusion detection framework. We used three benchmark IoT datasets to evaluate the proposed framework that are RT-IoT 2022, CIC-IoT 2023 and CIC-IoMT 2024 datasets. First, we show the MI distribution of each dataset to highlight the most informative features selected for analysis. Next, we examine the usage proportions of the four decision pathways namely Reduced Model, Base Model, LLM Refinement and Escalate to Analyst to understand the way that the adaptive policy allocates computational effort. After that, we evaluate the per-action performance of the proposed framework by computing the well known evaluation metrics and calculating the corresponding reward for each action. Finally, we analyze the trade-off between accuracy and computational cost in the four decision actions of the proposed framework.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Evaluation on RT-IoT-2022 Dataset</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Mutual Information (MI) Result on RT-IoT-2022 Dataset</title>
<p>MI was employed to estimate the relevance of each feature to the target class and to guide dimensionality reduction. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows the ranked MI scores for all features in the RT-IoT-2022 dataset. Each bar represents a feature&#x2019;s MI value, where higher scores indicate stronger dependence on the class label. Using the cumulative-fraction policy with a threshold of <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula>, the cutoff indicated by the red dashed line retains K &#x003D; 58 out of d &#x003D; 82 features. These 58 features account for at least 90% of the total MI, while the remaining variables contribute only marginal information and are removed. This reduction eliminates redundant or low-impact features and lowers computational cost by focusing the model on the most informative subset. To justify <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula>, we evaluated thresholds <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> &#x003D; 0.70, 0.80, 0.90, 0.95, and 0.99 as presented in <xref ref-type="table" rid="table-3">Table 3</xref>. The MI values drop quickly, so a small number of features carry most of the useful information. At <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> &#x003D; 0.70 and <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> &#x003D; 0.80, only 41 and 49 features are needed to cover most of the MI. The reference threshold <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> &#x003D; 0.90 selects 58 features, providing a stable balance between information coverage and dimensionality. Increasing the threshold beyond this point brings only minor benefit. For example, <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> &#x003D; 0.99 increases the count to 72 features, but the extra features add almost no meaningful information.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>MI distribution of RT-IoT-2022 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-6.tif"/>
</fig><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Effect of different cumulative MI thresholds on the number of selected features for the RT-IoT-2022 dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Threshold <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula></th>
<th>Selected Features (<italic>K</italic>)</th>
<th>Percentage of Features</th>
<th>Jaccard vs. <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>0.70</td>
<td>41</td>
<td>48.81%</td>
<td>0.707</td>
</tr>
<tr>
<td>0.80</td>
<td>49</td>
<td>58.33%</td>
<td>0.845</td>
</tr>
<tr>
<td>0.90</td>
<td>58</td>
<td>69.05%</td>
<td>1.000</td>
</tr>
<tr>
<td>0.95</td>
<td>64</td>
<td>76.19%</td>
<td>0.906</td>
</tr>
<tr>
<td>0.99</td>
<td>72</td>
<td>85.71%</td>
<td>0.806</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Jaccard similarity analysis also supports this choice. When using <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.70</mml:mn></mml:math></inline-formula> or <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.80</mml:mn></mml:math></inline-formula>, the selected features are simply smaller versions of the same important set obtained at <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula>. When increasing the threshold to <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.95</mml:mn></mml:math></inline-formula> or <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.99</mml:mn></mml:math></inline-formula>, the extra features added are mostly low-importance variables that do not change the core set. In short, <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula> preserves the meaningful features while avoiding unnecessary, low-impact variables.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Actions Usage and Performance Analysis on RT-IoT-2022 Dataset</title>
<p>The percentage of each action is calculated to determine action usage in the RT-IoT-2022 dataset. This helps identify which action dominates when the adaptive framework is used. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7a</xref>, the LinUCB controller can dynamically select among four possible inference actions based on prediction confidence and computational cost. There are four actions as mentioned in the methodology section named Reduced Model, Base Model, LLM Refinement, and Escalate to Analyst. The overall results demonstrate that the Reduced Model and Base Model together recorded <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mn>50</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of all decisions. The LLM Refinement was used in <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mn>27.55</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of the total actions. The Escalate to Analyst accounted for <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>22.45</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, which is desirable since relying less on this costly action improves overall efficiency.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Adaptive IDS behavior on the RT-IoT-2022 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-7.tif"/>
</fig>
<p>Therefore, the proposed framework reduces computation, as half of the actions involve no extra processing cost. The results also show that the Reduced Model being the most frequently chosen <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mn>27.93</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, which reflects the system&#x2019;s preference for computational efficiency inference when confidence is high.</p>
<p>The per-action performance on the RT-IoT-2022 dataset using 5-fold average is shown in <xref ref-type="table" rid="table-4">Table 4</xref>. The Reduced Model and Base Model achieve the best overall balance, with high accuracy at 0.9992 and 0.9986, zero penalty, and the lowest inference time per sample at 0.105 and 0.107 ms. These results confirm that lightweight actions handle most samples efficiently without compromising accuracy. For uncertain cases, the framework triggers LLM Refinement. ELECTRA reaches 0.9962 accuracy and 0.9981 F1-score, while DistilBERT achieves 0.9973 accuracy and 0.9986 F1-score. The two models deliver comparable refinement performance, with only minor differences in accuracy and F1, and both incur the same five percent penalty and inference time of about 0.238 ms per sample. Escalate to Analyst achieves the highest accuracy at 0.9999 and F1-score at 0.9948 but carries the highest penalty (10%) and therefore the lowest reward among the correct actions. Overall, the results show that the framework prioritizes fast, zero-penalty actions and uses costly steps only when uncertainty is high.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Per-action performance on the RT-IoT-2022 dataset (5-fold average).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Action</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>Penalty</th>
<th>Reward</th>
<th>Inference Time (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Reduced Model</td>
<td>0.9992</td>
<td>0.9995</td>
<td>0%</td>
<td>0.9992</td>
<td>0.105</td>
</tr>
<tr>
<td>Base Model</td>
<td>0.9986</td>
<td>0.9990</td>
<td>0%</td>
<td>0.9986</td>
<td>0.107</td>
</tr>
<tr>
<td>LLM Refinement (ELECTRA)</td>
<td>0.9962</td>
<td>0.9981</td>
<td>5%</td>
<td>0.9462</td>
<td>0.238</td>
</tr>
<tr>
<td>LLM Refinement (DistilBERT)</td>
<td>0.9973</td>
<td>0.9986</td>
<td>5%</td>
<td>0.9473</td>
<td>0.238</td>
</tr>
<tr>
<td>Escalate to Analyst</td>
<td>0.9999</td>
<td>1.0000</td>
<td>10%</td>
<td>0.8999</td>
<td>0.243</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It is worth mentioning that each experiment uses 5-fold cross-validation. We used 5-fold cross-validation to balance statistical robustness and computational efficiency, as higher fold counts (e.g., 10-fold) substantially increase training and evaluation cost and are commonly reported to provide only marginal performance gains when sufficient data is available.</p>
<p>In each fold, the dataset is first split into 80% training and 20% testing. The training portion is then further divided into 80% training and 20% validation, resulting in final proportions of 64% training, 16% validation, and 20% testing. The validation set is used for model calibration, feature selection, and adaptive decision tuning, while the test set remains completely unseen until final evaluation.</p>
<p>Different split ratios mainly affect the bias&#x2013;variance trade-off. Increasing the training portion may slightly improve fitting but reduces the reliability of validation and test estimates, whereas smaller training sets increase variance, particularly in imbalanced IoT traffic. The chosen configuration yields stable and consistent performance across folds.</p>
<p>We also analyzed the trade-off between accuracy and the action penalty across the four decision actions of the proposed framework on the RT-IoT-2022 dataset. The <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis represents the assigned penalty, while the <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis shows the average accuracy, as illustrated in <xref ref-type="fig" rid="fig-7">Fig. 7b</xref>. The optimal region appears on the far left, where high accuracy aligns with minimal penalty. Conversely, the far-right region reflects actions that achieve high accuracy but incur higher penalties.The Reduced Model and Base Model appear in the far-left region and achieve near-perfect accuracy with zero penalty, confirming their efficiency for standard predictions. The LLM Refinement action introduces a moderate penalty with a slight decrease in accuracy, and is triggered only for uncertain samples. The Escalate to Analyst action, which offers the highest accuracy, carries the highest penalty and appears in the upper-right corner, reflecting the cost of manual verification. Overall, the figure shows that the framework balances accuracy and efficiency. Zero-penalty actions handle most inputs reliably, while higher-penalty actions are used only when uncertainty is high.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Evaluation on CIC IoT 2023 Dataset</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Mutual Information (MI) Result on CIC IoT-2023 Dataset</title>
<p>The MI distribution for the CIC-IoT-2023 dataset, which contains <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>46</mml:mn></mml:math></inline-formula> features. At this point, <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>23</mml:mn></mml:math></inline-formula> features are retained, while the remaining 23 features, which contribute only marginally to the predictive signal, are excluded. This selection balances dimensionality reduction with information retention, ensuring that the model leverages the most informative subset of features while discarding less relevant ones to reduce complexity and noise.</p>
<p>The MI distribution of the CICIoT-2023 dataset reveals that among the total <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>46</mml:mn></mml:math></inline-formula> extracted features, the informative content is highly concentrated within the top-ranked subset. As presented in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the features are arranged in descending order of their MI scores, showing a rapid decline after the leading features. This indicates that only a limited number of them significantly contribute to class discrimination. Using the cumulative fraction policy with a threshold, a total of <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>23</mml:mn></mml:math></inline-formula> features are retained, capturing 90% of the total MI, while the remaining features are discarded due to their marginal relevance.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>MI distribution of CIC IoT 2023 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Actions Usage and Performance Analysis on CIC IoT 2023 Dataset</title>
<p>The distribution of action usage within the proposed framework on the CIC-IoT-2023 dataset is analyzed in this research. <xref ref-type="fig" rid="fig-9">Fig. 9a</xref> illustrates that the Reduced Model and Base Model account for around 37.52% and 36.70% of all forecasts, respectively and therefore lead the decision-making process. This highlights the preference of the framework for low-cost inference pathways, which confirms that most network traffic can be accurately classified using lightweight or full-feature models without the need for additional computation. The LLM Refinement action, which involves deeper semantic processing through ELECTRA embeddings and logistic regression, is used for only 9.72% of samples. This indicates that uncertainty levels in the baseline predictions are relatively low and the refinement mechanism is triggered selectively for uncertain cases. Finally, the Escalate to Analyst action was used for 16.06% of the cases, which represents that a small portion of samples had the highest prediction entropy. Overall, the adaptive framework successfully used the Reduced and Base Models for 74.22% of the cases and only about 25.78% were handled by LLM Refinement and Escalate to Analyst. This emphasizes that the model primarily relies on low-cost actions.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Adaptive IDS behavior on the CIC-IoT-2023 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-9.tif"/>
</fig>
<p>On the CIC-IoT-2023 dataset,the Reduced Model and Base Model deliver almost identical performance with accuracies of 0.9417 and 0.9416 and very low inference time in the range of 0.1020 to 0.1110 ms as reported in <xref ref-type="table" rid="table-5">Table 5</xref>. This indicates that the framework maintains strong predictive capability under both lightweight and full-feature inference without any penalty. The LLM Refinement actions powered by ELECTRA and DistilBERT produce closely matched results with accuracies near 0.916 to 0.919 and weighted F1-scores near 0.951 to 0.953. Both actions incur a five percent penalty, leading to rewards in the range of 0.865 to 0.869, and introduce a slightly higher latency of about 0.24 ms. The Escalate to Analyst action achieves the highest accuracy of 0.9459 and a weighted F1-score of 0.9698, though its ten percent penalty reduces the reward to 0.8459. Overall, the controller maintains a stable detection performance while using higher-cost actions only when uncertainty makes them necessary.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Per-action performance on the CIC-IoT-2023 Dataset (5-fold average).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Action</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>Penalty</th>
<th>Reward</th>
<th>Inference Time (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Reduced Model</td>
<td>0.9417</td>
<td>0.9666</td>
<td>0%</td>
<td>0.9417</td>
<td>0.102</td>
</tr>
<tr>
<td>Base Model</td>
<td>0.9416</td>
<td>0.9663</td>
<td>0%</td>
<td>0.9416</td>
<td>0.111</td>
</tr>
<tr>
<td>LLM Refinement (ELECTRA)</td>
<td>0.9155</td>
<td>0.9509</td>
<td>5%</td>
<td>0.8655</td>
<td>0.245</td>
</tr>
<tr>
<td>LLM Refinement (DistilBERT)</td>
<td>0.9189</td>
<td>0.9532</td>
<td>5%</td>
<td>0.8689</td>
<td>0.239</td>
</tr>
<tr>
<td>Escalate to Analyst</td>
<td>0.9459</td>
<td>0.9698</td>
<td>10%</td>
<td>0.8459</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-9">Fig. 9b</xref> shows the trade-off between accuracy and action penalty on the CIC-IoT-2023 dataset. The Reduced Model and Base Model appear in the top-left region, reaching about 0.94 accuracy with zero penalty, confirming their efficiency for most samples. LLM Refinement lies in the middle with a 5% penalty and lower accuracy (around 0.92), reflecting its selective use for uncertain cases. Escalate to Analyst action achieves accuracy around 0.945 but at the highest penalty (10%), indicating that escalation is reserved for the most difficult, high-risk samples.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation on CIC IoMT 2024</title>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Mutual Information (MI) Result on CIC IoMT 2024 Dataset</title>
<p><xref ref-type="fig" rid="fig-10">Fig. 10</xref> presents the MI distribution for the CIC-IoMT-2024 dataset, which has <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>45</mml:mn></mml:math></inline-formula> features. The blue bars correspond to MI scores and the solid line represents cumulative MI. The use of the <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.90</mml:mn></mml:math></inline-formula> cutoff, <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>25</mml:mn></mml:math></inline-formula> features are selected, as shown by the vertical dashed red line. The retained subset accounts for at least <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mn>90</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of the overall information, while 20 less-informative features are removed. This selection strategy reduces computational complexity while preserving the predictive richness of the dataset.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>MI distribution of CIC IoMT 2024 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-10.tif"/>
</fig>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Actions Usage and Performance Analysis on CIC IoMT 2024 Dataset</title>
<p>The distribution of action usage for the proposed framework in the CIC-IoMT-2024 dataset is shown in <xref ref-type="fig" rid="fig-11">Fig. 11a</xref>. The results show that the Reduced Model and Base Model together dominate the adaptive decisions as they accounted for 68.75% of all predictions. This reflects the framework&#x2019;s strong preference for computationally efficient because most samples are confidently classified without the need for additional processing. The LLM Refinement action was triggered for 15.48% of instances, which represents uncertain samples where the confidence of the primary models falls below the adaptive entropy threshold. Meanwhile, the Escalate to Analyst action was invoked for 15.77% of the inputs, which corresponds to the most complex samples with the highest uncertainty levels. These cases are routed for expert-level or external analysis to ensure accuracy and reliability. The near-balanced distribution between low-cost and high-cost actions confirms that the adaptive policy dynamically optimizes the trade-off between efficiency and precision, which ensures that the system maintains high detection accuracy while minimizing unnecessary computation.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Adaptive IDS behavior on the CIC-IoMT-2024 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76413-fig-11.tif"/>
</fig>
<p>The framework was also evaluated on CIC-IoMT-2024 dataset using 5-fold cross-validation, and the averaged per-action results as shown in <xref ref-type="table" rid="table-6">Table 6</xref>. The Reduced Model and Base Model actions provide the strongest overall balance, reaching accuracies of 0.9971 and 0.9976 with weighted F1-scores of 0.9984 and 0.9987, zero penalty, and small latency of about 0.105 to 0.109 ms.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Per-action performance on the CIC-IoMT-2024 dataset (5-fold average).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Action</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>Penalty</th>
<th>Reward</th>
<th>Inference Time (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Reduced Model</td>
<td>0.9971</td>
<td>0.9984</td>
<td>0%</td>
<td>0.9971</td>
<td>0.105</td>
</tr>
<tr>
<td>Base Model</td>
<td>0.9976</td>
<td>0.9987</td>
<td>0%</td>
<td>0.9976</td>
<td>0.109</td>
</tr>
<tr>
<td>LLM Refinement (ELECTRA)</td>
<td>0.9842</td>
<td>0.9913</td>
<td>5%</td>
<td>0.9342</td>
<td>0.239</td>
</tr>
<tr>
<td>LLM Refinement (DistilBERT)</td>
<td>0.9835</td>
<td>0.9909</td>
<td>5%</td>
<td>0.9335</td>
<td>0.240</td>
</tr>
<tr>
<td>Escalate to Analyst</td>
<td>0.9915</td>
<td>0.9955</td>
<td>10%</td>
<td>0.8915</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The LLM Refinement action achieves accuracies of 0.9842 with ELECTRA and 0.9835 with DistilBERT, with corresponding weighted F1-scores of 0.9909 and 0.9909. Both models deliver comparable accuracy, and the five percent penalty reduces the rewards to 0.9342 and 0.9335. Latency increases slightly due to the refinement step, reaching roughly 0.239 to 0.240 ms per sample. The Escalate to Analyst action attains 0.9915 accuracy and 0.8999 F1-score but carries the highest ten percent penalty.</p>
<p>The relationship between accuracy and action penalty on the CIC IoMT 2024 dataset is shown in <xref ref-type="fig" rid="fig-11">Fig. 11b</xref>. The Reduced Model and Base Model actions appear at the zero penalty region with accuracies of 0.9971 and 0.9976. They deliver the strongest efficiency to accuracy value and remain the dominant choices for most traffic. The LLM Refinement action carries a five percent penalty and achieves an accuracy of about 0.984. This reduction is expected because the action is triggered only for high-uncertainty samples, which are naturally more difficult to classify. The Escalate to Analyst action appears at the ten percent penalty level and reaches an accuracy of about 0.991. While not the highest, it provides the most reliable option for the hardest cases because it represents human verification.</p>
<p>To further validate the findings of our research, we compare the classification accuracy and F1-score of the proposed framework with existing intrusion detection approaches, as summarized in <xref ref-type="table" rid="table-7">Table 7</xref>. Unlike prior studies that report a single static model, the proposed framework is adaptive and operates through multiple inference paths. Depending on prediction uncertainty and computational constraints, the controller dynamically selects between lightweight (Reduced Model), full feature (Base Model), refinement, or escalation strategies. Therefore, <xref ref-type="table" rid="table-7">Table 7</xref> reports the performance of the proposed framework under its two primary zero-penalty inference paths (Reduced and Base Models), which represent the dominant operating modes during deployment.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of the proposed adaptive IDS framework with existing intrusion detection approaches on IoT benchmark datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>Dataset</th>
<th>Method</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>Computational Cost</th>
</tr>
</thead>
<tbody>
<tr>
<td>Hafid et al. (2025)</td>
<td>CIC-IoMT-2024</td>
<td>XGBoost&#x002B;Logistic Regression (Late Fusion)</td>
<td>97%</td>
<td>98%</td>
<td>Bandwidth reduced by 73%; latency &#x003C; 200 ms</td>
</tr>
<tr>
<td>Houichi et al. (2025)</td>
<td>CIC-IoT-2023</td>
<td>ML-Based IDS (KNN, RF, AdaBoost, SVM)</td>
<td>99%</td>
<td>91%</td>
<td>&#x2014;</td>
</tr>
<tr>
<td>Almohaimeed et al. (2024)</td>
<td>RT-IoT-2022</td>
<td>Feature Selection &#x002B; MLP</td>
<td>96%</td>
<td>92%</td>
<td>Runtime reduced by 66.4%</td>
</tr>
<tr>
<td>Prasad et al. (2025)</td>
<td>CIC-IoT-2023</td>
<td>RegNet &#x002B; FBNet &#x002B; CSMCR Balancing</td>
<td>98%</td>
<td>98%</td>
<td>Training time reduced by 53%</td>
</tr>
<tr>
<td>Mahdi et al. (2025)</td>
<td>CIC-IoT-2023</td>
<td>Incremental Learning &#x002B; Blockchain</td>
<td>99.53%</td>
<td>&#x2013;</td>
<td>Encryption overhead (0.0001&#x2013;0.0003 s); transaction time 2&#x2013;3 s</td>
</tr>
<tr>
<td rowspan="6"><bold>Proposed Framework (ours)</bold></td>  
<td>RT-IoT-2022</td>
<td>Reduced Model</td>
<td>99.92%</td>
<td>99.95%</td>
<td>0.105 ms</td>
</tr>
<tr>
<td>CIC-IoT-2023</td>
<td>Reduced Model</td>
<td>94.17%</td>
<td>96.66%</td>
<td>0.102 ms</td>
</tr>
<tr>
<td>CIC-IoMT-2024</td>
<td>Reduced Model</td>
<td>99.74%</td>
<td>99.84%</td>
<td>0.105 ms</td>
</tr>
<tr>
<td>RT-IoT-2022</td>
<td>Base Model</td>
<td>99.86%</td>
<td>99.90%</td>
<td>0.107 ms</td>
</tr>
<tr>
<td>CIC-IoT-2023</td>
<td>Base Model</td>
<td>94.16%</td>
<td>96.63%</td>
<td>0.111 ms</td>
</tr>
<tr>
<td>CIC-IoMT-2024</td>
<td>Base Model</td>
<td>99.76%</td>
<td>99.87%</td>
<td>0.109 ms</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Hafid et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] reported strong performance on CIC-IoMT-2024 using a late-fusion XGBoost and Logistic Regression model, achieving 97% accuracy and 98% F1-score. While both their approach and the proposed framework incorporate adaptivity, their interpretability relies on SHAP values and their evaluation follows a static cost-performance analysis. Their method reduces bandwidth usage by 73% with detection latency below 200 ms, whereas our framework achieves substantially lower inference time in the sub-millisecond range.</p>
<p>Houichi et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] reported 99% accuracy and a lower F1-score of 91% on CIC-IoT-2023 using traditional ML classifiers. However, their evaluation was limited to a single dataset and lacked adaptive decision control, limiting generalizability.</p>
<p>Prasad et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] achieved 98% accuracy and a 98% F1-score on CIC-IoT-2023 using a DL model with cosine-similarity&#x2013;based balancing. Their balancing method mitigated the inherent class imbalance of the dataset. In addition, their deep-learning pipeline followed a static inference process with no adaptive or cost-aware control. In contrast, the proposed adaptive framework maintains consistently strong performance across three benchmark datasets (RT-IoT-2022, CIC-IoT-2023, CIC-IoMT-2024) and introduces a LinUCB-based decision layer that dynamically balances accuracy and penalty under varying network conditions. They reported a 53% reduction in training time, whereas our framework delivers fast inference with an average cost near 0.1 ms per sample.</p>
<p>Almohaimeed and Albalwy [<xref ref-type="bibr" rid="ref-24">24</xref>] applied multiple feature selection methods to RT-IoT-2022 followed by an MLP classifier, achieving 96% accuracy and a 92% F1-score. While effective at reducing dimensionality, their approach does not incorporate adaptive or cost-aware decision logic. Our framework surpasses these results, achieving up to 99.92% accuracy and 99.95% F1-score on RT-IoT-2022 using the Reduced Model. Their method reduced runtime by 66.4%, whereas our framework additionally achieves sub-millisecond inference with higher predictive performance.</p>
<p>Finally, Mahdi et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] used the CIC-IoT-2023 dataset, which is inherently imbalanced, with attack traffic dominating normal flows. The authors did not apply any explicit data-balancing method, instead relying on the incremental learning process to adapt to changing distributions. Consequently, the reported 99.53% accuracy may be influenced by class imbalance, as no F1-score or weighted metrics were provided. Their system introduces blockchain-related overhead (0.0001 s&#x2013;0.0003 s per operation and 2&#x2013;3 s per transaction), whereas the proposed IDS maintains inference latency near 0.1 ms.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>In conclusion, we proposed an adaptive intrusion detection framework designed for IoT environments. This asymmetrical nature of IoT systems involves limited computational resources, high feature dimensionality, and prediction uncertainty. The proposed framework has the ability to intelligently balance detection accuracy and computational cost. This is accomplished by applying MI as a feature selection model, which effectively reduces the dimensionality of the data and keeps only the most informative features while eliminating redundancy. The adaptive decision dynamically selects between multiple inference paths, which ensures that additional computation is used only when uncertainty is high. Performance evaluation on the RT-IoT-2022, CIC-IoT-2023, and CIC-IoMT-2024 datasets indicates that the framework achieves F1-scores of 99.92%, 96.66%, and 99.84%, respectively, with an average inference time of approximately 0.105 ms per sample. A limitation of this work is that it focuses on binary intrusion classification. This choice was made to clearly study the balance between detection accuracy and computational cost. In future work, the framework will be extended to multiclass intrusion detection to support more detailed attack classification. In addition, future studies will evaluate the framework in real IoT environments and improve the adaptive controller by considering additional factors such as energy consumption, network latency, and device diversity.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This project was funded by the Deanship of Scientific Research (DSR) at King Abdulaziz University, Jeddah, Saudi Arabia under grant no. (IPP: 753-611-2025). The authors, therefore, acknowledge with thanks DSR for technical and financial support.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization, Abdulaziz A. Alsulami, Badraddin Alturki and Ahmad J. Tayeb; methodology, Abdulaziz A. Alsulami and Rayan A. Alsemmeari; software, Abdulaziz A. Alsulami; validation, Abdulaziz A. Alsulami, Rayan A. Alsemmeari and Ahmad J. Tayeb; formal analysis, Ahmad J. Tayeb; investigation, Badraddin Alturki; resources, Abdulaziz A. Alsulami; data curation, Abdulaziz A. Alsulami; writing&#x2014;original draft preparation, Abdulaziz A. Alsulami, Badraddin Alturki and Ahmad J. Tayeb; writing&#x2014;review and editing, Rayan A. Alsemmeari, Raed Alsini and Badraddin Alturki; visualization, Badraddin Alturki; supervision, Abdulaziz A. Alsulami; project administration, Badraddin Alturki, Ahmad J. Tayeb and Raed Alsini; funding acquisition, Abdulaziz A. Alsulami. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The benchmark datasets used in this study (RT-IoT2022, CIC-IoT2023, and CIC-IoMT2024) are publicly available from their respective sources.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>This study does not involve human participants, animal subjects, or the use of personal or sensitive data. Therefore, ethical approval was not required.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sadeghi-Niaraki</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Internet of Thing (IoT) review of review: bibliometric overview since its foundation</article-title>. <source>Future Gener Comput Syst</source>. <year>2023</year>;<volume>143</volume>(<issue>5</issue>):<fpage>361</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2023.01.016</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmetoglu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Che Cob</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>N</given-names></string-name></person-group>. <article-title>A systematic review of Internet of Things adoption in organizations: taxonomy, benefits, challenges and critical factors</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>9</issue>):<fpage>4117</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app12094117</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jouhari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Saeed</surname> <given-names>N</given-names></string-name>, <string-name><surname>Alouini</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Amhoud</surname> <given-names>EM</given-names></string-name></person-group>. <article-title>A survey on scalable LoRaWAN for massive IoT: recent advances, potentials, and challenges</article-title>. <source>IEEE Commun Surv Tutor</source>. <year>2023</year>;<volume>25</volume>(<issue>3</issue>):<fpage>1841</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2023.3274934</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Vailshery</surname> <given-names>LS</given-names></string-name></person-group>. <article-title>Number of IoT connections worldwide 2022&#x2013;2034</article-title>. <year>2025</year> <comment>[cited 2025 Oct 10]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.statista.com/statistics/1183457/iot-connected-devices-worldwide/">https://www.statista.com/statistics/1183457/iot-connected-devices-worldwide/</ext-link>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Vailshery</surname> <given-names>LS</given-names></string-name></person-group>. <article-title>IoT global annual revenue 2020&#x2013;2034</article-title>. <year>2025</year> <comment>[cited 2025 Oct 10]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.statista.com/statistics/1194709/iot-revenue-worldwide/">https://www.statista.com/statistics/1194709/iot-revenue-worldwide/</ext-link>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Othman</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yusuf</surname> <given-names>LM</given-names></string-name></person-group>. <article-title>Optimizing IoT intrusion detection system: feature selection versus feature extraction in machine learning</article-title>. <source>J Big Data</source>. <year>2024</year>;<volume>11</volume>(<issue>1</issue>):<fpage>36</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-024-00892-y</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Merlino</surname> <given-names>V</given-names></string-name>, <string-name><surname>Allegra</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Energy-based approach for attack detection in IoT devices: a survey</article-title>. <source>Internet Things</source>. <year>2024</year>;<volume>27</volume>(<issue>4</issue>):<fpage>101306</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2024.101306</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Altulaihan</surname> <given-names>E</given-names></string-name>, <string-name><surname>Almaiah</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Aljughaiman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Anomaly detection IDS for detecting DoS attacks in IoT networks based on machine learning algorithms</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>2</issue>):<fpage>713</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24020713</pub-id>; <pub-id pub-id-type="pmid">38276404</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Deep Q-learning intrusion detection system (DQ-IDS): a novel reinforcement learning approach for adaptive and self-learning cybersecurity</article-title>. <source>ICT Express</source>. <year>2025</year>;<volume>11</volume>(<issue>5</issue>):<fpage>875</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.icte.2025.05.007</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>A novel feature selection-driven ensemble learning approach for accurate botnet attack detection</article-title>. <source>Alex Eng J</source>. <year>2025</year>;<volume>118</volume>(<issue>13</issue>):<fpage>261</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aej.2025.01.042</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bakhsh</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Alshehri</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Enhancing IoT network security through deep learning-powered intrusion detection system</article-title>. <source>Internet Things</source>. <year>2023</year>;<volume>24</volume>(<issue>10</issue>):<fpage>100936</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2023.100936</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Orman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Cyberattack detection systems in industrial Internet of Things (IIoT) networks in big data environments</article-title>. <source>Appl Sci</source>. <year>2025</year>;<volume>15</volume>(<issue>6</issue>):<fpage>3121</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app15063121</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Salayma</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Risk and threat mitigation techniques in Internet of Things (IoT) environments: a survey</article-title>. <source>Front Internet Things</source>. <year>2024</year>;<volume>2</volume>:<fpage>1306018</fpage>. doi:<pub-id pub-id-type="doi">10.3389/friot.2023.1306018</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhukabayeva</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zholshiyeva</surname> <given-names>L</given-names></string-name>, <string-name><surname>Karabayev</surname> <given-names>N</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alnazzawi</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Cybersecurity solutions for industrial Internet of Things&#x2014;edge computing integration: challenges, threats, and future directions</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>1</issue>):<fpage>213</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25010213</pub-id>; <pub-id pub-id-type="pmid">39797003</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wakili</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bakkali</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Privacy-preserving security of IoT networks: a comparative analysis of methods and applications</article-title>. <source>Cyber Secur Appl</source>. <year>2025</year>;<volume>3</volume>(<issue>4</issue>):<fpage>100084</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.csa.2025.100084</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khayat</surname> <given-names>M</given-names></string-name>, <string-name><surname>Barka</surname> <given-names>E</given-names></string-name>, <string-name><surname>Serhani</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Sallabi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shuaib</surname> <given-names>K</given-names></string-name>, <string-name><surname>Khater</surname> <given-names>HM</given-names></string-name></person-group>. <article-title>Reinforcement learning with deep features: a dynamic approach for intrusion detection in IoT networks</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>92319</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3569312</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chiba</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Abghour</surname> <given-names>N</given-names></string-name>, <string-name><surname>Moussaid</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lifandali</surname> <given-names>O</given-names></string-name>, <string-name><surname>Kinta</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A deep study of novel intrusion detection systems and intrusion prevention systems for Internet of Things networks</article-title>. <source>Procedia Comput Sci</source>. <year>2022</year>;<volume>210</volume>(<issue>15</issue>):<fpage>94</fpage>&#x2013;<lpage>103</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2022.10.124</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mallidi</surname> <given-names>SKR</given-names></string-name>, <string-name><surname>Ramisetty</surname> <given-names>RR</given-names></string-name></person-group>. <article-title>Advancements in training and deployment strategies for AI-based intrusion detection systems in IoT: a systematic literature review</article-title>. <source>Discov Internet Things</source>. <year>2025</year>;<volume>5</volume>(<issue>1</issue>):<fpage>8</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s43926-025-00099-4</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Usama Tanveer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Munir</surname> <given-names>K</given-names></string-name>, <string-name><surname>Amjad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ali Jafar Zaidi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bermak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rehman</surname> <given-names>AU</given-names></string-name></person-group>. <article-title>Ensemble-guard IoT: a lightweight ensemble model for real-time attack detection on imbalanced dataset</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>168938</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3495708</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Musthafa</surname> <given-names>MB</given-names></string-name>, <string-name><surname>Huda</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kodera</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Araki</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mwaura</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Optimizing IoT intrusion detection using balanced class distribution, feature selection, and ensemble machine learning techniques</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>13</issue>):<fpage>4293</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24134293</pub-id>; <pub-id pub-id-type="pmid">39001072</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shafin</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Karmakar</surname> <given-names>G</given-names></string-name>, <string-name><surname>Mareels</surname> <given-names>I</given-names></string-name>, <string-name><surname>Balasubramanian</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kolluri</surname> <given-names>RR</given-names></string-name></person-group>. <article-title>Sensor self-declaration of numeric data reliability in Internet of Things</article-title>. <source>IEEE Trans Reliab</source>. <year>2025</year>;<volume>74</volume>(<issue>2</issue>):<fpage>2751</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2024.3416967</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Resource-constraint deep forest-based intrusion detection method in Internet of Things for consumer electronic</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2024</year>;<volume>70</volume>(<issue>2</issue>):<fpage>4976</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tce.2024.3373126</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Simioni</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Viegas</surname> <given-names>EK</given-names></string-name>, <string-name><surname>Santin</surname> <given-names>AO</given-names></string-name>, <string-name><surname>de Matos</surname> <given-names>E</given-names></string-name></person-group>. <article-title>An energy-efficient intrusion detection offloading based on DNN for edge computing</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>12</issue>):<fpage>20326</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2025.3544060</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Almohaimeed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Albalwy</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Enhancing IoT network security using feature selection for intrusion detection systems</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>24</issue>):<fpage>11966</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app142411966</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alturki</surname> <given-names>B</given-names></string-name>, <string-name><surname>Alsulami</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>Semi-supervised learning with entropy filtering for intrusion detection in asymmetrical IoT systems</article-title>. <source>Symmetry</source>. <year>2025</year>;<volume>17</volume>(<issue>6</issue>):<fpage>973</fpage>. doi:<pub-id pub-id-type="doi">10.3390/sym17060973</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Albalwy</surname> <given-names>F</given-names></string-name>, <string-name><surname>Almohaimeed</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Advancing artificial intelligence of things security: integrating feature selection and deep learning for real-time intrusion detection</article-title>. <source>Systems</source>. <year>2025</year>;<volume>13</volume>(<issue>4</issue>):<fpage>231</fpage>. doi:<pub-id pub-id-type="doi">10.3390/systems13040231</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prasad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mohammad Alenazy</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>G</given-names></string-name>, <string-name><surname>Abdallah</surname> <given-names>HA</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Optimizing IoT intrusion detection with cosine similarity based dataset balancing and hybrid deep learning</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>30939</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-15631-3</pub-id>; <pub-id pub-id-type="pmid">40847039</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Katsura</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Endo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Arai</surname> <given-names>I</given-names></string-name>, <string-name><surname>Fujikawa</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Efficient IDS for IoT networks using host-based data aggregation and multi-entropy analysis</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>21</issue>):<fpage>125406</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3589057</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vishwakarma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kesswani</surname> <given-names>N</given-names></string-name></person-group>. <article-title>StaEn-IDS: an explainable stacking ensemble deep neural network-based intrusion detection system for IoT</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>109713</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3582391</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ullah</surname> <given-names>F</given-names></string-name>, <string-name><surname>Turab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ullah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cacciagrano</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Enhanced network intrusion detection system for Internet of Things security using multimodal big data representation with transfer learning and game theory</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>13</issue>):<fpage>4152</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24134152</pub-id>; <pub-id pub-id-type="pmid">39000931</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mahdi</surname> <given-names>ZS</given-names></string-name>, <string-name><surname>Zaki</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Alzubaidi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A secure and adaptive framework for enhancing intrusion detection in IoT networks using incremental learning and blockchain</article-title>. <source>Secur Priv</source>. <year>2025</year>;<volume>8</volume>(<issue>4</issue>):<fpage>e70071</fpage>. doi:<pub-id pub-id-type="doi">10.1002/spy2.70071</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Houichi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jaidi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Bouhoula</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Enhancing smart city security: an intrusion detection system using machine learning methods with the UNB CIC IoT 2023 dataset</article-title>. <source>IET Smart Cities</source>. <year>2025</year>;<volume>7</volume>(<issue>1</issue>):<fpage>e70014</fpage>. doi:<pub-id pub-id-type="doi">10.1049/smc2.70014</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mahdi</surname> <given-names>ZS</given-names></string-name>, <string-name><surname>Zaki</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Alzubaidi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Advanced hybrid techniques for cyberattack detection and defense in IoT networks</article-title>. <source>Secur Priv</source>. <year>2025</year>;<volume>8</volume>(<issue>2</issue>):<fpage>e471</fpage>. doi:<pub-id pub-id-type="doi">10.1002/spy2.471</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dom&#x00E9;nech</surname> <given-names>J</given-names></string-name>, <string-name><surname>Le&#x00F3;n</surname> <given-names>O</given-names></string-name>, <string-name><surname>Siddiqui</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Pegueroles</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Evaluating and enhancing intrusion detection systems in IoMT: the importance of domain-specific datasets</article-title>. <source>Internet Things</source>. <year>2025</year>;<volume>32</volume>(<issue>1</issue>):<fpage>101631</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2025.101631</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hafid</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rahouti</surname> <given-names>M</given-names></string-name>, <string-name><surname>Aledhari</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Optimizing intrusion detection in IoMT networks through interpretable and cost-aware machine learning</article-title>. <source>Mathematics</source>. <year>2025</year>;<volume>13</volume>(<issue>10</issue>):<fpage>1574</fpage>. doi:<pub-id pub-id-type="doi">10.3390/math13101574</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mezina</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nurmi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ometov</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Novel hybrid UNet&#x002B;&#x002B; and LSTM model for enhanced attack detection and classification in IoMT traffic</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>1</issue>):<fpage>57589</fpage>&#x2013;<lpage>603</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3553966</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Haq</surname> <given-names>MIU</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Das</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Shetty</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Efficiently learning an encoder that classifies token replacements and masked permuted network-based BIGRU attention classifier for enhancing sentiment classification of scientific text</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>3</issue>):<fpage>190240</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3516946</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Regler</surname> <given-names>B</given-names></string-name>, <string-name><surname>Scheffler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ghiringhelli</surname> <given-names>LM</given-names></string-name></person-group>. <article-title>TCMI: a non-parametric mutual-dependence estimator for multivariate continuous distributions</article-title>. <source>Data Min Knowl Discov</source>. <year>2022</year>;<volume>36</volume>(<issue>5</issue>):<fpage>1815</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10618-022-00847-y</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>LinFuzz: program-sensitive seed scheduling greybox fuzzing based on LinUCB algorithm</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>10</issue>):<fpage>74843</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3404918</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sharmila</surname> <given-names>B</given-names></string-name>, <string-name><surname>Nagapadma</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Quantized autoencoder (QAE) intrusion detection system for anomaly detection in resource-constrained IoT devices using RT-IoT2022 dataset</article-title>. <source>Cybersecurity</source>. <year>2023</year>;<volume>6</volume>(<issue>1</issue>):<fpage>41</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s42400-023-00178-5</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Neto</surname> <given-names>ECP</given-names></string-name>, <string-name><surname>Dadkhah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zohourian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>CICIoT2023: a real-time dataset and benchmark for large-scale attacks in IoT environment</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>13</issue>):<fpage>5941</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23135941</pub-id>; <pub-id pub-id-type="pmid">37447792</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dadkhah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Neto</surname> <given-names>ECP</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>R</given-names></string-name>, <string-name><surname>Molokwu</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Sadeghi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>CICIoMT2024: a benchmark dataset for multi-protocol security assessment in IoMT</article-title>. <source>Internet Things</source>. <year>2024</year>;<volume>28</volume>:<fpage>101351</fpage>.</mixed-citation></ref>
</ref-list>
</back></article>