<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">74349</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2025.074349</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Hybrid Split-Attention and Transformer Architecture for High-Performance Network Intrusion Detection</article-title>
<alt-title alt-title-type="left-running-head">A Hybrid Split-Attention and Transformer Architecture for High-Performance Network Intrusion Detection</alt-title>
<alt-title alt-title-type="right-running-head">A Hybrid Split-Attention and Transformer Architecture for High-Performance Network Intrusion Detection</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Zhu</surname><given-names>Gan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yu</surname><given-names>Yongtao</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>csl@yxnu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Deng</surname><given-names>Xiaofan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Dai</surname><given-names>Yuanchen</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Zhenyuan</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Software, Yunnan University</institution>, <addr-line>Kunming, 650504</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Yunnan Key Laboratory of Smart City in Cyberspace Security, Yuxi Normal University</institution>, <addr-line>Yuxi, 653100</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>School of Information Science and Technology, Yunnan Normal University</institution>, <addr-line>Kunming, 650500</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yongtao Yu. Email: <email>csl@yxnu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>12</month><year>2025</year>
</pub-date>
<volume>145</volume>
<issue>3</issue>
<fpage>4317</fpage>
<lpage>4348</lpage>
<history>
<date date-type="received">
<day>09</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_74349.pdf"></self-uri>
<abstract>
<p>Existing deep learning Network Intrusion Detection Systems (NIDS) struggle to simultaneously capture fine-grained, multi-scale features and long-range temporal dependencies. To address this gap, this paper introduces TransNeSt, a hybrid architecture integrating a ResNeSt block (using split-attention for multi-scale feature representation) with a Transformer encoder (using self-attention for global temporal modeling). This integration of multi-scale and temporal attention was validated on four benchmarks: NSL-KDD, UNSW-NB15, CIC-IDS2017, and CICIOT2023. TransNeSt consistently outperformed its individual components and several state-of-the-art models, demonstrating significant quantitative gains. The model achieved high efficacy across all datasets, with F1-Scores of 99.04% (NSL-KDD), 91.92% (UNSW-NB15), 99.18% (CIC-IDS2017), and 97.85% (CICIOT2023), confirming its robustness.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Intrusion detection</kwd>
<kwd>transformer</kwd>
<kwd>resnest</kwd>
<kwd>split attention</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Opening Foundation of Yunnan Key Laboratory of Smart City in Cyberspace Security</funding-source>
<award-id>202105AG070010</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The expansion of the digital attack surface, driven by complex and distributed network architectures, has led to more sophisticated and rapid cyber intrusions. This evolving threat landscape erodes the efficacy of conventional defenses, such as signature-based intrusion detection systems (IDS). This reality demands a new generation of Network Intrusion Detection Systems (NIDS) with adaptive and predictive capabilities to discern both known and emergent (zero-day) threats in high-velocity data streams. Consequently, research has pivoted towards deep learning (DL), valued for its ability to autonomously extract salient feature hierarchies from high-dimensional network traffic [<xref ref-type="bibr" rid="ref-1">1</xref>].</p>
<p>Early DL-based NIDS highlighted a key trade-off. Convolutional Neural Networks (CNNs) identified localized patterns, but their limited receptive field struggled with long-range temporal dependencies [<xref ref-type="bibr" rid="ref-2">2</xref>]. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) models were designed for sequential analysis, but faced high computational overhead and gradient issues [<xref ref-type="bibr" rid="ref-3">3</xref>]. To address these limitations, the field has advanced significantly. One line of research involves sophisticated hybrid recurrent models, such as BiGRU (Bidirectional Gated Recurrent Unit)-LSTM with attention, which have proven robust for capturing complex temporal patterns [<xref ref-type="bibr" rid="ref-4">4</xref>]. A second, highly influential line of research adopted the Transformer architecture [<xref ref-type="bibr" rid="ref-5">5</xref>], leveraging its powerful self-attention mechanism to model global dependencies. As recent surveys confirm, Transformers and even Large Language Models (LLMs) are now a central focus for efficient intrusion detection [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>However, while these approaches excel at modeling temporal and global relationships, the direct application of standard Transformers, for example, is not optimized for extracting the granular, multi-scale feature hierarchies embedded within individual network flows. This creates an opportunity for an architecture that synergistically unifies the global contextual modeling of Transformers with the multi-scale feature extraction capabilities of advanced convolutional networks [<xref ref-type="bibr" rid="ref-7">7</xref>]. To resolve this specific trade-off, we designed TransNeSt, a hybrid architecture that integrates a ResNeSt (ResNet with Split-Attention) backbone [<xref ref-type="bibr" rid="ref-8">8</xref>] with a Transformer encoder. The core of our contribution is the novel adaptation of the Split-Attention mechanism, originally conceived for computer vision, to the one-dimensional structure of network traffic data. The ResNeSt backbone functions as a potent feature generator, employing its multi-path structure and cross-channel attention to produce a sequence of rich, cardinal-group feature representations from the raw traffic. This output sequence then serves as the input for the Transformer encoder, which leverages self-attention to explicitly model the global contextual relationships between these high-level feature sets.</p>
<p>We validated TransNeSt&#x2019;s performance and generalization capabilities through rigorous experimentation on four standard benchmarks: NSL-KDD, UNSW-NB15, CICIDS-2017 and CICIOT2023. These datasets were selected to ensure a comprehensive evaluation against a spectrum of threats, from legacy attack vectors to complex, modern intrusions. Across all four benchmarks, TransNeSt consistently achieved higher performance metrics than the other models tested, indicating its effectiveness for deep learning-based intrusion detection. A comparison of our proposed pipeline against traditional approaches is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Comparison of the Traditional IDS pipeline (left) and the proposed TransNeSt architecture (right)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-1.tif"/>
</fig>
<p>The research presented herein offers several significant contributions:
<list list-type="order">
<list-item>
<p>This paper proposes TransNeSt, a novel hybrid deep learning architecture designed to leverage ResNeSt for its multi-scale feature extraction prowess while employing the Transformer to capture global contextual relationships in network intrusion detection.</p></list-item>
<list-item>
<p>We have successfully adapted the Split-Attention mechanism, initially conceived for two-dimensional image analysis, for application to one-dimensional network traffic data, thereby demonstrating its considerable versatility and significant potential beyond the domain of computer vision.</p></list-item>
<list-item>
<p>We performed a comprehensive and rigorous evaluation using a diverse set of datasets. Our results demonstrate the detection efficacy and generalizability of our approach and provide strong performance results on these benchmarks.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Machine Learning and Optimization Approaches</title>
<p>The landscape of network intrusion detection continues to be substantively shaped by advancements in both classical and contemporary computational intelligence paradigms. Research has focused on optimizing classical models and addressing specific data challenges. For parameter refinement, Kolukisa et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] synergized logistic regression with a parallelized artificial bee colony (ABC) algorithm, achieving commendable detection performance. To confront high-dimensional feature spaces, Turukmane et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] integrated multi-layer Support Vector Machines (SVMs) with singular value decomposition. For the complexities of mixed-traffic environments, Rustam et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] introduced the Fully Automated Malicious Traffic Detection System (FAMTDS) framework, using a Moth Flame Optimizer (MFO) for automation. Addressing the critical problem of data imbalance, Hooshmand et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] developed a multi-stage model (K-means (SKM) hybrid, Denoising Autoencoder (DAE), and XGBoost classifier), achieving binary/multiclass accuracies up to 99.57% on NSL-KDD. Finally, for scalable stream processing, He et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] proposed a Bayesian gamma mixture model (GaMM) employing extended stochastic variational inference (ESVI).</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deep Learning Architectures for NIDS</title>
<p>The utility of deep learning in modern intrusion detection stems from its capacity for automated hierarchical feature extraction. Research has explored several integration strategies. For instance, novel architectures are being developed to work with limited labeled data; Xu et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed the Time-Space Separable Attention Network (TSSAN), which utilizes depth-wise separable convolution and a time-space self-attention mechanism to extract features in an unsupervised or semi-supervised context. Architectural synergy has also been exploited to model both spatial and temporal dependencies; Liu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed the Residual Memory Convolutional Neural Network (RMCNN), a hybrid model that first extracts spatial features with CNN layers, then captures temporal dynamics using a Gated Recurrent Unit (GRU), and reinforces critical features with a multi-head attention mechanism. To address the challenge of novel, unseen attacks in few-shot scenarios, meta-learning frameworks have been introduced. Xu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed Marrying Attention and Convolution-based Meta-learning (MACML), an optimization-based meta-learning method that &#x201C;marries&#x201D; a self-attention mechanism (for global dependencies) with a CNN (for local features), enabling rapid adaptation to new attack types with minimal samples. Hybrid dimensionality reduction has also been proposed for LSTM-based models, as seen in the two-stage fusion strategy by Thakkar et al. [<xref ref-type="bibr" rid="ref-17">17</xref>]. Their method, fusing a non-linear Autoencoder (AE) with linear Principal Component Analysis (PCA), yielded a performance gain of up to 3% over a standard AE&#x002B;LSTM baseline. Similarly, Alsoufi et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] addressed high-dimensional IoT data by employing a Sparse Autoencoder (SAE) for feature dimensionality reduction coupled with a CNN for classification, demonstrating significant accuracy improvements in IoT environments</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Hybrid and Ensemble Frameworks</title>
<p>A prominent research strategy is to couple DL feature engineering with classical classifiers or other learning paradigms. Chen et al. [<xref ref-type="bibr" rid="ref-19">19</xref>], for example, introduced Metric Learning Framework with Dual One-Class Units (MLF-DOU), which uses one-class classifiers and a triplet network-based metric learning stage to maximize inter-class separability, demonstrating superior accuracy and F1-score on UNSW-NB15. Hierarchical designs represent another approach to tackle specific challenges, such as identifying rare attack classes (a form of data imbalance). Wang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] devised a two-layer model where a CNN-BiLSTM layer (addressing temporal dependencies) performs coarse-grained classification, followed by a Stacking ensemble learner for fine-grained analysis of minority classes, which demonstrated superior recall. Reinforcement learning offers a third path; Hossain [<xref ref-type="bibr" rid="ref-21">21</xref>] ventured into this with their Deep Q-Learning Intrusion Detection System (DQ-IDS). This Deep Q-Network (DQN) autonomously learns classification policies using experience replay and adaptive exploration, achieving a notable 97.18% accuracy and 98.52% F1-score on the real-world CICIoT2023 dataset. Finally, to handle the dynamic nature of threats and avoid &#x201C;catastrophic forgetting,&#x201D; Wang et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] introduced CTWA, an incremental learning method that uses a Convolutional Autoencoder (CAE) and Temporal Convolutional Network (TCN) to extract features. This framework applies Weight Alignment (WA) techniques, enabling the model to learn new attack classes without losing knowledge of old ones.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Strategies for Data Imbalance</title>
<p>Persistent data imbalance remains a critical challenge. One solution is generative data enrichment; Arafah et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] confronted this with the Autoencoder-Wasserstein Generative Adversarial Network (AE-WGAN), where a Wasserstein Generative Adversarial Network synthesizes high-fidelity minority attack samples, yielding substantial performance gains. An alternative strategy integrates imbalance correction directly into the classification pipeline. The Network Intrusion Detection System-Convolutional Neural Network Random Forest (NIDS-CNNRF) model, proposed by Yang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>], exemplifies this by coupling a standard CNN extractor with a Random Forest (RF) classifier, but first applying PCA and the Adaptive Synthetic Sampling (ADASYN) algorithm to resample the data, which was validated across several benchmarks.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Key Challenges and Research Gaps</title>
<p>The literature reveals a clear trajectory towards hybrid models, but a persistent limitation is this reliance on standard CNNs for feature extraction. This highlights a key research gap: the need for models that integrate feature extractors capable of capturing more complex, multi-scale hierarchies. The proposed TransNeSt architecture is designed to fill this gap. To our knowledge, no existing NIDS research has adapted split-attention mechanisms (ResNeSt), originally designed for 2D computer vision, for 1D network traffic analysis. TransNeSt directly addresses the multi-scale feature challenge by integrating a state-of-the-art ResNeSt block and couples it with a Transformer encoder to effectively model long-term temporal dependencies, creating a more profound synergy between feature representation and temporal analysis.</p>
<p>This review of the literature, for which a detailed summary is presented in <xref ref-type="table" rid="table-1">Table 1</xref>, highlights three persistent challenges that motivate our work. First, data imbalance is a critical and prevalent challenge, as rare but dangerous attack classes like Remote-to-Local (R2L) and User-to-Root (U2R) are vastly underrepresented, causing models to develop a bias toward majority (benign) traffic. Second, effectively modeling complex temporal dependencies remains an open problem; while LSTMs were an improvement over CNNs, they suffer from computational overhead, and attacks are often sophisticated, multi-stage sequences that are difficult to model. Third, discerning subtle attacks requires capturing a multi-scale feature hierarchy, but the standard CNNs used in many hybrid models fail to capture the granular, hierarchical feature relationships embedded within network traffic.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of network intrusion detection methodologies</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>References</th>
<th>Datasets</th>
<th>Methodology</th>
<th>Limitations</th>
<th>Results</th>
</tr>
</thead>
<tbody>
<tr>
<td>Kolukisa et al. [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>UNSW-NB15, NSL-KDD</td>
<td>LR-ABC model (Logistic Regression &#x002B; parallel ABC)</td>
<td>Long training time.</td>
<td>Accuracy and F1 88.25%, 88.26% (UNSW-NB15), 90.11%, 90.15% (NSL-KDD)</td>
</tr>
<tr>
<td>Turukmane et al. [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>CSE-CIC-IDS 2018, UNSW-NB15</td>
<td>M-MultiSVM hybrid model; uses ASmoT, M-SvD and Mud Ring Optimization</td>
<td>Slow attack detection; unsuitable for large-scale AMI networks.</td>
<td>Acc: 99.89% (CSE-CIC-IDS 2018), 97.535% (UNSW-NB15)</td>
</tr>
<tr>
<td>Rustam et al. [<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>UNSW-NB15, IoTID-20</td>
<td>FAMTDS (Fully Automated system optimized by MFO)</td>
<td>Limited to simple network architectures; Small dataset size.</td>
<td>Accuracy: 0.85 (Multi-Environment) and F1: 0.84 (Multi-Environment)</td>
</tr>
<tr>
<td>Hooshmand et al. [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>NSL-KDD, UNSW-NB15</td>
<td>SMOTE with SKM and XGBoost (SKM-XGB)</td>
<td>Slightly longer training time than baselines.</td>
<td>Accuracy and F1: 99.01%, 99.02% (UNSW-NB15 binary), 99.37%, 99.37% (NSL-KDD binary)</td>
</tr>
<tr>
<td>He et al. [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>CICIDS2018, OPCUA, CIC-MalMem-2022</td>
<td>Bayesian Gamma Mixture Model (GaMM) with Extended Stochastic Variational Inference (ESVI)</td>
<td>Robustness against adversarial attacks was not validated.</td>
<td>F1: 0.995620 (CICMalmem2022), 0.999721 (OPCUA), 0.999940 (CICIDS2018)</td>
</tr>
<tr>
<td>Xu et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>UNSW-NB15, CICIDS-2017</td>
<td>TSSAN: Time-Space Separable Attention Network</td>
<td>High computational cost for pre-training.</td>
<td>F1: 0.86 (UNSW-NB15), 0.92 (CICIDS-2017) (Unsupervised)</td>
</tr>
<tr>
<td>Liu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>CICIOT2023</td>
<td>RMCNN: CNN and GRU, Multi-Head Attention</td>
<td>Limitations not explicitly stated.</td>
<td>Accuracy: 97.28%, F1-Score: 97.29%</td>
</tr>
<tr>
<td>Xu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>CICIoT2023</td>
<td>MACML: Few-shot meta-learning framework combining Self-Attention (global) and CNN (local)</td>
<td>Multiclass classification challenging; primarily evaluated on binary tasks.</td>
<td>Acc (10-shot): 91.05%, F1: 90.98%</td>
</tr>
<tr>
<td>Thakkar et al. [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>NSL-KDD, UNSW-NB15, CIC-IDS-2017</td>
<td>AE&#x002B;PCA&#x002B;LSTM: Fused AE (non-linear) and PCA (linear) reduction (to 30 dims), followed by LSTM classifier.</td>
<td>Reduction techniques (AE/PCA) not universally applicable; features may be uninterpretable.</td>
<td>Accuracy and F1: 82.22%, 82.82% (NSL-KDD), 76.28%, 81.41% (UNSW-NB15)</td>
</tr>
<tr>
<td>Alsoufi et al. [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Bot-IoT</td>
<td>SAE-CNN: Hybrid Sparse Autoencoder and CNN architecture</td>
<td>Evaluation limited to a single dataset.</td>
<td>Accuracy: 99.9%, F1-Score: 99.9%</td>
</tr>
<tr>
<td>Chen et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>NSL-KDD, KDD-99, UNSW-NB15</td>
<td>Metric Learning Framework with Dual One-Class Units (MLF-DOU) using a triplet network</td>
<td>High computational complexity and focused on binary classification.</td>
<td>Accuracy and F1: 93.14%, 94.16% (NSL-KDD), 90.68%, 91.47% (UNSW-NB15)</td>
</tr>
<tr>
<td>Wang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>CIC-IDS2017, NSL-KDD</td>
<td>Two-layer NIDS: L1 (CNN&#x002B;BiLSTM&#x002B;Attention), L2 Stacking ensemble learning for minority classes</td>
<td>Difficulty classifying extremely rare attacks (e.g., U2R).</td>
<td>Accuracy and F1: 0.99, 0.99 (NSL-KDD excl. U2R) 0.99, 0.99 (CIC-IDS2017)</td>
</tr>
<tr>
<td>Hossain [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>CICIoT2023</td>
<td>Deep Q-Learning Intrusion Detection System (DQ-IDS) using Deep Q-Networks (DQN)</td>
<td>High false positives; slow adaptation to new threats; high computational overhead; risk of overfitting.</td>
<td>Accuracy: 97.18% F1: 98.52%</td>
</tr>
<tr>
<td>Wang et al. [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>CICIoT2023, BoT-NetIoT</td>
<td>CTWA: Incremental Learning (CIL) framework using CAE &#x002B; TCN with Weight Alignment (WA)</td>
<td>High runtime; complex model; difficulty with small-sample unknown attacks.</td>
<td>Accuracy: 96.43%, F1-Score: 96.45%</td>
</tr>
<tr>
<td>Arafah et al. [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>NSL-KDD, CICIDS-2017</td>
<td>Denoising Autoencoder (AE) and Wasserstein GAN (WGAN) to generate synthetic attacks</td>
<td>Needs extensive real-time deployment for validation and generalization testing.</td>
<td>Accuracy and F1: 93.00%, 94.00% (NSL-KDD), 98.00%, 98.00% (CICIDS-2017)</td>
</tr>
<tr>
<td>Yang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>KDD CUP99, NSL-KDD, CIC-IDS2017</td>
<td>NIDS-CNNRF: CNN (feature extraction) and Random Forest (classification), with ADASYN and PCA</td>
<td>ADASYN may introduce noise; PCA may lose information; low accuracy for some attack classes.</td>
<td>Accuracy and F1: 99.75%, 99.75% (NSL-KDD) 98.07%, 98.02% (CIC-IDS2017)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Model</title>
<p>To address the challenge of capturing both fine-grained spatial patterns and complex temporal dependencies in network traffic data, this paper introduces an innovative hybrid deep learning architecture for Network Intrusion Detection, a framework we have termed TransNeSt. The architecture of this model, illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, is expressly designed to create a powerful synergy between high-resolution convolutional feature extraction and global, attention-based temporal analysis. The model&#x2019;s workflow is structured into two sequential stages. Initially, a ResNeSt Block functions as an advanced feature extractor, meticulously analyzing input network flows to capture fine-grained, multi-scale spatial patterns. Subsequently, the sequence of feature representations generated by this block is then passed to a Transformer Block, which captures the long-range temporal dependencies across the entire sequence. This integrated architecture facilitates a comprehensive understanding of network traffic, which is paramount for the robust detection of both simple and sophisticated, multi-stage cyber-attacks.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>TransNeSt architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-2.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Overall Architecture</title>
<p>The overall architecture of TransNeSt follows a two-stage sequential workflow. First, a ResNeSt Block functions as an advanced feature extractor, processing the input sequence to capture fine-grained, feature-based spatial patterns. Subsequently, the sequence of enriched feature representations generated by this block is passed to a Transformer Block, which models the long-range temporal dependencies across the entire sequence.</p>
<p>Formally, the input to our model is a sequence of pre-processed network traffic data, denoted as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>x</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where each <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula> is a <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>d</mml:mi></mml:math></inline-formula>-dimensional vector representing a single network flow or packet. This sequence is first processed by the ResNeSt Block. The principal role of this block is to transform the initial input into a sequence of enriched feature vectors, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>Z</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>z</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where each vector <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> contains a high-level representation of salient local characteristics. This sequence <italic>Z</italic> then serves as the input to the Transformer Block. The Transformer employs its self-attention mechanism to conduct a global analysis, modeling the temporal context and intricate inter-flow relationships across the entire sequence. The final output from the Transformer is then processed by a conventional classification head, which consists of a Linear layer and a Softmax activation function, in order to generate the final probability distribution across the designated classes of network activity (e.g., Normal, DoS, Probe).</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>ResNeSt Block</title>
<p>The ResNeSt block, originally designed for 2D image analysis, is a core component of our model. A key innovation of our work is the adaptation of its architecture to effectively process 1D sequential network traffic data by replacing 2D convolutions with their 1D counterparts (e.g., <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>). Our selection of the ResNeSt block is predicated on its empirically demonstrated advantages over incumbent architectures like ResNet and Squeeze-and-Excite Network (SE-Net). The seminal ResNeSt paper [<xref ref-type="bibr" rid="ref-8">8</xref>] provides extensive empirical validation, showing that ResNeSt models consistently outperform their ResNet and SE-Net counterparts across multiple benchmarks.</p>
<p>The process within a ResNeSt block begins with an input feature map of dimensions <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which is first divided into <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>k</mml:mi></mml:math></inline-formula> cardinal groups. Each of these groups is then further subdivided into <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>r</mml:mi></mml:math></inline-formula> radix splits, yielding a total of <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi></mml:math></inline-formula> parallel processing paths. These hyperparameters <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>k</mml:mi></mml:math></inline-formula> (cardinality) and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>r</mml:mi></mml:math></inline-formula> (radix) directly control the model&#x2019;s complexity and capacity; increasing them allows for more complex representations at the cost of higher computational load. Let the feature map for a single split be <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. This split-level feature map is then passed through a sequence of 1D convolutions, comprising a <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution that precedes a <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution, to produce a transformed feature map <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>U</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mtext>Conv</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The core innovation lies in the Split Attention Mechanism, which intelligently fuses the information from these transformed splits <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>.</p>
<p>The Split Attention mechanism, explicitly diagrammed in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, directs the flow of information and adaptively recalibrates feature responses within each cardinal group. The mechanism unfolds as follows:</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>A detailed view of the Split Attention mechanism</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-3.tif"/>
</fig>
<p>First, the transformed feature maps from the <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>r</mml:mi></mml:math></inline-formula> radix splits are aggregated via element-wise summation, producing a unified representation for the cardinal group:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mover><mml:mi>U</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>U</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>To generate a channel-wise summary of contextual information, global average pooling (<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) is applied to <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:mover><mml:mi>U</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, squeezing its spatial dimension to produce a channel descriptor vector <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>s</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>j</mml:mi></mml:math></inline-formula>-th component of this vector is given by:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>s</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>U</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>l</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mover><mml:mi>U</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>This vector <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>s</mml:mi></mml:math></inline-formula> is then passed through a small sub-network to generate split-specific attention weights.It is first passed through a fully-connected layer with weights <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and biases <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, then Batch Normalization (<inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mi>&#x0212C;</mml:mi></mml:mrow></mml:math></inline-formula>) and a Rectified Linear Unit (ReLU) activation (<inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>), producing the condensed feature vector <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>z</mml:mi></mml:math></inline-formula>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x0212C;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a bottleneck dimension, typically smaller than <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:math></inline-formula>. Subsequently, a second fully-connected layer with weights <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup></mml:math></inline-formula> and biases <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> takes <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>z</mml:mi></mml:math></inline-formula> as input and generates attention scores for all splits. This output is reshaped to produce <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>r</mml:mi></mml:math></inline-formula> distinct attention score vectors, <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:msub><mml:mi>s</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where each <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>These scores are normalized across the radix dimension for each channel using an r-Softmax function. The attention weight <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> for the <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>j</mml:mi></mml:math></inline-formula>-th channel of the <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>i</mml:mi></mml:math></inline-formula>-th split is given by:</p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:munderover><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>This ensures that for any channel <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>j</mml:mi></mml:math></inline-formula>, the attention weights across all <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>r</mml:mi></mml:math></inline-formula> splits sum to one, i.e., <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. The Split Attention block&#x2019;s final output, <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>V</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula>, for the cardinal group is a weighted fusion of the split feature maps, where each map <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msubsup><mml:mi>U</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is recalibrated by its corresponding attention vector <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2299;</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> denotes channel-wise multiplication, with the vector <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> being broadcast along the spatial dimension. The concatenated outputs from all <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>k</mml:mi></mml:math></inline-formula> cardinal groups are then passed through a final <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution, and a residual connection is added to the original block input to ensure stable and efficient training [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Transformer Block</title>
<p>Subsequent to the extraction of localized feature hierarchies by the ResNeSt block, the modeling of temporal dynamics across the sequence of network events is entrusted to a bespoke Transformer Block [<xref ref-type="bibr" rid="ref-5">5</xref>]. A pivotal architectural decision within this block is the normalization strategy. In a departure from the canonical Transformer, we substitute the conventional Layer Normalization with Root Mean Square Normalization (RMSNorm) [<xref ref-type="bibr" rid="ref-26">26</xref>]. This modification is deliberately implemented to enhance computational throughput and training efficiency. As demonstrated in [<xref ref-type="bibr" rid="ref-26">26</xref>], RMSNorm achieves comparable performance to LayerNorm but with significantly reduced computational overhead, observing speedups from 7% to 64% across different models. This efficiency gain comes from simplifying the operation: RMSNorm preserves the essential re-scaling invariance of LayerNorm but eschews the mean-centering operation, which contributes less to training stability while incurring computational costs. RMSNorm thus provides a more efficient drop-in replacement. For an input feature vector <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula>, the operation is formally defined as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>RMSNorm</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mi>x</mml:mi><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mi>d</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03F5;</mml:mi></mml:msqrt></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>g</mml:mi></mml:math></inline-formula> denotes a learnable gain parameter and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is a small constant for numerical stability.</p>
<p>Each Transformer block comprises two main sub-layers: a Multi-Head Attention (MHA) mechanism and a position-wise Feed-Forward Network (FFN).</p>
<p>In the MHA sub-layer, the input tensor <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>Z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is linearly mapped to form Query (<italic>Q</italic>), Key (<italic>K</italic>), and Value (<italic>V</italic>) matrices. The attention is then calculated using scaled dot-product attention:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>Attention</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>Q</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>V</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the dimensionality of the key vectors.</p>
<p>The output of the MHA sub-layer is subsequently processed by the FFN, which consists of two linear transformations with a ReLU activation:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>FFN</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></disp-formula></p>
<p>Each of these sub-layers is encapsulated with a residual connection and an RMSNorm layer. To build a deeper representation, this entire four-sub-layer block is stacked for two iterations.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Output and Classification Layer</title>
<p>The architectural design of the TransNeSt model culminates in a dedicated classification head, which performs the final, decisive conversion of the deeply learned feature vectors into a final probability distribution over the class labels. Adhering to the established paradigm for Transformer-based classification, the predictive inference is predicated on the terminal hidden state vector corresponding to a special classification token <monospace>[CLS]</monospace>. This vector, denoted as <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula>, functions as a holistic, aggregated representation of the entire input sequence. It is subsequently projected onto the logit space via a fully-connected linear layer, which is parameterized by a weight matrix <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>W</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and a bias vector <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>b</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msup></mml:math></inline-formula>, yielding a vector of unnormalized scores (logits), <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msup></mml:math></inline-formula>, where <italic>K</italic> signifies the number of target classes.</p>
<p>To transform these raw logits into a valid posterior probability distribution, the Softmax activation function is employed to yield the probability <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> for any given class <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>j</mml:mi></mml:math></inline-formula>, defined as:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The model&#x2019;s final prediction, <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, is obtained by selecting the class index with the highest posterior probability, a process performed using the argmax function:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mi>j</mml:mi></mml:munder><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>The model is trained to perform multi-class classification, as all datasets used in this study involve multiple classes of network activity (e.g., Normal and various attack types). To optimize the model&#x2019;s parameters, we employ the Categorical Cross-Entropy (CCE) loss function. This loss function is standard for multi-class problems and measures the dissimilarity between the model&#x2019;s predicted probability distribution <italic>P</italic> and the true one-hot encoded label vector <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>y</mml:mi></mml:math></inline-formula>. For a single sample, the CCE loss <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:math></inline-formula> is defined as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <italic>K</italic> is the total number of classes, <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the ground truth (1 if the sample belongs to class <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>k</mml:mi></mml:math></inline-formula>, 0 otherwise), and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>P</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the model&#x2019;s predicted probability for class <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>k</mml:mi></mml:math></inline-formula> from the Softmax output (<xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>).</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Setup</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Parameters</title>
<p>The feature extraction backbone, a stack of ResNeSt modules, was configured to create a robust and hierarchical feature representation. We selected four sequential blocks to ensure sufficient depth for learning complex feature representations from the input data. Within each block, we further optimized the split-attention mechanism to enhance representational power. Within each block, the split-attention mechanism was configured with a cardinality (<inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>k</mml:mi></mml:math></inline-formula>) of 2 and a radix (<inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>r</mml:mi></mml:math></inline-formula>) of 2. The feature processing pipeline commences with a 1D convolutional stem that projects the input into a 32-channel feature space. We maintain this channel width across all subsequent ResNeSt blocks to ensure consistent dimensionality for hierarchical feature learning.</p>
<p>In the temporal modeling stage, the Transformer Encoder was designed to effectively capture long-range dependencies. The encoder uses a 4-layer configuration, where each layer contains a Multi-Head Attention mechanism with 4 parallel heads and a Feed-Forward Network with an inner dimension of 256. To enhance regularization, a dropout rate of 0.3 was applied within the Transformer&#x2019;s sub-layers. As previously justified, standard LayerNorm was replaced with the more parameter-efficient RMSNorm, configured with a numerical stability term <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>&#x03F5;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, throughout the encoder.</p>
<p>To ensure consistent experimental outcomes and effectively manage the model&#x2019;s complexity, we actively prevented overfitting through a combination of architectural, data, and training-level regularization. This included applying Dropout, controlling the model&#x2019;s depth, employing the Adaptive Moment Estimation (Adam) optimizer with L2 Regularization (Weight Decay), and data augmentation. Furthermore, an Early Stopping was used to halt training at the point of peak generalization. The choice of key hyperparameters was determined through an empirical tuning process on the UNSW-NB15, with results summarized in <xref ref-type="table" rid="table-2">Table 2</xref>. This configuration was found to provide the best balance between model capacity and generalization, avoiding the underfitting of simpler models (e.g., 2 layers) and the diminishing returns or overfitting of more complex ones (e.g., 0.1 dropout). To ensure consistent experimental outcomes and effectively manage this final model&#x2019;s complexity, we incorporated several key training procedures.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Hyperparameter tuning results on the UNSW-NB15 dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
<th>Training accuracy</th>
<th>Test accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>Transformer layers</italic></td>
<td>2</td>
<td>0.9126</td>
<td>0.9012</td>
</tr>
<tr>
<td>(Heads &#x003D; 4, Dropout &#x003D; 0.3)</td>
<td>4</td>
<td>0.9394</td>
<td>0.9190</td>
</tr>
<tr>
<td></td>
<td>6</td>
<td>0.9432</td>
<td>0.9101</td>
</tr>
<tr>
<td><italic>Attention heads</italic></td>
<td>2</td>
<td>0.9355</td>
<td>0.9133</td>
</tr>
<tr>
<td>(Layers &#x003D; 4, Dropout &#x003D; 0.3)</td>
<td>4</td>
<td>0.9394</td>
<td>0.9190</td>
</tr>
<tr>
<td></td>
<td>8</td>
<td>0.9420</td>
<td>0.9109</td>
</tr>
<tr>
<td><italic>Dropout rate</italic></td>
<td>0.1</td>
<td>0.9468</td>
<td>0.8954</td>
</tr>
<tr>
<td>(Layers &#x003D; 4, Heads &#x003D; 4)</td>
<td>0.2</td>
<td>0.9441</td>
<td>0.9120</td>
</tr>
<tr>
<td></td>
<td>0.3</td>
<td>0.9394</td>
<td>0.9190</td>
</tr>
<tr>
<td></td>
<td>0.4</td>
<td>0.9281</td>
<td>0.9112</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The final architectural parameters derived from this selection process are summarized in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Architectural, training, and regularization parameters</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="4"><italic><bold>ResNeSt</bold></italic></td>
</tr>
<tr>
<td>&#x2003;Initial Input Channels</td>
<td>32</td>
<td>Number of ResNeSt Blocks</td>
<td>4</td>
</tr>
<tr>
<td>&#x2003;Cardinality (<inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>k</mml:mi></mml:math></inline-formula>)</td>
<td>2</td>
<td>Radix (<inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>r</mml:mi></mml:math></inline-formula>)</td>
<td>2</td>
</tr>
<tr>
<td>&#x2003;Output Channels per Block</td>
<td>32</td>
<td></td>
<td></td>
</tr>
<tr>
<td colspan="4"><italic><bold>Transformer</bold></italic></td>
</tr>
<tr>
<td>&#x2003;Number of Transformer Layers</td>
<td>4</td>
<td>Attention heads</td>
<td>4</td>
</tr>
<tr>
<td>&#x2003;FFN Inner Dimension</td>
<td>256</td>
<td>Dropout rate</td>
<td>0.3</td>
</tr>
<tr>
<td>&#x2003;RMSNorm Epsilon (<inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td></td>
<td></td>
</tr>
<tr>
<td colspan="4"><italic><bold>Training &#x0026; Regularization</bold></italic></td>
</tr>
<tr>
<td>&#x2003;Optimizer</td>
<td>Adam</td>
<td>Learning rate</td>
<td><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>&#x2003;Batch size</td>
<td>128</td>
<td>Max Epochs</td>
<td>100</td>
</tr>
<tr>
<td>&#x2003;Random seed</td>
<td>42</td>
<td>L2 Weight decay</td>
<td><inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>&#x2003;Early stopping patience</td>
<td>10 epochs</td>
<td></td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Datasets</title>
<p>To establish the empirical validity and generalization capabilities of the proposed model, our investigation was performed on a curated selection of four canonical benchmark datasets: NSL-KDD, UNSW-NB15, CIC-IDS2017, and the contemporary CICIoT2023. This collection ensures a rigorous evaluation across diverse network topologies, traffic profiles, and adversarial tactics. To ensure a consistent and unbiased evaluation, each dataset was partitioned into a training set (80%) and a testing set (20%) using a stratified sampling approach, thereby preserving the original class distribution in both subsets.</p>
<p>We first utilized the NSL-KDD dataset [<xref ref-type="bibr" rid="ref-27">27</xref>], a foundational benchmark for comparative analysis comprising a 41-feature vector. As detailed in <xref ref-type="table" rid="table-4">Table 4</xref>, the dataset exhibits severe class imbalance, with critical minority attack classes like U2R (0.17%) and R2L (2.50%) being heavily underrepresented.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>NSL-KDD dataset attack types and distribution</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attack type</th>
<th>Count</th>
<th>Percentage (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>77,054</td>
<td>52.75</td>
</tr>
<tr>
<td>DoS</td>
<td>53,385</td>
<td>36.54</td>
</tr>
<tr>
<td>Probe</td>
<td>14,077</td>
<td>9.63</td>
</tr>
<tr>
<td>R2L</td>
<td>3649</td>
<td>2.50</td>
</tr>
<tr>
<td>U2R</td>
<td>252</td>
<td>0.17</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To assess performance against modern threats, we incorporate the UNSW-NB15 dataset [<xref ref-type="bibr" rid="ref-28">28</xref>], which provides 42 features per record. Its class distribution, shown in <xref ref-type="table" rid="table-5">Table 5</xref>, presents a significant imbalance challenge, with minority classes such as Worms (0.08%) and Shellcode (0.69%) being exceptionally rare.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>UNSW-NB15 dataset attack types and distribution</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attack type</th>
<th>Count</th>
<th>Percentage (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>93,000</td>
<td>42.61</td>
</tr>
<tr>
<td>Generic</td>
<td>58,871</td>
<td>26.97</td>
</tr>
<tr>
<td>Exploits</td>
<td>44,525</td>
<td>20.41</td>
</tr>
<tr>
<td>Fuzzers</td>
<td>24,246</td>
<td>11.11</td>
</tr>
<tr>
<td>DoS</td>
<td>16,353</td>
<td>7.49</td>
</tr>
<tr>
<td>Reconnaissance</td>
<td>13,987</td>
<td>6.41</td>
</tr>
<tr>
<td>Analysis</td>
<td>2677</td>
<td>1.23</td>
</tr>
<tr>
<td>Backdoor</td>
<td>2329</td>
<td>1.07</td>
</tr>
<tr>
<td>Shellcode</td>
<td>1511</td>
<td>0.69</td>
</tr>
<tr>
<td>Worms</td>
<td>174</td>
<td>0.08</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For a test of real-world applicability, we employed the CIC-IDS2017 dataset [<xref ref-type="bibr" rid="ref-29">29</xref>], noted for its high-dimensional (80&#x002B; features) space. As shown in <xref ref-type="table" rid="table-6">Table 6</xref>, this dataset is heavily skewed, posing a challenge in detecting scarce minority attacks like Web and Bot, each representing only 0.07% of the data.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>CIC-IDS2017 dataset attack types and distribution</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attack type</th>
<th>Count</th>
<th>Percentage (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>2,271,320</td>
<td>77.26</td>
</tr>
<tr>
<td>DoS</td>
<td>379,737</td>
<td>12.92</td>
</tr>
<tr>
<td>PortScan</td>
<td>158,804</td>
<td>5.40</td>
</tr>
<tr>
<td>Patator</td>
<td>13,832</td>
<td>0.47</td>
</tr>
<tr>
<td>Web</td>
<td>2180</td>
<td>0.07</td>
</tr>
<tr>
<td>Bot</td>
<td>1956</td>
<td>0.07</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, to evaluate scalability on large-scale IoT traffic, we incorporate the CICIoT2023 dataset [<xref ref-type="bibr" rid="ref-30">30</xref>]. This benchmark consists of 46 features per record. It presents a unique imbalance structure, as shown in <xref ref-type="table" rid="table-7">Table 7</xref>, coexisting dominant attack classes with severe minority classes like BruteForce (0.23%) and Web (0.43%).</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>CICIot2023 dataset attack types and distribution</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attack type</th>
<th>Count</th>
<th>Percentage (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Benign</td>
<td>1,098,195</td>
<td>19.18</td>
</tr>
<tr>
<td>DDoS</td>
<td>1,250,000</td>
<td>21.83</td>
</tr>
<tr>
<td>DoS</td>
<td>1,250,000</td>
<td>21.83</td>
</tr>
<tr>
<td>Mirai</td>
<td>1,250,000</td>
<td>21.83</td>
</tr>
<tr>
<td>Spoofing</td>
<td>486,136</td>
<td>8.49</td>
</tr>
<tr>
<td>Recon</td>
<td>354,565</td>
<td>6.19</td>
</tr>
<tr>
<td>BruteForce</td>
<td>13,064</td>
<td>0.23</td>
</tr>
<tr>
<td>Web</td>
<td>24,829</td>
<td>0.43</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Data Preprocessing</title>
<p>We implemented a multi-stage data preprocessing pipeline involving feature selection, categorical encoding, and feature scaling. We performed feature selection to remove non-informative attributes. For NSL-KDD, all 41 original features were used. For UNSW-NB15, the id and attack_cat features were removed, retaining 42 features. For CIC-IDS2017, we removed 6 metadata features (Flow ID, Source IP, Source Port, Destination IP, Protocol, Timestamp), resulting in a 78-feature set. For CICIoT2023, all 46 features were utilized. Subsequently, categorical attributes were transformed using one-hot encoding. This step expanded the NSL-KDD dataset from 41 to 122 dimensions and the UNSW-NB15 dataset from 42 to 196 dimensions. The remaining 78 features of CIC-IDS2017 and all 46 features of CICIoT2023 are numerical and did not require this encoding step.</p>
<p>Data quality was addressed by imputing any missing or non-finite feature values with zero. To mitigate the biasing effects of disparate feature scales, all numerical data were standardized using Min&#x2013;Max normalization. This transformation was applied per-feature (i.e., column-wise), re-scaling each feature element <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> to <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> within the uniform interval <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> via the transformation:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represent the global minimum and maximum values for that specific feature across the dataset, respectively. This protocol is critical for stable model convergence.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Evaluation Metrics</title>
<p>We evaluate model performance based on the standard confusion matrix outcomes: True Positives (TP), True Negatives (TN), False Positives (FP, or Type I errors), and False Negatives (FN, or Type II errors), which are considered on a per-class basis for our multi-class task. While we report overall Accuracy, its utility is limited in imbalanced datasets where a high score may simply reflect strong performance on the dominant (benign) class. <disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Therefore, our primary evaluation focuses on Precision and Recall. Precision measures the fidelity of alerts, quantifying the proportion of predicted attacks that are genuinely malicious; this is critical for minimizing the operational cost of investigating false alarms.
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Recall (or Sensitivity) measures detection completeness, indicating the proportion of actual attacks the model successfully identifies; this is essential for ensuring threats are not missed.
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To account for the inherent trade-off between Precision and Recall, we employ the F1-Score. As the harmonic mean of Precision and Recall, the F1-Score provides a single, balanced measure of performance. This metric is particularly well-suited for imbalanced classification tasks (as opposed to Accuracy) because it does not depend on True Negatives, thus focusing evaluation on the positive (minority) classes.
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mtext>F1-Score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To visualize model performance independent of a specific classification threshold, we employ the Receiver Operating Characteristic (ROC) curve. The ROC curve is a graphical plot illustrating the trade-off between two key metrics&#x2014;the True Positive Rate (TPR) and the False Positive Rate (FPR)&#x2014;calculated across a range of decision thresholds. The TPR, which is mathematically equivalent to Recall, measures the proportion of actual attacks (positives) that are correctly identified.
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mtext>TPR (Recall)</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Conversely, the FPR measures the proportion of benign instances (negatives) that are incorrectly classified as attacks.
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mtext>FPR</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Data Imbalance Handling</title>
<p>A critical and prevalent challenge in network intrusion detection is the severe class imbalance inherent in realistic datasets, a characteristic clearly delineated in our datasets. On benchmark datasets, minority attack classes like R2L, U2R, and Worms are vastly underrepresented compared to normal traffic and common attacks. This imbalance can bias the model during training, leading to high accuracy on majority classes but dangerously poor recall on rare, yet often critical, threats [<xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<p>To mitigate this issue, we moved beyond conventional oversampling techniques like Synthetic Minority Over-sampling TEchnique (SMOTE) [<xref ref-type="bibr" rid="ref-32">32</xref>] or ADASYN [<xref ref-type="bibr" rid="ref-33">33</xref>], which generate synthetic samples through linear interpolation and may not fully capture the complex, non-linear distribution of attack data. Instead, we employed a sophisticated data augmentation strategy based on a Generative Adversarial Network (GAN) [<xref ref-type="bibr" rid="ref-34">34</xref>], specifically a variant of the Wasserstein GAN (WGAN). This approach is inspired by recent advancements that have successfully used generative models to create high-fidelity synthetic attack data for NIDS [<xref ref-type="bibr" rid="ref-23">23</xref>]. The WGAN architecture was chosen for its training stability and its ability to overcome the mode collapse problem common in standard GANs, thereby ensuring the generation of diverse and realistic minority samples [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>This process is visualized on a 2D plane, using the Web minority class from the CIC-IOT2023 dataset as a case study. As illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the visualization provides strong support for our methodological choice. The WGAN-generated samples (blue dots) demonstrate a distribution that closely mirrors the original &#x2018;Real Samples&#x2019; (red dots), clustering in similar regions and respecting the underlying data manifold. The SMOTE-generated samples (green dots), while also providing augmentation, show a distribution that does not align as closely with the true data structure compared to the WGAN samples.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title><italic>t</italic>-SNE (<italic>t</italic>-distributed Stochastic Neighbor Embedding) visualization comparing Real Samples (red), WGAN-generated Samples (blue), and SMOTE-generated Samples (green) for the &#x2018;Web&#x2019; attack class from the CIC-IOT2023 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-4.tif"/>
</fig>
<p>The WGAN framework replaces the Jensen-Shannon divergence with the Wasserstein-1 distance (<inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>), which provides a smoother and more meaningful loss metric, leading to more stable training and higher-quality synthetic samples. Our process for augmenting each minority class, illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, begins by training a dedicated WGAN model. The training follows an adversarial process governed by two distinct loss functions based on the Kantorovich-Rubinstein duality, which expresses the Wasserstein distance as: <disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>W</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">sup</mml:mo><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>f</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>L</mml:mi></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> is the real data distribution (the original minority class samples), <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> is the generator&#x2019;s distribution that models <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula>, and the supremum is taken over all 1-Lipschitz functions <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>f</mml:mi></mml:math></inline-formula>. In our implementation, the function <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>f</mml:mi></mml:math></inline-formula> is approximated by a neural network with weights <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>w</mml:mi></mml:math></inline-formula>, referred to as the critic. The 1-Lipschitz constraint is enforced through weight clipping, where the weights <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>w</mml:mi></mml:math></inline-formula> are clamped to a small range (e.g., <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>0.01</mml:mn><mml:mo>,</mml:mo><mml:mn>0.01</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>) after each gradient update.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>WGAN architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-5.tif"/>
</fig>
<p>Our process involved training a dedicated WGAN model for each minority attack class. The training follows an adversarial process governed by two distinct loss functions. The critic is trained to maximize the objective in <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref>, which corresponds to minimizing the following loss:
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Critic</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula></p>
<p>Conversely, the generator is trained to produce samples that are indistinguishable from real data, which corresponds to minimizing its loss function:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Generator</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">P</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula></p>
<p>Training for each per-class WGAN was continued until the Wasserstein distance estimate converged. The trained generator was then used to synthesize the new samples to create a more balanced class distribution for training. This targeted augmentation ensures that our TransNeSt model is exposed to a richer and more varied representation of rare attack patterns, enhancing its ability to learn their distinguishing features.</p>
<p>To experimentally quantify the impact of our WGAN-based augmentation. We evaluated the performance of our TransNeSt model on the minority attack classes across all four datasets under these three augmentation strategies. The results of this comparison are detailed in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Comparative analysis of data augmentation techniques on minority class performance</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> 
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Class</th>
<th colspan="3">Precision</th>
<th colspan="3">Recall</th>
<th colspan="3">F1-Score</th>
</tr>
<tr>
<th></th>
<th></th>
<th>Original</th>
<th>SMOTE</th>
<th>WGAN (Ours)</th>
<th>Original</th>
<th>SMOTE</th>
<th>WGAN (Ours)</th>
<th>Original</th>
<th>SMOTE</th>
<th>WGAN (Ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">NSL-KDD</td>
<td>R2L</td>
<td>0.7732</td>
<td>0.8711</td>
<td><bold>0.9175</bold></td>
<td>0.7898</td>
<td>0.8711</td>
<td><bold>0.8715</bold></td>
<td>0.7814</td>
<td>0.8711</td>
<td><bold>0.8939</bold></td>
</tr>
<tr>

<td>U2R</td>
<td>0.5455</td>
<td>0.5833</td>
<td><bold>0.7083</bold></td>
<td>0.4167</td>
<td>0.5417</td>
<td><bold>0.6071</bold></td>
<td>0.4728</td>
<td>0.5617</td>
<td><bold>0.6538</bold></td>
</tr>
<tr>
<td rowspan="4">UNSW-NB15</td>
<td>Worms</td>
<td><bold>0.7692</bold></td>
<td>0.5429</td>
<td>0.6667</td>
<td>0.2857</td>
<td>0.5143</td>
<td><bold>0.6286</bold></td>
<td>0.4167</td>
<td>0.5282</td>
<td><bold>0.6471</bold></td>
</tr>
<tr>

<td>Shellcode</td>
<td>0.6058</td>
<td>0.6167</td>
<td><bold>0.7169</bold></td>
<td>0.5497</td>
<td>0.6167</td>
<td><bold>0.6457</bold></td>
<td>0.5764</td>
<td>0.6167</td>
<td><bold>0.6794</bold></td>
</tr>
<tr>

<td>Backdoor</td>
<td>0.2584</td>
<td>0.4096</td>
<td><bold>0.6007</bold></td>
<td>0.1652</td>
<td>0.4096</td>
<td><bold>0.6910</bold></td>
<td>0.2016</td>
<td>0.4096</td>
<td><bold>0.6427</bold></td>
</tr>
<tr>

<td>Analysis</td>
<td><bold>0.6427</bold></td>
<td>0.4458</td>
<td>0.6106</td>
<td>0.1159</td>
<td><bold>0.6957</bold></td>
<td>0.6449</td>
<td>0.1959</td>
<td>0.5434</td>
<td><bold>0.6273</bold></td>
</tr>
<tr>
<td>CICIDS-2017</td>
<td>Bot</td>
<td>0.6151</td>
<td>0.7262</td>
<td><bold>0.8028</bold></td>
<td>0.4783</td>
<td><bold>0.7801</bold></td>
<td>0.7527</td>
<td>0.5381</td>
<td>0.7522</td>
<td><bold>0.7769</bold></td>
</tr>
<tr>
<td rowspan="2">CIC-IOT2023</td>
<td>BruteForce</td>
<td>0.6621</td>
<td>0.7921</td>
<td><bold>0.8954</bold></td>
<td>0.5978</td>
<td>0.7081</td>
<td><bold>0.7635</bold></td>
<td>0.6283</td>
<td>0.7477</td>
<td><bold>0.8242</bold></td>
</tr>
<tr>

<td>Web</td>
<td>0.6329</td>
<td>0.7183</td>
<td><bold>0.8439</bold></td>
<td>0.6221</td>
<td>0.7183</td>
<td><bold>0.7479</bold></td>
<td>0.6275</td>
<td>0.7183</td>
<td><bold>0.7930</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-8fn1" fn-type="other">
<p>Note: Bold values indicate the best performance among the comparison methods.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The data clearly demonstrate the superiority of the WGAN (Ours) approach. Across the vast majority of minority classes, our WGAN-augmented model achieves the highest scores in Precision, Recall, and F1-Score. By addressing the data imbalance at its source, we aim to significantly improve the model&#x2019;s recall on minority classes without compromising precision, ultimately fostering a more robust and reliable detection system.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experiment and Result Analysis</title>
<p>To rigorously evaluate the proposed TransNeSt architecture, our experimental design incorporates a controlled ablation study. We benchmarked TransNeSt against its core constituent components&#x2014;a standalone ResNeSt and a standalone Transformer&#x2014;to empirically validate that the hybrid design offers a synergistic improvement over its individual parts. Furthermore, to demonstrate that our specific fusion strategy is superior to conventional methods, we included a generic ResNet-Transformer hybrid as an additional challenging baseline. This comparative analysis was performed on four standard network intrusion detection datasets (NSL-KDD, UNSW-NB15, CICIDS-2017 and CICIOT2023), with model efficacy quantified by Accuracy, Precision, Recall, and F1-Score. It is critical to note that to ensure a fair comparison, all models discussed in this section were trained and evaluated on the same WGAN-augmented datasets.</p>
<p>All experiments were conducted on a server running Ubuntu 22.04. The software environment was built upon Python 3.10, PyTorch 2.1.2, and CUDA 11.8. The hardware platform consisted of a 12-core Intel(R) Xeon(R) Silver 4214R CPU @ 2.40 GHz and a single NVIDIA GeForce RTX 3080 Ti GPU with 12 GB of VRAM.</p>
<p>Before presenting the detection performance results, we first provide a quantitative analysis of each model&#x2019;s computational complexity and efficiency, including its total parameters and Floating Point Operations (FLOPs). The key metrics for this benchmark are summarized in <xref ref-type="table" rid="table-9">Table 9</xref>.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Comparison of model complexity and efficiency metrics</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Total Parameters (M)</th>
<th>Inference Time per 1M Samples (s)</th>
<th>FLOPs per 1M Samples (MFLOPs)</th>
</tr>
</thead>
<tbody>
<tr>
<td>ResNeSt</td>
<td>0.478</td>
<td>8.322</td>
<td>2374.24</td>
</tr>
<tr>
<td>Transformer</td>
<td>0.523</td>
<td>9.320</td>
<td>2667.98</td>
</tr>
<tr>
<td>ResNet-Transformer</td>
<td>0.764</td>
<td>11.610</td>
<td>3393.45</td>
</tr>
<tr>
<td>TransNeSt</td>
<td>0.780</td>
<td>11.961</td>
<td>3483.71</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The data in <xref ref-type="table" rid="table-9">Table 9</xref> establishes the computational cost for each architecture. As observed, our proposed TransNeSt model maintains a complexity profile comparable to the standard ResNet-Transformer hybrid.</p>

<sec id="s5_1">
<label>5.1</label>
<title>Evaluation Using the NSL-KDD Dataset</title>
<p>We initiated our empirical evaluation on the foundational NSL-KDD benchmark. The comprehensive performance comparison, illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, clearly establishes the superiority of the proposed TransNeSt architecture. Our model achieved the highest performance across all metrics, attaining an F1-Score of 99.04% and an accuracy of 99.05%.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Performance comparison of all models on the NSL-KDD dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-6.tif"/>
</fig>
<p>The results also serve as a compelling ablation study, isolating the contributions of our model&#x2019;s core components. The standalone ResNeSt model, acting as a powerful feature extraction baseline, achieved a remarkable F1-score of 97.30%, demonstrating the inherent effectiveness of the split-attention mechanism in capturing salient traffic features. The standalone Transformer, while also performing well, was surpassed by the ResNeSt, underscoring the importance of robust initial feature representation. Crucially, our integrated TransNeSt model outperformed both individual components, validating the synergistic benefit of combining ResNeSt&#x2019;s fine-grained feature extraction with the Transformer&#x2019;s sequential modeling capabilities.</p>
<p>A deeper analysis of per-class performance, detailed in the confusion matrix (<xref ref-type="fig" rid="fig-7">Fig. 7</xref>), ROC curves (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>), and performance metrics (<xref ref-type="table" rid="table-10">Table 10</xref>), highlights the model&#x2019;s robust performance. TransNeSt achieves outstanding F1-scores on majority classes like &#x2018;DoS&#x2019; (99.84%) and &#x2018;Normal&#x2019; (99.11%). More critically, this strength extends to the challenging minority attack classes, a success largely attributable to our augmentation strategy. As the comparative analysis in <xref ref-type="table" rid="table-8">Table 8</xref> empirically demonstrates, our WGAN-based approach yields performance gains over both the baseline and traditional SMOTE. For the extremely rare U2R class, WGAN elevates the F1-Score to 0.6538, surpassing the baseline&#x2019;s 0.4728 and SMOTE&#x2019;s 0.5617, thereby validating its efficacy in generating high-fidelity synthetic data.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix for TransNeSt on the NSL-KDD dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>ROC curve analysis for TransNeSt on the NSL-KDD dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-8.tif"/>
</fig><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Performance metrics on the NSL-KDD dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Class</th>
<th>TP</th>
<th>FP</th>
<th>FN</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>15,243</td>
<td>168</td>
<td>105</td>
<td>0.9891</td>
<td>0.9932</td>
<td>0.9911</td>
<td>15,348</td>
</tr>
<tr>
<td>DoS</td>
<td>10,662</td>
<td>16</td>
<td>18</td>
<td>0.9985</td>
<td>0.9983</td>
<td>0.9984</td>
<td>10,680</td>
</tr>
<tr>
<td>Probe</td>
<td>2788</td>
<td>27</td>
<td>43</td>
<td>0.9904</td>
<td>0.9848</td>
<td>0.9876</td>
<td>2831</td>
</tr>
<tr>
<td>R2L</td>
<td>712</td>
<td>64</td>
<td>105</td>
<td>0.9175</td>
<td>0.8715</td>
<td>0.8939</td>
<td>817</td>
</tr>
<tr>
<td>U2R</td>
<td>17</td>
<td>7</td>
<td>11</td>
<td>0.7083</td>
<td>0.6071</td>
<td>0.6538</td>
<td>28</td>
</tr>
<tr>
<td>Macro-Avg</td>
<td>29,422</td>
<td>282</td>
<td>282</td>
<td>0.9208</td>
<td>0.8910</td>
<td>0.9050</td>
<td>29,704</td>
</tr>
<tr>
<td>Micro-Avg</td>
<td>29,422</td>
<td>282</td>
<td>282</td>
<td>0.9904</td>
<td>0.9905</td>
<td>0.9904</td>
<td>29,704</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Performance on R2L (89.39% F1) and U2R (65.38% F1) remain comparatively lower than that of the high-support classes. This challenge is attributed to the extreme data scarcity and low feature diversity of these attack instances. It suggests that while WGAN provides a critical performance boost, fully modeling the feature boundaries of these subtle attacks remains the primary difficulty and a key focus for subsequent research.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Evaluation Using the UNSW-NB15 Dataset</title>
<p>To assess the model&#x2019;s efficacy against more contemporary and diverse threats, we conducted evaluations on the UNSW-NB15 dataset. As shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, TransNeSt once again significantly outperformed all baseline models, achieving a top-tier F1-Score of 91.92% and an accuracy of 91.90%. This F1-Score, while lower than the 99.04% achieved on NSL-KDD, reflects the distinctly higher complexity, greater class imbalance, and more subtle feature patterns inherent in this modern and more challenging dataset.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Performance comparison of all models on the UNSW-NB15 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-9.tif"/>
</fig>
<p>This dataset provides further evidence of our architecture&#x2019;s successful design. The standalone ResNeSt again proved to be a potent feature extractor with an F1-score of 88.03%, reinforcing the value of its multi-path, channel-aware attention structure for discerning modern attack patterns. The subsequent integration with the Transformer block in TransNeSt yielded a substantial performance gain of nearly 4 percentage points in F1-score. This improvement highlights the critical role of the Transformer in modeling the temporal relationships between the rich feature sets generated by ResNeSt, a synergy that is essential for handling the complexity of the UNSW-NB15 dataset.</p>
<p>The detailed per-class metrics are presented in <xref ref-type="table" rid="table-11">Table 11</xref>, with corresponding confusion matrix in <xref ref-type="fig" rid="fig-10">Fig. 10</xref> and the ROC curves in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. The model demonstrates high F1-scores on prevalent classes like &#x2018;Normal&#x2019; (97.67%) and &#x2018;Generic&#x2019; (97.46%). For the challenging low-support classes, the WGAN-based augmentation provides a clear advantage, as detailed in <xref ref-type="table" rid="table-8">Table 8</xref>. WGAN, for example, raised the F1-Score for &#x2018;Backdoor&#x2019; to 0.6427 (from 0.2016 baseline) and &#x2018;Shellcode&#x2019; to 0.6794 (from 0.5764 baseline). Despite this improvement, these minority classes (e.g., &#x2018;Analysis&#x2019;, &#x2018;Backdoor&#x2019;, &#x2018;Shellcode&#x2019;, &#x2018;Worms&#x2019;) still yield the lowest F1-scores (62%&#x2013;68% range). This indicates that while WGAN provides a critical boost, the low instance count and feature overlap remain the primary challenge for model refinement on this dataset.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Performance metrics on the UNSW-NB15 dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Class</th>
<th>TP</th>
<th>FP</th>
<th>FN</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>18,037</td>
<td>298</td>
<td>563</td>
<td>0.9837</td>
<td>0.9697</td>
<td>0.9767</td>
<td>18,600</td>
</tr>
<tr>
<td>Generic</td>
<td>11,580</td>
<td>410</td>
<td>194</td>
<td>0.9658</td>
<td>0.9835</td>
<td>0.9746</td>
<td>11,774</td>
</tr>
<tr>
<td>Exploits</td>
<td>7789</td>
<td>906</td>
<td>1116</td>
<td>0.8958</td>
<td>0.8747</td>
<td>0.8851</td>
<td>8905</td>
</tr>
<tr>
<td>Fuzzers</td>
<td>4188</td>
<td>853</td>
<td>661</td>
<td>0.8308</td>
<td>0.8637</td>
<td>0.8469</td>
<td>4849</td>
</tr>
<tr>
<td>DoS</td>
<td>2507</td>
<td>653</td>
<td>764</td>
<td>0.7934</td>
<td>0.7664</td>
<td>0.7797</td>
<td>3271</td>
</tr>
<tr>
<td>Reconnaissance</td>
<td>2376</td>
<td>532</td>
<td>422</td>
<td>0.8171</td>
<td>0.8492</td>
<td>0.8328</td>
<td>2798</td>
</tr>
<tr>
<td>Analysis</td>
<td>345</td>
<td>220</td>
<td>190</td>
<td>0.6106</td>
<td>0.6449</td>
<td>0.6273</td>
<td>535</td>
</tr>
<tr>
<td>Backdoor</td>
<td>322</td>
<td>214</td>
<td>144</td>
<td>0.6007</td>
<td>0.6910</td>
<td>0.6427</td>
<td>466</td>
</tr>
<tr>
<td>Shellcode</td>
<td>195</td>
<td>77</td>
<td>107</td>
<td>0.7169</td>
<td>0.6457</td>
<td>0.6794</td>
<td>302</td>
</tr>
<tr>
<td>Worms</td>
<td>22</td>
<td>11</td>
<td>13</td>
<td>0.6667</td>
<td>0.6286</td>
<td>0.6471</td>
<td>35</td>
</tr>
<tr>
<td>Macro-Avg</td>
<td>47,361</td>
<td>4174</td>
<td>4174</td>
<td>0.7882</td>
<td>0.7917</td>
<td>0.7892</td>
<td>51,535</td>
</tr>
<tr>
<td>Micro-Avg</td>
<td>47,361</td>
<td>4174</td>
<td>4174</td>
<td>0.9198</td>
<td>0.9190</td>
<td>0.9192</td>
<td>51,535</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Confusion matrix for TransNeSt on the UNSW-NB15 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>ROC curve analysis for TransNeSt on the UNSW-NB15 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-11.tif"/>
</fig>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Evaluation Using the CIC-IDS2017 Dataset</title>
<p>A rigorous evaluation was performed on the large-scale and highly realistic CIC-IDS2017 dataset. The results, illustrated in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>, demonstrate the model&#x2019;s high performance. TransNeSt achieved an F1-Score of 99.18% and an accuracy of 99.19%, outperforming the other tested models. The ablation analysis on this dataset further solidifies our central thesis. The ResNeSt component, with an F1-score of 97.52%, proves itself to be a highly effective feature extractor even on high-volume, complex traffic. The Transformer component also performs well, but the fusion in TransNeSt again elevates the performance to a new level, surpassing the next-best model (ResNet-Transformer) by over 1.5 percentage points. This margin is significant at such high performance levels and underscores the superiority of our specific architectural combination.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Performance comparison of all models on the CIC-IDS2017 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-12.tif"/>
</fig>
<p>Given the severe class imbalance of this dataset, a detailed per-class analysis is crucial (<xref ref-type="table" rid="table-12">Table 12</xref>, <xref ref-type="fig" rid="fig-13">Figs. 13</xref>, and <xref ref-type="fig" rid="fig-14">14</xref>). TransNeSt exhibits near-perfect detection on dominant classes like &#x2018;Normal&#x2019; (99.59% F1) and &#x2018;DoS&#x2019; (97.64% F1). Critically, it also shows strong performance on less frequent attack types, such as the 98.05% F1-score on &#x2018;Web&#x2019;. For the &#x2018;Bot&#x2019; class (less than 0.07% of data), the WGAN augmentation strategy proved essential. As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, WGAN raised the &#x2018;Bot&#x2019; F1-Score to 0.7769, outperforming both the baseline (0.5381) and SMOTE (0.7522).</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Performance metrics on the CICIDS-2017 dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Class</th>
<th>TP</th>
<th>FP</th>
<th>FN</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Normal</td>
<td>452,422</td>
<td>2442</td>
<td>1305</td>
<td>0.9946</td>
<td>0.9971</td>
<td>0.9959</td>
<td>453,727</td>
</tr>
<tr>
<td>DoS</td>
<td>73,842</td>
<td>1647</td>
<td>1917</td>
<td>0.9782</td>
<td>0.9747</td>
<td>0.9764</td>
<td>75,759</td>
</tr>
<tr>
<td>PortScan</td>
<td>30,705</td>
<td>414</td>
<td>1236</td>
<td>0.9867</td>
<td>0.9613</td>
<td>0.9738</td>
<td>31,941</td>
</tr>
<tr>
<td>Patator</td>
<td>2761</td>
<td>5</td>
<td>20</td>
<td>0.9982</td>
<td>0.9928</td>
<td>0.9955</td>
<td>2781</td>
</tr>
<tr>
<td>Web</td>
<td>428</td>
<td>8</td>
<td>9</td>
<td>0.9817</td>
<td>0.9794</td>
<td>0.9805</td>
<td>437</td>
</tr>
<tr>
<td>Bot</td>
<td>350</td>
<td>86</td>
<td>115</td>
<td>0.8028</td>
<td>0.7527</td>
<td>0.7769</td>
<td>465</td>
</tr>
<tr>
<td>Macro-Avg</td>
<td>560,508</td>
<td>4602</td>
<td>4602</td>
<td>0.9570</td>
<td>0.9430</td>
<td>0.9498</td>
<td>565,110</td>
</tr>
<tr>
<td>Micro-Avg</td>
<td>560,508</td>
<td>4602</td>
<td>4602</td>
<td>0.9918</td>
<td>0.9919</td>
<td>0.9918</td>
<td>565,110</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Confusion matrix for TransNeSt on the CICIDS-2017 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-13.tif"/>
</fig><fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>ROC curve analysis for TransNeSt on the CIC-IDS2017 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-14.tif"/>
</fig>
<p>The &#x2018;Bot&#x2019; class performance, while superior to other augmentation methods, remains the primary area for refinement. The 115 False Negatives identified for this class in (<xref ref-type="fig" rid="fig-13">Fig. 13</xref>) are a key indicator. These are primarily misclassifications against the high-volume &#x2018;Normal&#x2019; class. This highlights the persistent difficulty of distinguishing extremely low-support malicious flows from benign traffic, a challenge that persists even with effective data augmentation.</p>

</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Evaluation Using the CICIoT2023 Dataset</title>
<p>To validate our model&#x2019;s scalability and performance on a very recent, large-scale IoT benchmark, we performed a final evaluation on the CICIoT2023 dataset. The comprehensive performance comparison, illustrated in <xref ref-type="fig" rid="fig-15">Fig. 15</xref>, confirms the superiority of our proposed architecture. TransNeSt achieved the highest F1-Score of 97.85% and an accuracy of 97.85%, again clearly outperforming the other models. This dataset&#x2019;s ablation results further reinforce our design. The standalone ResNeSt (95.23% F1) and Transformer (95.44% F1) components performed well. However, their fusion in TransNeSt (97.85% F1) yielded a performance gain of over 1.6 percentage points compared to the next-best ResNet-Transformer (96.21% F1), underscoring the synergistic benefit of combining ResNeSt&#x2019;s fine-grained feature extraction with the Transformer&#x2019;s sequential modeling.</p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Performance comparison of all models on the CICIOT2023 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-15.tif"/>
</fig>
<p>A detailed per-class analysis is crucial for this highly imbalanced dataset (<xref ref-type="table" rid="table-13">Table 13</xref>, <xref ref-type="fig" rid="fig-16">Figs. 16</xref> and <xref ref-type="fig" rid="fig-17">17</xref>). The model demonstrates exceptional performance on high-volume attack classes, achieving F1-scores of 99.68% (DDoS), 99.62% (DoS), and 99.87% (Mirai). For the challenging minority classes, &#x2018;BruteForce&#x2019; and &#x2018;Web&#x2019;, the WGAN-based augmentation strategy was essential. As quantified in <xref ref-type="table" rid="table-8">Table 8</xref>, WGAN elevated the F1-Score for &#x2018;BruteForce&#x2019; to 0.8242 (from a 0.6283 baseline) and &#x2018;Web&#x2019; to 0.7930 (from a 0.6275 baseline), validating its efficacy. While WGAN provides a clear boost, these minority classes remain the primary area for refinement. The metrics in <xref ref-type="table" rid="table-13">Table 13</xref> identify 618 False Negatives for &#x2018;BruteForce&#x2019; and 1252 for &#x2018;Web&#x2019;. A corresponding analysis of the confusion matrix (<xref ref-type="fig" rid="fig-17">Fig. 17</xref>) reveals that these misclassifications are not random; they are primarily confused with &#x2018;Spoofing&#x2019; and &#x2018;Recon&#x2019;. This indicates a high degree of feature overlap between these specific IoT attack types, posing a persistent challenge even with effective data augmentation.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Performance metrics on the CICIOT2023 dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Class</th>
<th>TP</th>
<th>FP</th>
<th>FN</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Benign</td>
<td>212,784</td>
<td>6151</td>
<td>6855</td>
<td>0.9719</td>
<td>0.9688</td>
<td>0.9703</td>
<td>219,639</td>
</tr>
<tr>
<td>DDoS</td>
<td>249,293</td>
<td>886</td>
<td>707</td>
<td>0.9965</td>
<td>0.9972</td>
<td>0.9968</td>
<td>250,000</td>
</tr>
<tr>
<td>DoS</td>
<td>249,253</td>
<td>1135</td>
<td>747</td>
<td>0.9955</td>
<td>0.9970</td>
<td>0.9962</td>
<td>250,000</td>
</tr>
<tr>
<td>Mirai</td>
<td>249,619</td>
<td>293</td>
<td>381</td>
<td>0.9988</td>
<td>0.9985</td>
<td>0.9987</td>
<td>250,000</td>
</tr>
<tr>
<td>Spoofing</td>
<td>89,776</td>
<td>6850</td>
<td>7525</td>
<td>0.9291</td>
<td>0.9227</td>
<td>0.9259</td>
<td>97,301</td>
</tr>
<tr>
<td>Recon</td>
<td>64,341</td>
<td>8422</td>
<td>6572</td>
<td>0.8843</td>
<td>0.9073</td>
<td>0.8956</td>
<td>70,913</td>
</tr>
<tr>
<td>BruteForce</td>
<td>1995</td>
<td>233</td>
<td>618</td>
<td>0.8954</td>
<td>0.7635</td>
<td>0.8242</td>
<td>2613</td>
</tr>
<tr>
<td>Web</td>
<td>3714</td>
<td>687</td>
<td>1252</td>
<td>0.8439</td>
<td>0.7479</td>
<td>0.7930</td>
<td>4966</td>
</tr>
<tr>
<td>Macro-Avg</td>
<td>1,120,775</td>
<td>24,657</td>
<td>24,657</td>
<td>0.9394</td>
<td>0.9129</td>
<td>0.9251</td>
<td>1,145,432</td>
</tr>
<tr>
<td>Micro-Avg</td>
<td>1,120,775</td>
<td>24,657</td>
<td>24,657</td>
<td>0.9785</td>
<td>0.9785</td>
<td>0.9785</td>
<td>1,145,432</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>ROC curve analysis for TransNeSt on the CICIOT2023 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-16.tif"/>
</fig><fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>Confusion matrix for TransNeSt on the CICIOT2023 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74349-fig-17.tif"/>
</fig>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Comparison with Existing Models</title>
<p>To contextualize the advancement offered by our proposed model, we compare its performance against several recent NIDS models on the same four benchmark datasets. This comparison serves to contextualize the performance of the TransNeSt architecture relative to other contemporary approaches. The comprehensive results of this comparative analysis are summarized in <xref ref-type="table" rid="table-14">Table 14</xref>.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Comparison results of TransNeSt with existing models</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Model</th>
<th>Year</th>
<th>Multiclass Accuracy (%)</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F1-Score (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4">NSL-KDD</td>
<td>AE&#x002B;PCA&#x002B;LSTM [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>2024</td>
<td>82.22</td>
<td>92.01</td>
<td>75.30</td>
<td>82.62</td>
</tr>
<tr>

<td>AE-WGAN [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>2024</td>
<td>93.00</td>
<td>96.00</td>
<td>94.00</td>
<td>94.00</td>
</tr>
<tr>

<td>MLF-DOU [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>2025</td>
<td>93.14</td>
<td>91.26</td>
<td>97.26</td>
<td>94.16</td>
</tr>
<tr>

<td>TransNeSt</td>
<td>&#x2013;</td>
<td><bold>99.05</bold></td>
<td><bold>99.04</bold></td>
<td><bold>99.05</bold></td>
<td><bold>99.04</bold></td>
</tr>
<tr>
<td rowspan="4">UNSW-NB15</td>
<td>AE&#x002B;PCA&#x002B;LSTM</td>
<td>2024</td>
<td>76.28</td>
<td>71.58</td>
<td><bold>94.39</bold></td>
<td>81.41</td>
</tr>
<tr>

<td>TSSAN [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>2024</td>
<td>&#x2013;</td>
<td>88.00</td>
<td>84.00</td>
<td>86.00</td>
</tr>
<tr>

<td>MLF-DOU</td>
<td>2025</td>
<td>90.68</td>
<td><bold>92.18</bold></td>
<td>90.78</td>
<td>91.47</td>
</tr>
<tr>

<td>TransNeSt</td>
<td>&#x2013;</td>
<td><bold>91.90</bold></td>
<td>91.98</td>
<td>91.90</td>
<td><bold>91.92</bold></td>
</tr>
<tr>
<td rowspan="5">CICIDS-2017</td>
<td>AE&#x002B;PCA&#x002B;LSTM</td>
<td>2024</td>
<td>93.78</td>
<td>96.57</td>
<td>95.65</td>
<td>96.11</td>
</tr>
<tr>

<td>TSSAN</td>
<td>2024</td>
<td>&#x2013;</td>
<td>88.00</td>
<td>97.00</td>
<td>92.00</td>
</tr>
<tr>

<td>AE-WGAN</td>
<td>2024</td>
<td>98.00</td>
<td>98.00</td>
<td>99.00</td>
<td>98.00</td>
</tr>
<tr>

<td>NIDS-CNNRF [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>2025</td>
<td>98.07</td>
<td>98.21</td>
<td>97.83</td>
<td>98.02</td>
</tr>
<tr>

<td>TransNeSt</td>
<td>&#x2013;</td>
<td><bold>99.19</bold></td>
<td><bold>99.18</bold></td>
<td><bold>99.19</bold></td>
<td><bold>99.18</bold></td>
</tr>
<tr>
<td rowspan="4">CICIOT2023</td>
<td>MACML [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>2025</td>
<td>91.05</td>
<td>92.23</td>
<td>91.05</td>
<td>90.98</td>
</tr>
<tr>

<td>RMCNN [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>2024</td>
<td>97.28</td>
<td>97.28</td>
<td>97.35</td>
<td>97.29</td>
</tr>
<tr>

<td>CTWA [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>2025</td>
<td>96.43</td>
<td>96.59</td>
<td>96.43</td>
<td>96.45</td>
</tr>
<tr>

<td>TransNeSt</td>
<td>&#x2013;</td>
<td><bold>97.85</bold></td>
<td><bold>97.85</bold></td>
<td><bold>97.85</bold></td>
<td><bold>97.85</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-14fn1" fn-type="other">
<p>Note: Bold values indicate the best performance in each category.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>On the foundational NSL-KDD benchmark, TransNeSt achieves the highest F1-Score (99.04%), demonstrating a superior balance of precision and recall compared to models like MLF-DOU [<xref ref-type="bibr" rid="ref-19">19</xref>]. This robust management of the precision-recall trade-off is also evident on the more complex UNSW-NB15 dataset, where TransNeSt again secures the top F1-Score (91.92%), surpassing other specialized models [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>The model&#x2019;s advantage is particularly pronounced on large-scale, modern datasets. On CICIDS-2017, TransNeSt (99.18% F1) establishes a clear margin of over 1.1 percentage points against the next-strongest competitor, NIDS-CNNRF [<xref ref-type="bibr" rid="ref-24">24</xref>]. This robust performance is confirmed on the recent CICIOT2023 IoT dataset, where our model (97.85% F1) again outperforms all contemporary methods, including the recent RMCNN [<xref ref-type="bibr" rid="ref-15">15</xref>] (97.29%) and CTWA [<xref ref-type="bibr" rid="ref-22">22</xref>] (96.45%).</p>
<p>In summary, the consistent outperformance of TransNeSt across these four diverse datasets against strong, contemporary models validates its innovative architectural design. The synergy between the powerful multi-scale feature representation from the ResNeSt module and the global contextual understanding from the Transformer allows our model to set a new standard for performance in intrusion detection.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This work confirms that the integration of 1D split-attention mechanisms with Transformer encoders presents a highly effective architecture for NIDS. Our key insight is that the ResNeSt block&#x2019;s fine-grained feature representation serves as a more potent input for a Transformer&#x2019;s global temporal analysis, directly addressing the common trade-off between multi-scale and temporal modeling.</p>
<p>Empirically, TransNeSt demonstrated highly effective and robust detection capabilities, achieving strong F1-Scores of 99.04% on NSL-KDD, 91.92% on UNSW-NB15, 99.18% on CIC-IDS2017, and 97.85% on CICIOT2023. These results were consistently favorable when compared to both its constituent components and existing state-of-the-art methods across all four datasets. This analysis presents an actionable takeaway: while our analysis identified TransNeSt as the most computationally intensive model tested. Its complexity remains comparable to standard hybrid architectures. This finding suggests the computational cost is a justified trade-off for achieving this high level of detection performance.</p>
<p>Despite these promising results, this study has limitations. The primary limitation concerns the high computational complexity of the split-attention mechanism, which may pose challenges for real-time deployment on resource-constrained edge devices. Additionally, as a deep learning-based approach, the model currently operates as a &#x2018;black box,&#x2019; lacking intrinsic interpretability to explain the rationale behind specific classification decisions. Looking forward, we identify several key directions. The primary limitation is computational cost, which motivates model optimization. Future work will investigate advanced compression techniques, including network pruning, quantization, and knowledge distillation, to create a lightweight, deployment-ready version for resource-constrained IoT networks. A second direction is to enhance model interpretability by integrating Explainable AI (XAI) techniques, specifically SHapley Additive exPlanations (SHAP) or Local Interpretable Model-agnostic Explanations (LIME), to demystify its decision-making process. Finally, we will investigate the adaptation of TransNeSt for federated learning (FL), enabling collaborative, privacy-preserving training across distributed networks.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to acknowledge that this work was sponsored by the Opening Foundation of Yunnan Key Laboratory of Smart City in Cyberspace Security.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Opening Foundation of Yunnan Key Laboratory of Smart City in Cyberspace Security (No. 202105AG070010).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Study conception and design: Gan Zhu, Yongtao Yu; data collection: Gan Zhu, Xiaofan Deng, Yuanchen Dai, Zhenyuan Li; analysis and interpretation of results: Gan Zhu, Yongtao Yu; draft manuscript preparation: Gan Zhu, Xiaofan Deng,Yuanchen Dai, Zhenyuan Li. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Study data can be obtained by contacting the corresponding author.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aldweesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Derhab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Emam</surname> <given-names>AZ</given-names></string-name></person-group>. <article-title>Deep learning approaches for anomaly-based intrusion detection systems: a survey, taxonomy, and open issues</article-title>. <source>Knowl Based Syst</source>. <year>2020</year>;<volume>189</volume>:<fpage>105124</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2019.105124</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fei</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A deep learning approach for intrusion detection using recurrent neural networks</article-title>. <source>IEEE Access</source>. <year>2017</year>;<volume>5</volume>:<fpage>21954</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2017.2762418</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Roy</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cheung</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A deep learning approach for intrusion detection in internet of things using bi-directional long short-term memory recurrent neural network</article-title>. In: <conf-name>2018 28th International Telecommunication Networks and Applications Conference (ITNAC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ATNAC.2018.8615294</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gueriani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kheddar</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mazari</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Ghanem</surname> <given-names>MC</given-names></string-name></person-group>. <article-title>A robust cross-domain IDS using BiGRU-LSTM-attention for medical and industrial IoT security</article-title>. <source>ICT Express</source>. <year>2025</year>;<volume>8</volume>(<issue>11</issue>):<fpage>8707</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.icte.2025.08.011</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vaswani</surname></string-name> <string-name> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname></string-name> <string-name> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname></string-name> <string-name> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname></string-name> <string-name> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname></string-name> <string-name> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname></string-name> <string-name> <given-names>AN</given-names></string-name></person-group>, <etal>et al.</etal> <source>Attention is all you need. In: Advances in neural information processing systems</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017</year>. <fpage>30</fpage> p.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kheddar</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Transformers and large language models for efficient intrusion detection systems: a comprehensive survey</article-title>. <source>Inf Fusion</source>. <year>2025</year>;<volume>124</volume>:<fpage>103347</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2025.103347</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A novel multi-scale network intrusion detection model with transformer</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>23239</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-74214-w</pub-id>; <pub-id pub-id-type="pmid">39369065</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Split-attention networks</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>2736</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW56347.2022.00309</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kolukisa</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dedeturk</surname> <given-names>BK</given-names></string-name>, <string-name><surname>Hacilar</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gungor</surname> <given-names>VC</given-names></string-name></person-group>. <article-title>An efficient network intrusion detection approach based on logistic regression model and parallel artificial bee colony algorithm</article-title>. <source>Comput Standards Interfaces</source>. <year>2024</year>;<volume>89</volume>:<fpage>103808</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.csi.2023.103808</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Turukmane</surname> <given-names>AV</given-names></string-name>, <string-name><surname>Devendiran</surname> <given-names>R</given-names></string-name></person-group>. <article-title>M-MultiSVM: an efficient feature selection assisted network intrusion detection system using machine learning</article-title>. <source>Comput Secur</source>. <year>2024</year>;<volume>137</volume>:<fpage>103587</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2023.103587</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rustam</surname> <given-names>F</given-names></string-name>, <string-name><surname>Aljedaani</surname> <given-names>W</given-names></string-name>, <string-name><surname>Elsayed</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Jurcut</surname> <given-names>AD</given-names></string-name></person-group>. <article-title>FAMTDS: a novel MFO-based fully automated malicious traffic detection system for multi-environment networks</article-title>. <source>Comput Netw</source>. <year>2024</year>;<volume>251</volume>:<fpage>110603</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comnet.2024.110603</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hooshmand</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Huchaiah</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Alzighaibi</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Hashim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Atlam</surname> <given-names>ES</given-names></string-name>, <string-name><surname>Gad</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Robust network anomaly detection using ensemble learning approach and explainable artificial intelligence (XAI)</article-title>. <source>Alexandria Eng J</source>. <year>2024</year>;<volume>94</volume>:<fpage>120</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aej.2024.03.041</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>X</given-names></string-name></person-group>. <article-title>ESVI-GaMM: a fast network intrusion detection approach based on the Bayesian gamma mixture model</article-title>. <source>Inf Sci</source>. <year>2024</year>;<volume>678</volume>:<fpage>121001</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ins.2024.121001</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>TSSAN: time-space separable attention network for intrusion detection</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>98734</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3429420</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>An intrusion detection model based on a residual memory convolutional neural network with attention mechanism</article-title>. <source>J Phys Conf Series</source>. <year>2024</year>;<volume>2833</volume>(<issue>1</issue>):<fpage>012009</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1742-6596/2833/1/012009</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name></person-group>. <article-title>MACML: marrying attention and convolution-based meta-learning method for few-shot IoT intrusion detection</article-title>. <source>PLoS One</source>. <year>2025</year>;<volume>20</volume>(<issue>8</issue>):<fpage>e0331065</fpage>. doi:<pub-id pub-id-type="doi">10.1371/journal.pone.0331065</pub-id>; <pub-id pub-id-type="pmid">40880393</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thakkar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kikani</surname> <given-names>N</given-names></string-name>, <string-name><surname>Geddam</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Fusion of linear and non-linear dimensionality reduction techniques for feature reduction in LSTM-based Intrusion Detection System</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>154</volume>:<fpage>111378</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.111378</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alsoufi</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Siraj</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Ghaleb</surname> <given-names>FA</given-names></string-name>, <string-name><surname>Al-Razgan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Al-Asaly</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Alfakih</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Anomaly-based intrusion detection model using deep learning for IoT networks</article-title>. <source>Comput Model Eng Sci</source>. <year>2024</year>;<volume>141</volume>(<issue>1</issue>):<fpage>823</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2024.052112</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>N</given-names></string-name></person-group>. <article-title>MLF-DOU: a metric learning framework with dual one-class units for network intrusion detection</article-title>. <source>Neurocomputing</source>. <year>2025</year>;<volume>649</volume>:<fpage>130754</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2025.130754</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>A two-layer network intrusion detection method incorporating LSTM and stacking ensemble learning</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>83</volume>(<issue>3</issue>):<fpage>5129</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.062094</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Deep Q-learning intrusion detection system (DQ-IDS): a novel reinforcement learning approach for adaptive and self-learning cybersecurity</article-title>. <source>ICT Express</source>. <year>2025</year>;<volume>11</volume>(<issue>5</issue>):<fpage>875</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.icte.2025.05.007</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>P</given-names></string-name></person-group>. <article-title>CTWA: a novel incremental deep learning-based intrusion detection method for the Internet of Things</article-title>. <source>Artif Intell Rev</source>. <year>2025</year>;<volume>58</volume>(<issue>12</issue>):<fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-025-11358-9</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arafah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Phillips</surname> <given-names>I</given-names></string-name>, <string-name><surname>Adnane</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hadi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Alauthman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Al-Banna</surname> <given-names>AK</given-names></string-name></person-group>. <article-title>Anomaly-based network intrusion detection using denoising autoencoder and Wasserstein GAN synthetic attacks</article-title>. <source>Appl Soft Comput</source>. <year>2025</year>;<volume>168</volume>:<fpage>112455</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112455</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>NIDS-CNNRF integrating CNN and random forest for efficient network intrusion detection model</article-title>. <source>Internet of Things</source>. <year>2025</year>;<volume>32</volume>:<fpage>101607</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2025.101607</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sennrich</surname> <given-names>R</given-names></string-name></person-group>. <chapter-title>Root mean square layer normalization</chapter-title>. In: <source>Advances in neural information processing systems</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2019</year>. <fpage>32</fpage> p.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tavallaee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bagheri</surname> <given-names>E</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>A detailed analysis of the KDD CUP 99 data set</article-title>. In: <conf-name>2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2009</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CISDA.2009.5356528</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Moustafa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Slay</surname> <given-names>J</given-names></string-name></person-group>. <article-title>UNSW-NB15: a comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set)</article-title>. In: <conf-name>2015 Military Communications and Information Systems Conference (MilCIS) 2015</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MilCIS.2015.7348942</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Panigrahi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Borah</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A detailed analysis of CICIDS2017 dataset for designing intrusion detection systems</article-title>. <source>Int J Eng Technol</source>. <year>2018</year>;<volume>7</volume>(<issue>3.24</issue>):<fpage>479</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.14419/ijet.v7i3.24.227971</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Neto</surname> <given-names>EC</given-names></string-name>, <string-name><surname>Dadkhah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zohourian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>CICIoT2023: a real-time dataset and benchmark for large-scale attacks in IoT environment</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>13</issue>):<fpage>5941</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23135941</pub-id>; <pub-id pub-id-type="pmid">37447792</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Qarni</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Al-Asmari</surname> <given-names>GA</given-names></string-name></person-group>. <article-title>Addressing imbalanced data in network intrusion detection: a review and survey</article-title>. <source>Int J Adv Comput Sci Appl</source>. <year>2024</year>;<volume>15</volume>(<issue>2</issue>):<fpage>136</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.14569/IJACSA.2024.0150215</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chawla</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Bowyer</surname> <given-names>KW</given-names></string-name>, <string-name><surname>Hall</surname> <given-names>LO</given-names></string-name>, <string-name><surname>Kegelmeyer</surname> <given-names>WP</given-names></string-name></person-group>. <article-title>SMOTE: synthetic minority over-sampling technique</article-title>. <source>J Artif Intell Res</source>. <year>2002</year>;<volume>16</volume>:<fpage>321</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1613/jair.953</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name></person-group>. <article-title>ADASYN: adaptive synthetic sampling approach for imbalanced learning</article-title>. In: <conf-name>2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence) 2008</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>1322</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/IJCNN.2008.4633969</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Pouget-Abadie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Warde-Farley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ozair</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>Generative adversarial nets</chapter-title>. In: <source>Advances in neural information processing systems</source>. <comment>arXiv:1406.2661. 2014</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1406.2661</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Arjovsky</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chintala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Wasserstein generative adversarial networks</article-title>. <comment>arXiv:1701.07875. 2017</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1701.07875</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>