<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">82810</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.082810</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Explainable Hierarchical Mamba for Edge-Based IoT Traffic Classification</article-title>
<alt-title alt-title-type="left-running-head">Explainable Hierarchical Mamba for Edge-Based IoT Traffic Classification</alt-title>
<alt-title alt-title-type="right-running-head">Explainable Hierarchical Mamba for Edge-Based IoT Traffic Classification</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Yu</surname><given-names>Jiangyong</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Hu</surname><given-names>Chuanping</given-names></name><email>cphu@zzu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Runnan</given-names></name></contrib>
<aff id="aff-1"><institution>School of Cyber Science and Engineering, Zhengzhou University</institution>, <addr-line>Zhengzhou</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Chuanping Hu. Email: <email>cphu@zzu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>59</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_82810.pdf"></self-uri>
<abstract>
<p>With the proliferation of Internet of Things (IoT) devices, accurate device fingerprinting of highly encrypted traffic has emerged as a critical challenge for ensuring network security. Existing deep learning models are either difficult to deploy in real-time due to excessive computational complexity (e.g., Transformers) or are limited in performance because their structure does not match the inherent hierarchy of traffic data (e.g., flattened state space models). Furthermore, a general lack of transparency in their decision-making processes restricts their trustworthiness in security-critical scenarios. To address these challenges, this paper proposes a Hierarchical Mamba with Gated Attribution Fingerprinting (HMX-GAF) framework. The framework explicitly models the intrinsic hierarchical structure of traffic data via a bespoke packet-flow dual-layer Mamba encoder, resolving the issue of architectural mismatch. Concurrently, it pioneers a zero-overhead Gated Attribution Fingerprinting (GAF) mechanism by leveraging the internal gating signals of the Mamba model, achieving high-fidelity intrinsic explainability. Comprehensive experiments on the CIC-IoT-2024 dataset demonstrate that HMX-GAF significantly outperforms current state-of-the-art models in both device identification and anomaly detection tasks, while maintaining the millisecond-level inference efficiency required for edge deployment.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Traffic classification</kwd>
<kwd>IoT security</kwd>
<kwd>device fingerprinting</kwd>
<kwd>state space model</kwd>
<kwd>Mamba</kwd>
<kwd>explainable AI</kwd>
<kwd>hierarchical modeling</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Smart speakers, security cameras, and network-enabled home appliances are rapidly transforming domestic spaces into highly networked micro-ecosystems. Market forecasts project that global smart-home revenue will leap from US $<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>84.5</mml:mn></mml:math></inline-formula> billion in 2024 to more than US $116 billion by 2029, reflecting both an extraordinary penetration rate and vast growth potential [<xref ref-type="bibr" rid="ref-1">1</xref>]. This trend not only drives an exponential increase in household network traffic but also lays a data foundation for subsequent intelligent services.</p>
<p>Beyond convenience, security and privacy threats proliferate in lockstep: malware-controlled cameras, continuously listening voice assistants, and vulnerability-prone smart hubs have all been exploited for large-scale attacks and data exfiltration. Zero Trust Architecture (ZTA) [<xref ref-type="bibr" rid="ref-2">2</xref>], therefore, has been widely adopted; its &#x201C;never trust, always verify&#x201D; maxim demands precise, real-time device fingerprints to support least-privilege access control and enable dynamic defence.</p>
<p>Within highly encrypted, heterogeneous home networks, obtaining accurate fingerprints faces two core challenges: (i) Transport Layer Security (TLS) 1.3 [<xref ref-type="bibr" rid="ref-3">3</xref>] and Encrypted Client Hello (ECH) [<xref ref-type="bibr" rid="ref-4">4</xref>] markedly diminish Deep Packet Inspection visibility; (ii) IoT traffic is highly dynamic and diverse, rendering rule-based methods largely ineffective. These macro-level obstacles push researchers toward deep-learning solutions that rely solely on traffic metadata to compensate for diminished visibility and accommodate complex behaviour.</p>
<p>Among deep-learning frameworks for encrypted traffic, Transformers [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>] achieve excellent performance but are hindered in real-time deployment by quadratic time complexity; State Space Models (SSMs) [<xref ref-type="bibr" rid="ref-7">7</xref>] such as Mamba [<xref ref-type="bibr" rid="ref-8">8</xref>] reduce complexity to linear time and, in NetMamba [<xref ref-type="bibr" rid="ref-9">9</xref>], set multiple new benchmarks, yet reveal fresh technical bottlenecks: first, flattening hierarchical traffic produces an architectural mismatch; second, the model remains a &#x201C;black box,&#x201D; lacking intrinsic explainability [<xref ref-type="bibr" rid="ref-10">10</xref>]. Consequently, a lightweight model that both captures hierarchical structure and delivers real-time explanations is urgently needed.</p>
<p>In response, this paper proposes HMX-GAF, a hierarchical Mamba framework that models packet-flow dependencies and leverages selective gating for zero-cost explainability. Experiments on CIC-IoT-2024 confirm significant improvements in both accuracy and interpretability. The main contributions are: (1) an HMX hierarchical architecture capturing intra- and inter-packet dependencies; (2) a zero-overhead GAF explainability mechanism. The remainder of this paper is organised as follows: <xref ref-type="sec" rid="s2">Section 2</xref> surveys related work; <xref ref-type="sec" rid="s3">Section 3</xref> details the method; <xref ref-type="sec" rid="s4">Section 4</xref> reports experiments; <xref ref-type="sec" rid="s5">Section 5</xref> concludes. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> contrasts flattening with our hierarchical approach.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Conceptual comparison of traffic representation: (<bold>a</bold>) conventional flattening approach vs. (<bold>b</bold>) proposed hierarchical HMX approach.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82810-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>The Evolution of IoT Device Fingerprinting</title>
<p>IoT device fingerprinting has evolved through three stages [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Early work leveraged physical-layer characteristics such as clock skew [<xref ref-type="bibr" rid="ref-13">13</xref>] and Radio Frequency (RF) imbalances [<xref ref-type="bibr" rid="ref-14">14</xref>], but these methods are susceptible to environmental conditions and require specialized hardware. The second stage focused on plaintext protocol tokens (e.g., Dynamic Host Configuration Protocol (DHCP), HyperText Transfer Protocol (HTTP) User-Agent), exemplified by IoT Sentinel [<xref ref-type="bibr" rid="ref-15">15</xref>]. However, with TLS now encrypting over 90% of home-network traffic, plaintext-based methods have become largely ineffective [<xref ref-type="bibr" rid="ref-16">16</xref>], driving a shift toward encryption-agnostic behavioral analysis.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deep Learning Methods for Encrypted Traffic</title>
<p>Current approaches analyze sequences of traffic metadata (packet size, direction, inter-arrival times) using deep learning [<xref ref-type="bibr" rid="ref-17">17</xref>]. Convolutional Neural Networks (CNNs) [<xref ref-type="bibr" rid="ref-18">18</xref>] and Long Short-Term Memory networks (LSTMs) [<xref ref-type="bibr" rid="ref-19">19</xref>] capture local temporal patterns [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>] but struggle with long-range dependencies. Transformers improved accuracy via global self-attention, yet their <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> complexity precludes millisecond-level gateway deployment. Pre-training paradigms such as ET-BERT [<xref ref-type="bibr" rid="ref-22">22</xref>] and YaTC [<xref ref-type="bibr" rid="ref-23">23</xref>] further advanced the state of the art through contextualized representation learning.</p>
<p>To overcome the complexity bottleneck, Structured State Space models (S4 [<xref ref-type="bibr" rid="ref-7">7</xref>], H3 [<xref ref-type="bibr" rid="ref-24">24</xref>]) reduce complexity to <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Mamba [<xref ref-type="bibr" rid="ref-8">8</xref>] adds input-dependent selective gating for context-aware state evolution, matching Transformer quality at linear cost. Mamba-2 [<xref ref-type="bibr" rid="ref-25">25</xref>] further establishes a theoretical SSM&#x2013;attention duality with 2&#x2013;8<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> speedups. In the traffic domain, NetMamba [<xref ref-type="bibr" rid="ref-9">9</xref>] first applied Mamba to network traffic classification, raising Macro-F1 from 96.1% to 97.8% while cutting Graphics Processing Unit (GPU) memory by 40%, and ET-Mamba [<xref ref-type="bibr" rid="ref-26">26</xref>] extends the paradigm with ultra-low parameter counts. SSMs have also shown promise for traffic generation [<xref ref-type="bibr" rid="ref-27">27</xref>]. Despite these successes, existing methods flatten traffic into one-dimensional sequences, disregarding the inherent &#x201C;packet-flow&#x201D; hierarchy and lacking intrinsic interpretability [<xref ref-type="bibr" rid="ref-9">9</xref>]. In parallel, lightweight classification architectures have been explored for resource-constrained IoT environments, including Multi-Layer Perceptron (MLP)-based feature extractors with linear classifiers [<xref ref-type="bibr" rid="ref-28">28</xref>] and hybrid ensemble frameworks for industrial IoT intrusion detection [<xref ref-type="bibr" rid="ref-29">29</xref>]. However, these approaches do not address the hierarchical structure of traffic data or provide intrinsic explainability.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Interpretability in Network Security</title>
<p>Interpretability is essential for deploying deep learning in safety-critical security scenarios [<xref ref-type="bibr" rid="ref-30">30</xref>]: Security Operations Center (SOC) analysts need to understand alert root causes [<xref ref-type="bibr" rid="ref-31">31</xref>], and regulations such as the General Data Protection Regulation (GDPR) mandate auditability. Post-hoc methods like LIME [<xref ref-type="bibr" rid="ref-32">32</xref>] and SHAP [<xref ref-type="bibr" rid="ref-33">33</xref>] rely on perturbation-based sampling whose fidelity has been questioned [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>], while attention-based visualizations correlate poorly with model gradients [<xref ref-type="bibr" rid="ref-35">35</xref>]. Our proposed GAF overcomes these limitations by leveraging the Mamba architecture&#x2019;s internal gating signals to generate fine-grained, input-aligned attributions at zero additional cost [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Summary and Research Motivation</title>
<p>In summary, physical-layer and protocol-token methods are ineffective under encryption [<xref ref-type="bibr" rid="ref-3">3</xref>]; Transformers face quadratic complexity bottlenecks [<xref ref-type="bibr" rid="ref-5">5</xref>]; and Mamba-based models lack hierarchical modeling and interpretability [<xref ref-type="bibr" rid="ref-9">9</xref>]. Meanwhile, edge deployment demands lightweight solutions [<xref ref-type="bibr" rid="ref-36">36</xref>]. This paper addresses these gaps with a hierarchical Mamba framework and gated attribution mechanism, targeting two objectives at linear complexity: (1) modeling the &#x201C;packet-flow&#x201D; hierarchy, and (2) providing real-time explanations.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Method</title>
<p>Formally, given a network flow <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>F</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> of <italic>L</italic> packets, the objectives are to learn a classifier <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>f</mml:mi><mml:mo>:</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> that maps flows to one of <italic>C</italic> device classes, and simultaneously produce a per-token attribution map <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow></mml:math></inline-formula> that explains each decision, all within a single forward pass. This section details the HMX-GAF framework (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>), which addresses these objectives via a hierarchical Mamba encoder complemented by an intrinsic attribution mechanism [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The architecture of the HMX-GAF framework. (<bold>a</bold>) Hierarchical representation learning; (<bold>b</bold>) downstream application &#x0026; explainability.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82810-fig-2.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Hierarchical Traffic Representation</title>
<p>Conventional models flatten all packet features into a single sequence, discarding the inherent two-level structure where packets are nested within flows [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>], inflating the sequence length to <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>L</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>K</mml:mi></mml:math></inline-formula> and conflating intra-packet and inter-packet patterns. We instead propose a multi-level representation that explicitly preserves the packet-flow hierarchy.</p>
<p>A network flow is defined as a sequence of <italic>L</italic> packets:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mtext>Flow</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Each packet <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is represented as a sequence of <italic>K</italic> intra-packet feature tokens. Specifically, we extract a fixed set of metadata fields from each packet header&#x2014;including source port, destination port, protocol type, Transmission Control Protocol (TCP) flags, payload length, and inter-arrival time&#x2014;and discretize continuous values into categorical bins. Notably, the inclusion of <italic>protocol type</italic> as an explicit input token allows the shared Packet-Mamba encoder to naturally differentiate TCP from User Datagram Protocol (UDP) semantics: because Mamba&#x2019;s gating parameters (<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">C</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>) are all input-dependent (<xref ref-type="sec" rid="s3_2">Section 3.2</xref>), the encoder implicitly conditions its behavior on the protocol context without requiring separate per-protocol branches. The tokenized representation is:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:mtext>pkt</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mi>K</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Each token <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mtext>tok</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is mapped to a dense vector via a learnable embedding layer <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>E</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:math></inline-formula>:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>E</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>E</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>E</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mi>K</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>To preserve the ordering semantics among tokens within a packet, a learnable positional encoding <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mtext mathvariant="bold">PE</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is added element-wise to the embedded sequence:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="bold">PE</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>The position-augmented sequence <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msubsup><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> serves as the input to the hierarchical encoder.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Hierarchical Mamba (HMX) Encoder</title>
<p>The HMX encoder employs a dual-layer stacked Mamba architecture to explicitly model the hierarchical nature of network traffic. Before describing each layer, we first review the core Selective State Space mechanism that underpins both.</p>
<p><bold>Selective State Space Mechanism.</bold> Each Mamba layer is built upon a state space model that maps input <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> to output <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> through a latent state <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The continuous system <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>y</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">C</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is discretized via zero-order hold with step size <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula>:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow></mml:math></disp-formula>yielding the recurrence <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">C</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>. Crucially, Mamba makes <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mtext mathvariant="bold">C</mml:mtext></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> input-dependent via learned projections from <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-8">8</xref>]: the gate <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> modulates how strongly each token updates the hidden state, enabling the model to emphasize informative tokens while suppressing noise. This selectivity is the theoretical cornerstone upon which our GAF mechanism (<xref ref-type="sec" rid="s3_3">Section 3.3</xref>) is built.</p>
<p><bold>Packet-Mamba (Packet-Level Encoder).</bold> The first layer processes <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> to capture intra-packet token dependencies. The selective scan is applied over the <italic>K</italic> tokens, and the final hidden state serves as the context-aware packet embedding:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>PacketMamba</mml:mi><mml:mspace width="1pt" /><mml:mspace width="negativethinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>The resultant packet embedding <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> encodes rich intra-packet contextual information, such as the semantic association between TCP flags and port numbers. All <italic>L</italic> packets in a flow are processed by a shared Packet-Mamba with tied parameters, ensuring parameter efficiency. We adopt last-state extraction rather than alternative pooling (mean-pooling, attention-pooling) because Mamba&#x2019;s selective scan progressively accumulates all token information into each hidden state via input-dependent gating. Unlike vanilla Recurrent Neural Networks (RNNs), Mamba&#x2019;s <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>-controlled update ensures <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>h</mml:mi><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> retains full intra-packet context, making additional pooling redundant. This design choice is empirically validated in <xref ref-type="sec" rid="s4_4">Section 4.4</xref>.</p>
<p><bold>Flow-Mamba (Flow-Level Encoder).</bold> The second layer, termed Flow-Mamba, takes the sequence of packet embeddings <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>S</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:math></inline-formula> as its input to model inter-packet temporal dependencies:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>S</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mi>L</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>Flow-Mamba computes the final hidden state <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, which serves as the comprehensive embedding for the entire network flow:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>FlowMamba</mml:mi><mml:mspace width="1pt" /><mml:mspace width="negativethinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>S</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>Although both layers share the same mathematical formulation, they operate at different semantic granularities: the former captures field-level correlations within a packet header, while the latter captures behavioral patterns across the temporal evolution of a flow. This hierarchical design mirrors the physical generation process of network traffic, enabling more robust and semantically meaningful representations.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Gated Attribution Fingerprinting (GAF)</title>
<p>To achieve intrinsic explainability without post-hoc computation, GAF repurposes the selective gating signals already computed during the Mamba forward pass.</p>
<p><bold>Theoretical Foundation.</bold> As shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, the step size <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> controls how strongly input <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> updates the hidden state. Under a first-order approximation (<inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2248;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">B</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>), the state innovation scales linearly with <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, making it a natural, gradient-aligned [<xref ref-type="bibr" rid="ref-35">35</xref>] measure of each token&#x2019;s contribution. For token <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>z</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, its attribution score is:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>This hypothesis&#x2014;that <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> serves as a high-fidelity proxy for feature importance&#x2014;is empirically validated in <xref ref-type="sec" rid="s4_5">Section 4.5</xref> through comparison with gradient-based ground truth.</p>
<p><bold>Hierarchical Attribution Pipeline.</bold> Leveraging the dual-layer HMX design, GAF enables a structured top-down attribution process across two semantic levels:<list list-type="bullet">
<list-item>
<p><bold>Flow-level attribution:</bold> The <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> values computed by Flow-Mamba quantify the relative importance of each packet <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> to the final flow embedding <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. These scores are normalized across the flow to produce a packet-importance distribution:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:math></disp-formula></p></list-item>
<list-item><p><bold>Packet-level attribution:</bold> For each identified key packet, <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> values from Packet-Mamba localize the influential intra-packet tokens (e.g., which specific header field drove the decision):
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pkt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:math></disp-formula></p></list-item>
</list></p>
<p><bold>Composite Attribution Score.</bold> The two levels are composed multiplicatively to trace a decision down to a specific token:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi>&#x1D49C;</mml:mi></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pkt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>tok</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>This composite score traces decisions from flow level to individual packet features at <italic>zero</italic> additional inference cost, since all <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> values are already computed during the forward pass.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Joint Training and Anomaly Scoring</title>
<p>The model is trained with a composite loss to jointly optimize classification accuracy and embedding space geometry (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>):<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>triplet</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denotes the standard cross-entropy loss over the device label space.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Impact of triplet loss on embedding space quality. (<bold>a</bold>) A poor embedding space with overlapping clusters vs. (<bold>b</bold>) an ideal embedding space with compact and well-separated clusters.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82810-fig-3.tif"/>
</fig>
<p><bold>Triplet Loss and Hard Mining.</bold> The triplet loss encourages compact, well-separated clusters for each device class:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>triplet</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula>where <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the Euclidean distance and <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0.2</mml:mn></mml:math></inline-formula> is the margin. We employ online semi-hard negative mining to select negatives satisfying <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>m</mml:mi></mml:math></inline-formula>, balancing gradient strength with training stability. Concretely, within each mini-batch of 256 flows (sampled to contain at least 4 flows per device class), we form all valid anchor-positive pairs from same-class flows and search the batch for semi-hard negatives from different classes. When no semi-hard negative exists for a given pair, we fall back to the hardest negative (smallest <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>); pairs lacking any valid negative are skipped. This yields approximately 1200 triplets per batch on average. The coefficient <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula> is determined via grid search.</p>
<p><bold>Anomaly Scoring.</bold> After training, a class centroid <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mi>c</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is computed for each known device class. At inference, a new flow <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>new</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is scored by its distance to the nearest centroid:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>AnomalyScore</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo symmetric="true" maxsize="1.2em" minsize="1.2em">&#x2016;</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>new</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo symmetric="true" maxsize="1.2em" minsize="1.2em">&#x2016;</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>Flows exceeding a threshold <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> (calibrated on the validation set) are flagged as anomalous; GAF (<xref ref-type="sec" rid="s3_3">Section 3.3</xref>) then identifies the responsible features. We choose Euclidean distance for its simplicity and edge efficiency; a comparison with Mahalanobis and kNN scoring is provided in <xref ref-type="sec" rid="s4_4">Section 4.4</xref>. Note that <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>flow</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is the final hidden state of the Flow-Mamba recurrence, which is inherently order-sensitive: two flows with identical per-packet statistics but different temporal orderings produce distinct embeddings, ensuring that temporal dynamics are captured in the anomaly score.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results and Analysis</title>
<p>Our evaluation addresses three research questions: <bold>RQ1</bold> (Effectiveness)&#x2014;does HMX&#x2013;GAF significantly outperform flat-sequence baselines? <bold>RQ2</bold> (Architectural Contribution)&#x2014;what do the hierarchy and triplet loss each contribute? <bold>RQ3</bold> (Explainability)&#x2014;does GAF provide high-fidelity, analyst-meaningful insights?</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Datasets and Splitting Strategy</title>
<p>Our study is primarily based on the CIC-IoT-2024 dataset [<xref ref-type="bibr" rid="ref-19">19</xref>], a large-scale and realistic collection of traffic from 97 distinct smart home devices. This dataset provides a challenging benchmark due to its diversity and scale. For the anomaly detection task, we augment the test set with malicious traffic samples from the Mirai and BashLite portions of the IoT-23 dataset [<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
<p>To ensure the model learns generalizable device fingerprints, we strictly adhere to a device-disjoint splitting strategy: the 97 device identities are randomly shuffled and partitioned into training/validation/test sets with a 70%/15%/15% ratio, guaranteeing mutually exclusive device sets across partitions. The final distribution is summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Dataset split statistics based on the device-disjoint principle.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Set</th>
<th>Purpose</th>
<th># of Devices</th>
<th>Device ID Examples</th>
<th># of Flows</th>
<th>Percentage</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Training</bold></td>
<td>Model Training</td>
<td>68</td>
<td>D01&#x2013;D68</td>
<td>284,521</td>
<td><inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow><mml:mn>70</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>Validation</bold></td>
<td>Hyperparameter Tuning</td>
<td>15</td>
<td>D69&#x2013;D83</td>
<td>61,156</td>
<td><inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow><mml:mn>15</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>Test (Benign)</bold></td>
<td>Final Performance Eval.</td>
<td>14</td>
<td>D84&#x2013;D97</td>
<td>59,782</td>
<td><inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow><mml:mn>15</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>Test (Malicious)</bold></td>
<td>Anomaly Detection Eval.</td>
<td>External dataset</td>
<td>IoT-23 Mirai/BashLite</td>
<td>5210</td>
<td>External AD test only</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Data Preprocessing and Feature Extraction</title>
<p>The raw network traffic data (in <monospace>.pcap</monospace> format) was processed through a multi-stage pipeline to generate feature sequences suitable for the HMX&#x2013;GAF model.
<list list-type="simple">
<list-item>
<label>1.</label>
<p><italic>Flow Generation:</italic> Raw packets were first assembled into bidirectional traffic flows using a 5-tuple identifier (source Internet Protocol (IP), destination IP, source port, destination port, protocol) with a 60-s inactivity timeout.</p></list-item>
<list-item>
<label>2.</label>
<p><italic>Flow Truncation/Padding:</italic> To create fixed-length sequences for batch processing, each flow was truncated to the first <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula> packets. Flows with fewer than 50 packets were padded with zero vectors to ensure uniform length.</p></list-item>
<list-item>
<label>3.</label>
<p><italic>Feature Extraction:</italic> For each packet within a flow, we extracted a sequence of <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> intra-packet features (tokens), including packet length, TCP window size, IP flags, and protocol type. These categorical and numerical features were then mapped to a unified integer vocabulary.</p></list-item>
<list-item>
<label>4.</label>
<p><italic>Normalization:</italic> All numerical features, such as packet length, were normalized to a <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> range using min&#x2013;max scaling based on statistics computed <italic>only</italic> from the training set to prevent data leakage.</p></list-item>
</list></p>
<p>This pipeline results in each network flow being represented as a tensor of shape <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which serves as the input to the model&#x2019;s embedding layer. The same fixed input shape (<inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>) is used during inference timing to avoid kernel selection variance.</p>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Evaluation Metrics and Baselines</title>
<p><bold>Evaluation Metrics:</bold> We employ Accuracy and Macro-F1 for device classification, Area Under the Receiver Operating Characteristic Curve (AUC-ROC) for anomaly detection, parameter count and inference latency on NVIDIA Jetson Xavier NX for efficiency, and intra-class variance <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>/inter-class distance <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>inter</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> for embedding space geometry.</p>
<p><bold>Baseline Models:</bold> We compare against three categories under the <italic>same pipeline</italic>: (1) <bold>Random Forest (RF)</bold> using statistical features; (2) <bold>Transformer-Tiny</bold>, <bold>Transformer (Vanilla)</bold>, <bold>Transformer-XL</bold> [<xref ref-type="bibr" rid="ref-6">6</xref>], and <bold>TabTransformer</bold> [<xref ref-type="bibr" rid="ref-39">39</xref>]; (3) <bold>Mamba (Generic)</bold> [<xref ref-type="bibr" rid="ref-8">8</xref>] and <bold>NetMamba-U</bold> [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p><bold>Implementation Details.</bold> All models use the Adam optimizer. For HMX&#x2013;GAF, <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0.2</mml:mn></mml:math></inline-formula>; early stopping is applied on validation Macro-F1. We report mean <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std over five runs. Efficiency metrics are measured on Jetson Xavier NX (FP32, batch &#x003D; 1, 200 runs after 50 warm-ups).</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Main Performance Evaluation (RQ1)</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> summarizes the performance comparison on the CIC-IoT-2024 test set (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>). HMX&#x2013;GAF achieved a Macro-F1 of 0.992 (&#x002B;0.3% over NetMamba-U) with low standard deviation (<inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:math></inline-formula>), and the highest AUC-ROC of 0.987. Under the throughput setting (<italic>FP32</italic>, batch <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mo>=</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>) on Jetson Xavier NX, it reaches <bold>20.1 ms</bold> per sample, maintaining competitive efficiency. The edge setting (batch &#x003D; 1) is analyzed in <xref ref-type="sec" rid="s4_3">Section 4.3</xref> (<xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Main performance comparison on the CIC-IoT-2024 dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Parameters (M)</th>
<th>Latency (ms)</th>
<th>Accuracy (%)</th>
<th>Macro-F1</th>
<th>AUC-ROC</th>
</tr>
</thead>
<tbody>
<tr>
<td>RF (Baseline)</td>
<td>0.02</td>
<td>11.5</td>
<td><inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mn>82.3</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>0.785</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.008</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mn>0.815</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.007</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>TabTransformer</td>
<td>2.6</td>
<td>26.3</td>
<td><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mn>97.2</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mn>0.972</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mn>0.958</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.006</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Transformer-Tiny</td>
<td>2.1</td>
<td>25.8</td>
<td><inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mn>98.50</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.4</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mn>0.985</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mn>0.965</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Transformer (Vanilla)</td>
<td>3.0</td>
<td>27.8</td>
<td><inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mn>98.3</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mn>0.983</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mn>0.962</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Transformer-XL</td>
<td>3.4</td>
<td>28.6</td>
<td><inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mn>98.6</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mn>0.986</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mn>0.969</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Mamba (Generic)</td>
<td>1.7</td>
<td>18.4</td>
<td><inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mn>98.9</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mn>0.988</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mn>0.973</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>NetMamba-U</td>
<td>1.8</td>
<td>18.9</td>
<td><inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mn>98.95</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mn>0.989</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mn>0.974</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>HMX-GAF (Ours)</bold></td>
<td>2.0</td>
<td>20.1</td>
<td><bold>99.18 </bold><inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mo>&#x00B1;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.2</mml:mn></mml:math></inline-formula></td>
<td><bold>0.992</bold> (&#x002B;0.3%)<sup>&#x2217;</sup></td>
<td><bold>0.987</bold> (&#x002B;1.3%)<sup>&#x2217;</sup></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: Results are the mean <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation of 5 runs. The best results are shown in bold. <sup>&#x2217;</sup>Values in parentheses denote the relative improvement over the NetMamba-U baseline.</p>
</fn>
</table-wrap-foot>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Performance comparison of HMX-GAF and baseline models on the CIC-IoT-2024 dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82810-fig-4.tif"/>
</fig><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Edge inference efficiency on NVIDIA Jetson Xavier NX (B &#x003D; 1, FP32).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Latency (ms, mean <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std)</th>
<th>P95 (ms)</th>
<th>Peak Mem (MB)</th>
<th>Energy (mJ)</th>
</tr>
</thead>
<tbody>
<tr>
<td>RF (Baseline, CPU)</td>
<td>11.5 <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.69</td>
<td>12.6</td>
<td>CPU only, not profiled</td>
<td>58</td>
</tr>
<tr>
<td>TabTransformer</td>
<td>36.8 <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.58</td>
<td>41.0</td>
<td>394</td>
<td>442</td>
</tr>
<tr>
<td>Transformer-Tiny</td>
<td>36.1 <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.53</td>
<td>40.2</td>
<td>359</td>
<td>433</td>
</tr>
<tr>
<td>Transformer (Vanilla)</td>
<td>38.9 <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.72</td>
<td>43.4</td>
<td>423</td>
<td>467</td>
</tr>
<tr>
<td>Transformer-XL</td>
<td>40.0 <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.80</td>
<td>44.6</td>
<td>497</td>
<td>480</td>
</tr>
<tr>
<td>Mamba (Generic)</td>
<td>25.8 <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.81</td>
<td>28.8</td>
<td>280</td>
<td>310</td>
</tr>
<tr>
<td>NetMamba-U</td>
<td>26.5 <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.86</td>
<td>29.6</td>
<td>286</td>
<td>318</td>
</tr>
<tr>
<td><bold>HMX&#x2013;GAF (Ours)</bold></td>
<td><bold>28.1 <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.97</bold></td>
<td><bold>31.3</bold></td>
<td><bold>352</bold></td>
<td><bold>337</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-3fn1" fn-type="other">
<p>Note: Measured on Jetson Xavier NX (FP32, batch &#x003D; 1, fixed input shapes). Latency: mean <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std over 200 runs after 50 warm-ups; P95: 95th-percentile.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Inference Efficiency on Edge (B &#x003D; 1)</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> reports single-sample inference on NVIDIA Jetson Xavier NX (FP32, B &#x003D; 1, 200 runs after 50 warm-ups). SSM models are consistently more efficient than Transformers. HMX&#x2013;GAF lies on the Pareto frontier: while slightly slower than NetMamba-U (28.1 vs. 26.5 ms), it achieves higher Macro-F1 (0.992 vs. 0.989), offering a superior accuracy&#x2013;efficiency trade-off. Tail latency (P95) is well controlled at &#x002B;2&#x2013;4 ms above the mean.</p>

</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Ablation Study (RQ2)</title>
<p>We conduct ablations under the same pipeline (device-disjoint split, identical preprocessing/training schedule); <xref ref-type="table" rid="table-4">Table 4</xref> reports mean <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std over five seeds. Three findings stand out: (i) <italic>Hierarchy is the dominant contributor</italic>&#x2014;removing it lowers Macro-F1 by 1.4 pp and AUC by 1.2 pp; (ii) <italic>Triplet loss acts as a geometry regularizer</italic>&#x2014;it brings a modest Macro-F1 gain but a larger AUC gain; (iii) <italic>GAF provides small yet consistent improvements</italic> without extra training overhead. Replacing the Flow-Mamba backbone with Transformer or Gated Recurrent Unit (GRU) further degrades both metrics, confirming the advantage of SSM-style backbones.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation results on the CIC-IoT-2024 dataset (RQ2).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Variant</th>
<th>Accuracy (%)</th>
<th>Macro-F1</th>
<th><inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> vs. Full (F1)</th>
<th>AUC-ROC</th>
<th><inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> vs. Full (AUC)</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>HMX-GAF (Full)</bold></td>
<td><bold>99.18 </bold><inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mo>&#x00B1;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.20</mml:mn></mml:math></inline-formula></td>
<td><bold>0.992 </bold><inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mo>&#x00B1;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.002</mml:mn></mml:math></inline-formula></td>
<td>&#x2014;</td>
<td><bold>0.987 </bold><inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mo>&#x00B1;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td>&#x2014;</td>
</tr>
<tr>
<td>w/o Hierarchy (Removal of Hierarchy)</td>
<td><inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mn>98.70</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.30</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mn>0.978</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mo>&#x2212;</mml:mo><mml:mn>0.014</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mn>0.975</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mo>&#x2212;</mml:mo><mml:mn>0.012</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>w/o Triplet Loss (Removal of Triplet)</td>
<td><inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mn>98.90</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.30</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mn>0.987</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mo>&#x2212;</mml:mo><mml:mn>0.005</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mn>0.981</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mo>&#x2212;</mml:mo><mml:mn>0.006</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>w/o GAF (Disable Gating/Attribution)</td>
<td><inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mn>99.00</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.20</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mn>0.989</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mo>&#x2212;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mn>0.983</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mo>&#x2212;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Backbone <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> Transformer</td>
<td><inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mn>98.70</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.30</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mn>0.986</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mo>&#x2212;</mml:mo><mml:mn>0.006</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mn>0.979</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mo>&#x2212;</mml:mo><mml:mn>0.008</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Backbone <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> GRU</td>
<td><inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mn>98.50</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.30</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mn>0.981</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mo>&#x2212;</mml:mo><mml:mn>0.011</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mn>0.976</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mo>&#x2212;</mml:mo><mml:mn>0.011</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="6"><italic>Packet Aggregation Strategy</italic></td>
</tr>
<tr>
<td>Aggregation <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> Mean-Pool</td>
<td><inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mn>98.95</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.24</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:mn>0.989</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mo>&#x2212;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mn>0.983</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mo>&#x2212;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Aggregation <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> Attn-Pool</td>
<td><inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mn>99.04</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.22</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mn>0.991</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mo>&#x2212;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mn>0.985</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mo>&#x2212;</mml:mo><mml:mn>0.002</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="6"><italic>Anomaly Scoring Function (classification metrics unchanged)</italic></td>
</tr>
<tr>
<td>Scoring <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> Mahalanobis</td>
<td>&#x2014;</td>
<td>&#x2014;</td>
<td>&#x2014;</td>
<td><inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mn>0.988</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mo>+</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Scoring <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> kNN (<inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>)</td>
<td>&#x2014;</td>
<td>&#x2014;</td>
<td>&#x2014;</td>
<td><inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mn>0.986</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.004</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mo>&#x2212;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-4fn1" fn-type="other">
<p>Note: Results are reported as mean <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation over 5 runs with different seeds. The best results are shown in bold. <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> denotes the absolute change w.r.t. the Full model.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>We further evaluate two additional design choices. First, we compare three packet aggregation strategies: Mean-Pool <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mrow><mml:mtext mathvariant="italic">ep</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>K</mml:mi></mml:mfrac><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:msubsup><mml:mi>h</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></inline-formula>, Attention-Pool (learned weighted sum), and our default Last-State <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>h</mml:mi><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></inline-formula>. As shown in the middle portion of <xref ref-type="table" rid="table-4">Table 4</xref>, Last-State achieves the best Macro-F1 (0.992) with no additional parameters, confirming that the Mamba recurrence effectively accumulates full token context into its final state. Second, we compare anomaly scoring functions: Euclidean distance to the nearest centroid (default), Mahalanobis distance, and <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbor (<inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>) in embedding space. Mahalanobis yields a marginal AUC-ROC improvement (&#x002B;0.001) at 1.12<inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> overhead, while kNN slightly underperforms (<inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mo>&#x2212;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>) at 3.5&#x2013;6<inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> overhead. The Euclidean baseline thus offers the best trade-off for edge deployment.</p>

</sec>
<sec id="s4_5">
<label>4.5</label>
<title>GAF Explainability Analysis (RQ3)</title>
<p>We validate GAF through three complementary analyses.</p>
<p><bold>(1) Qualitative Case Study.</bold> As illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, GAF correctly attributed the classification of a &#x201C;Google Home Mini&#x201D; to Domain Name System (DNS) queries for <monospace>&#x002A;.google.com</monospace> and subsequent QUIC bursts&#x2014;a known behavioral signature. For a Mirai-infected camera, GAF identified repeated TCP connection attempts to port 23 (Telnet) as the root cause of the high anomaly score, aligning with expert domain knowledge.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Ablation study: Performance impact of removing key components from the HMX-GAF model. Note: Statistical significance was evaluated against the full model over five independent runs using a two-sided paired <italic>t</italic>-test; &#x002A;indicates <italic>p</italic> &#x003C; 0.05, and &#x002A;&#x002A;indicates <italic>p</italic> &#x003C; 0.01. Error bars represent standard deviations.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82810-fig-5.tif"/>
</fig>
<p><bold>(2) Perturbation Test.</bold> We injected a known irrelevant token into 10,000 flows; GAF achieved 96.2% top-1 precision in identifying it as least important (<xref ref-type="table" rid="table-5">Table 5</xref>).</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Quantitative validation results for GAF&#x2019;s attribution precision.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Evaluation Metric</th>
<th>Attribution Precision (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Top-1 Precision</td>
<td>96.2</td>
</tr>
<tr>
<td>Top-3 Precision</td>
<td>99.5</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>(3) Gradient Cross-Validation.</bold> We computed Spearman&#x2019;s rank correlation between GAF scores and gradient-based importance (L2-norm of output gradients w.r.t. input embeddings) over 10,000 test samples, obtaining a mean correlation of 0.88 (<inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.06</mml:mn></mml:math></inline-formula>). This confirms that GAF provides a faithful, zero-overhead alternative to expensive post-hoc gradient methods. For reference, generating SHAP explanations for a single flow requires <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow></mml:math></inline-formula>500 forward passes (<inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow></mml:math></inline-formula>14 s on Jetson Xavier NX), while LIME requires <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow></mml:math></inline-formula>1000 perturbation samples (<inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow></mml:math></inline-formula>28 s); GAF produces attributions within the same single forward pass used for classification.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion and Conclusion</title>
<sec id="s5_1">
<label>5.1</label>
<title>Discussion</title>
<p><bold>Generalization.</bold> Our evaluation uses CIC-IoT-2024 (97 device types). Generalizability to enterprise or industrial IoT environments remains to be verified through cross-dataset evaluation.</p>
<p><bold>Scalability.</bold> The hierarchical architecture introduces a moderate memory increase over flat models (352 vs. 286 MB on Jetson Xavier NX). Scaling to ultra-high-throughput (<inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mo>&#x2265;</mml:mo></mml:math></inline-formula>10 Gbps) core networks may require FP16/INT8 quantization. Additionally, the default truncation to <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula> packets covers 87% of flows in the CIC-IoT-2024 dataset; preliminary experiments at <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mi>L</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>20</mml:mn><mml:mo>,</mml:mo><mml:mn>50</mml:mn><mml:mo>,</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mn>200</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> show diminishing returns beyond <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula> (Macro-F1: 0.984, 0.992, 0.993, 0.993), and adaptive-length processing remains a direction for future work.</p>
<p><bold>Explainability Scope.</bold> GAF attributions are tied to the model&#x2019;s internal <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> dynamics (cf. <xref ref-type="sec" rid="s3_2">Section 3.2</xref>) [<xref ref-type="bibr" rid="ref-40">40</xref>]. Integrating external domain knowledge could further enhance actionability for SOC analysts.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Conclusion</title>
<p>This paper proposed HMX-GAF, a Hierarchical Mamba framework with Gated Attribution Fingerprinting for IoT traffic classification. The HMX hierarchical encoder achieves a Macro-F1 of 0.992 on CIC-IoT-2024 via dual-layer Mamba architecture; the GAF mechanism provides zero-cost attributions with 96.2% top-1 precision and 0.88 Spearman correlation with gradient-based importance; and edge evaluations on Jetson Xavier NX confirm 28.1 ms inference latency at batch size 1.</p>
<p>Future work will focus on: (1) cross-domain validation on heterogeneous datasets; (2) model compression via knowledge distillation and quantization-aware training; and (3) integrating GAF attributions with automated threat response systems.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>The authors express their gratitude to Professor Chuanping Hu for his guidance and support throughout the study. His advice on academic exploration and manuscript writing has been a source of inspiration.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Jiangyong Yu; data collection: Jiangyong Yu; analysis and interpretation of results: Jiangyong Yu, Runnan Wang; draft manuscript preparation: Jiangyong Yu; supervision: Jiangyong Yu, Chuanping Hu. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on request from the authors.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>MarketsandMarkets</collab></person-group>. <article-title>Smart home market size, share &#x0026; trends report, 2029 [Internet]</article-title>. <comment>2024 [cited 2026 Mar 15]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.marketsandmarkets.com/Market-Reports/smart-homes-and-assisted-living-advanced-technologies-and-global-market-121.html">https://www.marketsandmarkets.com/Market-Reports/smart-homes-and-assisted-living-advanced-technologies-and-global-market-121.html</ext-link>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rose</surname> <given-names>S</given-names></string-name>, <string-name><surname>Borchert</surname> <given-names>O</given-names></string-name>, <string-name><surname>Mitchell</surname> <given-names>S</given-names></string-name>, <string-name><surname>Connelly</surname> <given-names>S</given-names></string-name></person-group>. <source>Zero trust architecture (NIST special publication 800-207)</source>. <publisher-loc>Gaithersburg, MD, USA</publisher-loc>: <publisher-name>National Institute of Standards and Technology</publisher-name>; <year>2020</year>. doi:<pub-id pub-id-type="doi">10.6028/NIST.SP.800-207</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rescorla</surname> <given-names>E</given-names></string-name></person-group>. <source>The transport layer security (TLS) protocol version 1.3 (RFC 8446)</source>. <publisher-loc>Fremont, CA, USA</publisher-loc>: <publisher-name>Internet Engineering Task Force</publisher-name>; <year>2018</year>. doi:<pub-id pub-id-type="doi">10.17487/RFC8446</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rescorla</surname> <given-names>E</given-names></string-name>, <string-name><surname>Oku</surname> <given-names>K</given-names></string-name>, <string-name><surname>Sullivan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Wood</surname> <given-names>CA</given-names></string-name></person-group>. <source>TLS encrypted client hello (RFC 9849)</source>. <publisher-loc>Fremont, CA, USA</publisher-loc>: <publisher-name>Internet Engineering Task Force</publisher-name>; <year>2026</year>. doi:<pub-id pub-id-type="doi">10.17487/RFC9849</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Carbonell</surname> <given-names>J</given-names></string-name>, <string-name><surname>Le</surname> <given-names>QV</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Transformer-XL: attentive language models beyond a fixed-length context</article-title>. In: <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; 2019 Jul 28&#x2013;Aug 2</conf-name>; <publisher-loc>Florence, Italy</publisher-loc>. p. <fpage>2978</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/P19-1285</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goel</surname> <given-names>K</given-names></string-name>, <string-name><surname>R&#x00E9;</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Efficiently modeling long sequences with structured state spacesl</article-title>. In: <conf-name>Proceedings of the Tenth International Conference on Learning Representations; 2022 Apr 25&#x2013;29; Virtual</conf-name>. p. <fpage>14323</fpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dao</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Mamba: linear-time sequence modeling with selective state spaces</article-title>. <comment>arXiv:2312.00752. 2024</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>NetMamba: efficient network traffic classification via pre-training unidirectional Mamba</article-title>. In: <conf-name>Proceedings of the IEEE 32nd International Conference on Network Protocols (ICNP); 2024 Oct 28&#x2013;31</conf-name>; <publisher-loc>Charleroi, Belgium</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICNP61940.2024.10858569</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Das</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rad</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Opportunities and challenges in explainable artificial intelligence (XAI): a survey</article-title>. <comment>arXiv:2006.11371. 2020</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Safi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dadkhah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shoeleh</surname> <given-names>F</given-names></string-name>, <string-name><surname>Mahdikhani</surname> <given-names>H</given-names></string-name>, <string-name><surname>Molyneaux</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>A survey on IoT profiling, fingerprinting, and identification</article-title>. <source>ACM Trans Internet Things</source>. <year>2022</year>;<volume>3</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3539736</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aouedi</surname> <given-names>O</given-names></string-name>, <string-name><surname>Piamrat</surname> <given-names>K</given-names></string-name>, <string-name><surname>Viho</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A survey of machine learning methods for IoT and smart environments device identification</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>10</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3578935</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kohno</surname> <given-names>T</given-names></string-name>, <string-name><surname>Broido</surname> <given-names>A</given-names></string-name>, <string-name><surname>Claffy</surname> <given-names>KC</given-names></string-name></person-group>. <article-title>Remote physical device fingerprinting</article-title>. <source>IEEE Trans Dependable Secur Comput</source>. <year>2005</year>;<volume>2</volume>(<issue>2</issue>):<fpage>93</fpage>&#x2013;<lpage>108</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TDSC.2005.26</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Radio frequency fingerprint identification for Internet of Things: a survey</article-title>. <source>Secur Saf</source>. <year>2024</year>;<volume>3</volume>:<fpage>2023022</fpage>. doi:<pub-id pub-id-type="doi">10.1051/sands/2023022</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Miettinen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Marchal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hafeez</surname> <given-names>I</given-names></string-name>, <string-name><surname>Asokan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Sadeghi</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Tarkoma</surname> <given-names>S</given-names></string-name></person-group>. <article-title>IoT SENTINEL: automated device-type identification for security enforcement in IoT</article-title>. In: <conf-name>Proceedings of the IEEE 37th International Conference on Distributed Computing Systems (ICDCS); 2017 Jun 5&#x2013;8</conf-name>; <publisher-loc>Atlanta, GA, USA</publisher-loc>. p. <fpage>2177</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICDCS.2017.283</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Okonkwo</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Foo</surname> <given-names>E</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A CNN-based encrypted network traffic classifier</article-title>. In: <conf-name>Proceedings of the 2022 Australasian Computer Science Week; 2022 Feb 14&#x2013;18</conf-name>; <publisher-loc>Brisbane, Australia</publisher-loc>. p. <fpage>74</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3511616.3513107</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rezaei</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Deep learning for encrypted traffic classification: an overview</article-title>. <source>IEEE Commun Mag</source>. <year>2019</year>;<volume>57</volume>(<issue>5</issue>):<fpage>76</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MCOM.2019.1800819</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>End-to-end encrypted traffic classification with one-dimensional convolution neural networks</article-title>. In: <conf-name>Proceedings of the 2017 IEEE International Conference on Intelligence and Security Informatics (ISI); 2017 Jul 22&#x2013;24</conf-name>; <publisher-loc>Beijing, China</publisher-loc>. p. <fpage>43</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ISI.2017.8004872</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rabbani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gui</surname> <given-names>J</given-names></string-name>, <string-name><surname>Nejati</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Kaniyamattam</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mirani</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Device identification and anomaly detection in IoT environments</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>10</issue>):<fpage>13625</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3522863</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aceto</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ciuonzo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Montieri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pescap&#x00E9;</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Mobile encrypted traffic classification using deep learning: experimental evaluation, lessons learned, and challenges</article-title>. <source>IEEE Trans Netw Serv Manag</source>. <year>2019</year>;<volume>16</volume>(<issue>2</issue>):<fpage>445</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNSM.2019.2899085</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lotfollahi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jafari Siavoshani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shirali Hossein Zade</surname> <given-names>R</given-names></string-name>, <string-name><surname>Saberian</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep packet: a novel approach for encrypted traffic classification using deep learning</article-title>. <source>Soft Comput</source>. <year>2020</year>;<volume>24</volume>(<issue>3</issue>):<fpage>1999</fpage>&#x2013;<lpage>2012</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00500-019-04030-2</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>G</given-names></string-name>, <string-name><surname>Gou</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ET-BERT: a contextualized datagram representation with pre-training transformers for encrypted traffic classification</article-title>. In: <conf-name>Proceedings of the ACM Web Conference (WWW); 2022 Apr 25&#x2013;29; Virtual</conf-name>. p. <fpage>633</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3485447.3512217</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gui</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Yet another traffic classifier: a masked autoencoder based traffic transformer with multi-level flow representation</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2023</year>;<volume>37</volume>(<issue>4</issue>):<fpage>5420</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v37i4.25674</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fu</surname> <given-names>DY</given-names></string-name>, <string-name><surname>Dao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Saab</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Thomas</surname> <given-names>AW</given-names></string-name>, <string-name><surname>Rudra</surname> <given-names>A</given-names></string-name>, <string-name><surname>R&#x00E9;</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Hungry hungry hippos: towards language modeling with state space models</article-title>. In: <conf-name>Proceedings of the Eleventh International Conference on Learning Representations; 2023 May 1&#x2013;3</conf-name>; <publisher-loc>Kigali, Rwanda</publisher-loc>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Transformers are SSMs: generalized models and efficient algorithms through structured state space duality</article-title>. <comment>arXiv:2405.21060. 2024</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>ET-Mamba: a Mamba model for encrypted traffic classification</article-title>. <source>Information</source>. <year>2025</year>;<volume>16</volume>(<issue>4</issue>):<fpage>314</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info16040314</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhagoji</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Bronzino</surname> <given-names>F</given-names></string-name>, <string-name><surname>Schmitt</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Feasibility of state space models for network traffic generation</article-title>. In: <conf-name>Proceedings of the 2024 SIGCOMM Workshop on Networks for AI Computing; 2024 Aug 4&#x2013;8</conf-name>; <publisher-loc>Sydney, NSW, Australia</publisher-loc>. p. <fpage>9</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3672198.3673792</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chandroth</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A lightweight MLP-based feature extraction with linear classifier for intrusion detection system in Internet of Things</article-title>. <source>Electronics</source>. <year>2026</year>;<volume>15</volume>(<issue>8</issue>):<fpage>1604</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics15081604</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Govindarajan</surname> <given-names>V</given-names></string-name>, <string-name><surname>Selvam</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ravi</surname> <given-names>V</given-names></string-name>, <string-name><surname>Sowmya</surname> <given-names>V</given-names></string-name>, <string-name><surname>Soman</surname> <given-names>KP</given-names></string-name></person-group>. <article-title>Aegis-5: a hybrid ensemble framework for intrusion detection in Industry 5.0 driven smart manufacturing environment</article-title>. <source>ACM Trans Auton Adapt Syst</source>. <year>2026</year>:<fpage>3787224</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3787224</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rjoub</surname> <given-names>G</given-names></string-name>, <string-name><surname>Bentahar</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wahab</surname> <given-names>OA</given-names></string-name>, <string-name><surname>Mizouni</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cohen</surname> <given-names>R</given-names></string-name>, <string-name><surname>Otrok</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on explainable artificial intelligence for cybersecurity</article-title>. <source>IEEE Trans Netw Serv Manag</source>. <year>2023</year>;<volume>20</volume>(<issue>4</issue>):<fpage>5115</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNSM.2023.3282740</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khani</surname> <given-names>P</given-names></string-name>, <string-name><surname>Moeinaddini</surname> <given-names>E</given-names></string-name>, <string-name><surname>Dehghan Abnavi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Shahraki</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Explainable artificial intelligence for feature selection in network traffic classification: a comparative study</article-title>. <source>Trans Emerg Telecomm Technol</source>. <year>2024</year>;<volume>35</volume>(<issue>4</issue>):<fpage>e4970</fpage>. doi:<pub-id pub-id-type="doi">10.1002/ett.4970</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ribeiro</surname> <given-names>MT</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guestrin</surname> <given-names>C</given-names></string-name></person-group>. <article-title>&#x201C;Why should I trust you?&#x201D;: explaining the predictions of any classifier</article-title>. In: <conf-name>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016 Aug 13&#x2013;17</conf-name>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. p. <fpage>1135</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2939672.2939778</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lundberg</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>SI</given-names></string-name></person-group>. <article-title>A unified approach to interpreting model predictions</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>4765</fpage>&#x2013;<lpage>74</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gaspar</surname> <given-names>D</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>P</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Explainable AI for intrusion detection systems: LIME and SHAP applicability on multi-layer perceptron</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>11</issue>):<fpage>30164</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3368377</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jain</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wallace</surname> <given-names>BC</given-names></string-name></person-group>. <article-title>Attention is not explanation</article-title>. In: <conf-name>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; 2019 Jun 2&#x2013;7</conf-name>; <publisher-loc>Minneapolis, MN, USA</publisher-loc>. p. <fpage>3543</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/N19-1357</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A lightweight intrusion detection method for IoT based on deep learning and dynamic quantization</article-title>. <source>PeerJ Comput Sci</source>. <year>2023</year>;<volume>9</volume>(<issue>19</issue>):<fpage>e1569</fpage>. doi:<pub-id pub-id-type="doi">10.7717/peerj-cs.1569</pub-id>; <pub-id pub-id-type="pmid">37810346</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sivanathan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gharakheili</surname> <given-names>HH</given-names></string-name>, <string-name><surname>Loi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wiber</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vishwanath</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Classifying IoT devices in smart environments using network traffic characteristics</article-title>. <source>IEEE Trans Mob Comput</source>. <year>2019</year>;<volume>18</volume>(<issue>8</issue>):<fpage>1745</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TMC.2018.2866249</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garcia</surname> <given-names>S</given-names></string-name>, <string-name><surname>Parmisano</surname> <given-names>A</given-names></string-name>, <string-name><surname>Erquiaga</surname> <given-names>MJ</given-names></string-name></person-group>. <article-title>IoT-23: a labeled dataset with malicious and benign IoT network traffic (Version 1.0.0) [Internet]</article-title>. <year>2020 [cited 2026 Jan 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/records/4743746">https://zenodo.org/records/4743746</ext-link>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Khetan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cvitkovic</surname> <given-names>M</given-names></string-name>, <string-name><surname>Karnin</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>TabTransformer: tabular data modeling using contextual embeddings</article-title>. <comment>arXiv:2012.06678. 2020</comment>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Manna</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sett</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Reconciling privacy and explainability in high-stakes: a systematic inquiry</article-title>. <source>Trans Mach Learn Res</source>. <comment>arXiv:2412.20798. 2025</comment>.</mixed-citation></ref>
</ref-list>
</back></article>