<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">77084</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.077084</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>HMF-Net: Hierarchical Multi-Feature Network for IIoT Malware Detection</article-title>
<alt-title alt-title-type="left-running-head">HMF-Net: Hierarchical Multi-Feature Network for IIoT Malware Detection</alt-title>
<alt-title alt-title-type="right-running-head">HMF-Net: Hierarchical Multi-Feature Network for IIoT Malware Detection</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Alamri</surname><given-names>Faten S.</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Raza</surname><given-names>Muhammad Amjad</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Mirdad</surname><given-names>Abeer Rashad</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Saleem</surname><given-names>Adil Ali</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Saba</surname><given-names>Tanzila</given-names></name><xref ref-type="aff" rid="aff-4">4</xref><email>tsaba@psu.edu.sa</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Mathematical Sciences, College of Science, Princess Nourah bint Abdulrahman University</institution>, <addr-line>Riyadh</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Institute of Computer Science, Khwaja Fareed University of Engineering and Information Technology, Abu Dhabi Road</institution>, <addr-line>Rahim Yar Khan, Punjab</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Science and Information Technology, University of Lahore, 1-km Defense Road</institution>, <addr-line>Lahore, Punjab</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-4"><label>4</label><institution>Artificial Intelligence &#x0026; Data Analytics Lab, CCIS, Prince Sultan University</institution>, <addr-line>Riyadh</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Tanzila Saba. Email: <email>tsaba@psu.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>9</day><month>4</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>3</issue>
<elocation-id>84</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>12</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>3</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_77084.pdf"></self-uri>
<abstract>
<p>Rapid expansion of Industrial Internet of Things (IIoT) systems has heightened the vulnerability of critical infrastructure to sophisticated malware attacks. Traditional signature-based detection methods are ineffective against evolving threats, and many machine learning models fail to capture temporal behavior, offer interpretability, or operate efficiently in resource-constrained environments. This study proposes HMF-Net, a Hierarchical Multi-Feature Network, for accurate, interpretable, and efficient IIoT malware detection. HMF-Net combines hierarchical VT-Tag embedding (HVTE) to model semantic behavioral information, temporal detection ratio analysis (TDRA) to capture confidence variations for polymorphic malware, and static structural binary features. These features are fused using an adaptive attention mechanism that dynamically prioritizes the most informative modalities during classification. The framework is evaluated on an IIoT malware dataset with 2515 samples from six malware families using five-fold cross-validation. Results show HMF-Net achieves 92.47% accuracy, outperforming Gradient Boosting (90.57%), Random Forest (88.52%), DeepMLP (87.26%), and SimpleMLP (84.34%) with <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>p</mml:mi></mml:math></inline-formula> &#x003C; 0.05. Ablation studies reveal HVTE as the most influential component, while TDRA and adaptive fusion further enhance performance. Attention-weight analysis highlights feature importance, especially for polymorphic behavior. The compact HMF-Net architecture (4.2 MB, 2.1 M parameters) with a 3.5 ms inference time supports real-time deployment in edge environments, balancing precision and recall for security applications.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>IIoT security</kwd>
<kwd>malware detection</kwd>
<kwd>deep learning</kwd>
<kwd>attention mechanism</kwd>
<kwd>hierarchical embedding</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Princess Nourah bint Abdulrahman University Researchers Supporting Project</funding-source>
<award-id>PNURSP2026R346</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Advancements in the Industrial Internet of Things (IIoT) are transforming manufacturing, energy, healthcare, and transportation by enhancing efficiency and enabling real-time monitoring. However, this connectivity exposes industrial control systems to sophisticated cyberattacks. Unlike conventional IT systems, IIoT faces unique security challenges due to real-time constraints, long equipment lifecycles, and limited computing resources, which traditional cybersecurity solutions cannot fully address [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Timely malware detection is critical in IIoT, as interruptions cannot be tolerated. Advanced malware such as Stuxnet, BlackEnergy, and Triton exploit legacy protocols and unpatched hardware, causing significant disruption and even loss of life [<xref ref-type="bibr" rid="ref-4">4</xref>]. Signature-based antivirus tools fail against polymorphic, encrypted, or obfuscated malware; static analysis misses runtime behavior, and dynamic analysis can be evaded by sandboxes [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. Machine learning offers potential but struggles with evolving polymorphic malware, heterogeneous features, interpretability, and IIoT edge constraints [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Current IIoT malware detection approaches, static learning, dynamic behavior analysis, and network modeling, often rely on a single modality, limiting accuracy, interpretability, and edge deployability. Effective solutions must perform across diverse malware families, remain interpretable for analysts, operate efficiently at the edge, capture temporal polymorphic behavior, and handle heterogeneous features.</p>
<p>This work proposes HMF-Net, a Hierarchical Multi-Feature Network for IIoT malware classification, integrating semantic behavioral embeddings, temporal analysis, and structural feature modeling through attention-based multimodal fusion. The main contributions of this work are summarized as follows:
<list list-type="bullet">
<list-item>
<p>A hierarchical vocabulary embedding scheme for 371 behavioral tags that jointly models fine-grained actions and eight semantic categories, improving behavioral representation and interpretability over flat feature encodings.</p></list-item>
<list-item>
<p>A temporal detection ratio analysis module that captures confidence variations across time, enabling effective detection of polymorphic and evolving malware.</p></list-item>
<list-item>
<p>A multi-scale binary signature extraction mechanism that combines global and localized structural features to improve robustness against obfuscation.</p></list-item>
<list-item>
<p>An adaptive attention-based fusion strategy that dynamically weights heterogeneous feature modalities, enhancing classification accuracy and explainability.</p></list-item>
<list-item>
<p>Comprehensive experiments demonstrating statistically significant performance gains (92.47% accuracy, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>) with a compact model size (4.2 MB) and low inference latency (3.5 ms), suitable for real-time IIoT edge deployment.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> presents the related work relevant to this study. <xref ref-type="sec" rid="s3">Section 3</xref> describes the proposed methodology in detail. The experimental results and performance analysis are reported in <xref ref-type="sec" rid="s4">Section 4</xref>. <xref ref-type="sec" rid="s5">Section 5</xref> provides a discussion of the obtained results, and finally, <xref ref-type="sec" rid="s6">Section 6</xref> concludes the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Evolution of malware detection has moved away to signature-based detection to machine learning and later on to deep learning, with each generation being more focused on solving the shortcomings of the earlier generation and introducing challenges of its own. Visualization-based methods proposed by [<xref ref-type="bibr" rid="ref-11">11</xref>] transformed binaries to grayscale images, with a 98% accuracy on 9458 samples. Nevertheless, these approaches do not store the semantic behavioral information, are susceptible to packed malware as well as computationally expensive.</p>
<p>Deep learning (DL) improved feature representation through various architectures. Saxe and Berlin [<xref ref-type="bibr" rid="ref-12">12</xref>] used a four-layer network on PE file headers, achieving 95.4% detection on 400,000 samples, though static analysis cannot capture temporal dynamics while Ref. [<xref ref-type="bibr" rid="ref-13">13</xref>] achieved 91.5% accuracy, but tree-based methods can&#x2019;t model temporal dependencies. Raff et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed Recurrent Neural Networks (RNNs), achieving 93.2% accuracy with high computational cost, and Vinayakumar et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] improved results with hybrid Convolutional Neural Network&#x2013;Recurrent Neural Network (CNN-RNN) frameworks, though without interpretability. Recent advances in autonomous diagnostic systems [<xref ref-type="bibr" rid="ref-16">16</xref>] show that adaptive learning frameworks are effective in complex distributed environments, demonstrating the broad applicability of intelligent anomaly detection for industrial security and reliability.</p>
<p>Domain-specific methods like Azmoodeh et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] used energy consumption for IoT ransomware detection (97.5% accuracy), while Cui et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] applied variational autoencoders (89.7%). Kitsune [<xref ref-type="bibr" rid="ref-19">19</xref>] provided unsupervised IoT intrusion detection, and Demetrio et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] showed that adversarial perturbations could bypass commercial detectors.</p>
<p>Recent IoT/IIoT approaches have addressed deployment constraints. Ta&#x015F;c&#x0131; [<xref ref-type="bibr" rid="ref-21">21</xref>] used 1D-CNN with self-attention, achieving 98.36%&#x2013;99.99% accuracy. Maddali [<xref ref-type="bibr" rid="ref-22">22</xref>] combined (Convolutional Neural Network Next Generation) ConvNeXt and (Deep Convolutional Generative Adversarial Network) DCGAN for 96.8% accuracy in edge IIoT deployment.</p>
<p>Addressing class imbalance, Li et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed multimodal fusion (95.3% on imbalanced data), and Kim et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] implemented a three-tier architecture for 98.93% accuracy in smart factories. Imran et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] used (Synthetic Minority Over-sampling Technique) SMOTE for APT (Advanced Persistent Threat) defenses in (Industrial Control System) ICS (96.5% accuracy).</p>
<p>Recent surveys [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-29">29</xref>] highlight adversarial robustness, cross-domain generalization, and interpretability as critical research challenges. Existing IIoT malware detection approaches often suffer from several limitations, including reliance on single-modality features, loss of semantic behavioral context, limited capability to capture temporal dynamics of polymorphic malware, and lack of interpretability in decision-making. Moreover, many models are computationally heavy and unsuitable for deployment in resource-constrained industrial environments. The proposed HMF-Net mitigates these drawbacks by: (i) integrating multi-modal features, including behavioral tags, temporal detection ratios, and structural statistics, to preserve both semantic and temporal information; (ii) employing hierarchical embeddings and attention-based fusion to enhance interpretability and dynamically weight feature modalities on a per-sample basis; and (iii) providing a lightweight and scalable framework tailored for efficient IIoT malware classification.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>The study used a machine learning pipeline, shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, for systematic malware classification on 2515 samples from six families. After preprocessing and feature engineering, stratified 5-fold cross-validation ensured robust evaluation. The HMF Net model employed hierarchical multimodal fusion to integrate behavioral embeddings, temporal analysis, and structural features. Performance was measured using accuracy, precision, recall, and F1 score, with statistical significance assessed via paired <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>t</mml:mi></mml:math></inline-formula>-tests against baseline methods.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the proposed machine learning pipeline for IIoT malware classification.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset Description</title>
<p>The IIoT malware dataset [<xref ref-type="bibr" rid="ref-30">30</xref>] contains 2515 labeled samples from six malware families: Trojan (850), Adware (620), Backdoor (480), Worm (340), Ransomware (180), and Rootkit (45), as shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>. Each sample includes VirusTotal detection outputs from over seventy engines, temporal detection ratios at three points (T1, T2, T3), file metadata (size, type, Shannon entropy), and a set of 371 fine-grained behavioral tags derived from dynamic malware analysis reports. These tags, extracted from sandbox execution traces and VirusTotal behavioral summaries, capture low-level actions such as system calls (Application Programming Interface) API invocations, file system operations, registry changes, process events, and network communications.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Overview of the IIoT malware dataset: (<bold>a</bold>) distribution of malware families, (<bold>b</bold>) breakdown of feature types, (<bold>c</bold>) hierarchical taxonomy of the included families, and (<bold>d</bold>) temporal distribution of collected samples.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-2.tif"/>
</fig>
<p>To enhance interpretability and reduce feature sparsity, the 371 behavioral tags are grouped into eight high-level semantic categories: file system interaction, network communication, registry activity, process manipulation, memory and code injection, data exfiltration, persistence mechanisms, and evasion techniques, following established malware taxonomies and IIoT attack lifecycles. The hierarchical behavioral embedding module encodes tags individually and then at the category level, capturing both fine-grained actions and higher-level semantics. This approach preserves detailed behavioral cues while generating robust representations of malware families.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2b</xref> illustrates the dataset&#x2019;s feature space composition, with behavioral tags forming the largest portion. Temporal ratios, file metadata, and structural properties provide complementary descriptors. The hierarchical structure of the six malware families is shown in <xref ref-type="fig" rid="fig-2">Fig. 2c</xref>, and the month-by-month acquisition pattern is in <xref ref-type="fig" rid="fig-2">Fig. 2d</xref>.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Preprocessing</title>
<p>Prior to feature extraction, the dataset is preprocessed to ensure consistency and remove artifacts. Approximately 8.3% of missing values in numerical attributes are imputed using the median, while missing categorical values are replaced by the mode. Formally, for a feature <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> in a sample,
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>imputed</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>median</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>num</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mtext>numerical</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>mode</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>cat</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mtext>categorical</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mtext>num</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mtext>cat</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> represent the sets of numerical and categorical features, respectively. Numerical features are standardized to zero mean and unit variance:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>with <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> being the mean and standard deviation of the feature. Outlier instances are removed using the interquartile range (IQR) method:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>IQR</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mtext>&#xA0;is removed if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1.5</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>IQR</mml:mtext></mml:mrow><mml:mrow><mml:mtext>&#xA0;or&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>1.5</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>IQR</mml:mtext></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Stratified five-fold cross-validation ensures that the proportion of malware families is preserved in both training and validation splits.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>HMF-Net</title>
<p>The IIoT malware has multifaceted and heterogeneous characteristics, such as file manipulations, network interactions, registry and process control, and evasion. The conventional single-modality models do not reflect this diversity thus giving a low detection accuracy. HMF-Net fills this by hierarchically learning multi-feature modalities using special modules: Fine-grained and Category-level Behavioral Tag Acquisition module (Hierarchical VT-Tag Embedding&#x2014;HVTE), Temporal Dynamic Detection Signal Analysis (Temporal Detection Ratio Analysis&#x2014;TDRA), and Multi-scale Binary Signature Extraction modules (MBSE). These modules produce complementary embeddings, which are fused with an attention mechanism adaptively to enhance malware classification.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Hierarchical VT-Tag Embedding (HVTE)</title>
<p>Each sample includes up to 371 behavioral tags organized into eight semantic categories. To encode these, HVTE employs a learnable embedding matrix <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mtext mathvariant="bold">E</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>371</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, mapping each tag <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> to a dense vector:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">E</mml:mtext></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>m</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>371</mml:mn></mml:math></inline-formula> is the number of tags present in the sample. The tag embeddings are aggregated using mean pooling to obtain the tag-level feature representation:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Category-level embeddings are computed by averaging the embeddings of the tags within each of the eight semantic categories. This provides a hierarchical representation that reflects both fine-grained behaviors and broader malicious patterns.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Temporal Detection Ratio Analysis (TDRA)</title>
<p>For each sample, detection ratios from VirusTotal engines are collected at three time points, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> corresponding to T1, T2, and T3. TDRA summarizes temporal dynamics into statistical features. The mean detection ratio is calculated as
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>3</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>and the standard deviation captures temporal variance:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mn>3</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:msqrt><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The temporal trend, representing the rate of change between the first and last observation, is given by
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>while the range
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>encodes the extrema in detection over time. These handcrafted features are transformed into a learned representation <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mtext>temp</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> through a feedforward network with ReLU activations and hidden dimension 64, capturing non-linear temporal patterns specific to polymorphic malware in the IIoT dataset.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Multi-Scale Binary Signature Extraction (MBSE)</title>
<p>MBSE characterizes file structural and statistical properties. Global features include the log-transformed file size, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mtext>size</mml:mtext><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and Shannon entropy computed from the byte distribution <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> :
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>H</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>log</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Additionally, multi-scale block statistics are extracted at block sizes <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>l</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>256</mml:mn><mml:mo>,</mml:mo><mml:mn>512</mml:mn><mml:mo>,</mml:mo><mml:mn>1024</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> bytes. For each scale, the mean, standard deviation, and skewness of byte values <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> are computed as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:msqrt><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mrow><mml:mtext>skew</mml:mtext></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> is the number of blocks at scale <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>l</mml:mi></mml:math></inline-formula>. Coarse blocks reveal global structural patterns, whereas finer blocks highlight localized anomalies, such as obfuscated payloads or injected code.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Heterogeneous Feature Fusion</title>
<p>The HVTE, TDRA, MBSE, and other behavioral features are projected to a shared space of dimension <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula>. Attention weights are computed to adaptively fuse modalities:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">w</mml:mtext></mml:mrow><mml:mi>i</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">w</mml:mtext></mml:mrow><mml:mi>j</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is shared, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">w</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is modality-specific, and the softmax ensures normalized weights.</p>
<p>The attention-weighted modality embeddings are concatenated to form a joint representation:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>concat</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>M</mml:mi></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mi>M</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>resulting in a 256-dimensional fused vector.</p>
<p>This representation is then passed through two fully connected layers (256 <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> 128) with ReLU activations and dropout (<inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>) before classification.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Model Training</title>
<p>The model is trained using categorical cross-entropy with L2 regularization:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>N</italic> is the number of training samples, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mn>6</mml:mn></mml:math></inline-formula> is the number of malware families, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the ground truth label, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the predicted probability, and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.0001</mml:mn></mml:math></inline-formula>. Optimization is performed with Adam (<inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>&#x03B7;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.9</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.999</mml:mn></mml:math></inline-formula>), combined with a ReduceLROnPlateau scheduler (factor 0.1 after 5 stagnant validation steps) and early stopping (patience 10). Gradient clipping (<inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>5.0</mml:mn></mml:math></inline-formula>) prevents exploding updates, and Xavier-uniform initialization stabilizes signal flow. Training typically converges in 50&#x2013;60 epochs, evaluated using stratified five-fold cross-validation to maintain family distributions. Model performance is assessed using standard classification metrics, including accuracy, precision, recall, and F1-score, defined as follows:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>F1-score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Algorithm</title>
<p>Algorithm 1 summarizes the HMF-Net pipeline tailored for the IIoT malware dataset. Preprocessing handles missing values, scaling, and outlier removal. HVTE, TDRA, and MBSE extract modality-specific features, which are fused via attention, then mapped through the classifier. Training is performed with cross-entropy loss and L2 regularization using Adam optimization.</p>
<fig id="fig-7">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-7.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results</title>
<p>Evaluation employed stratified 5-fold cross-validation with 80&#x2013;20 train-validation splits maintaining class distribution. Five models were compared: HMF-Net (proposed), Gradient Boosting (100 estimators, learning rate 0.1, max depth 5, subsample 0.8), Random Forest (100 trees, max depth 20, min samples split 5), DeepMLP, and SimpleMLP. Metrics include accuracy, precision, recall, and F1-score computed via macro-averaging giving equal weight per class. Statistical significance employed paired <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>t</mml:mi></mml:math></inline-formula>-tests with <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>. Implementation used PyTorch 2.0.1 on NVIDIA RTX 3090 GPUs with CUDA 11.8. Random seed 42 ensures reproducibility.</p>
<p>As evident from <xref ref-type="table" rid="table-1">Table 1</xref>, HMF-Net achieves 92.47% accuracy, outperforming Gradient Boosting by 1.9 percentage points with statistical significance (<inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.031</mml:mn></mml:math></inline-formula>, paired <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>t</mml:mi></mml:math></inline-formula>-test). The balanced precision-recall trade-off (91.83%&#x2013;92.12%) minimizes both false positives (reducing analyst alert fatigue) and false negatives (maximizing threat detection) critical for security applications. Low standard deviation (<inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>1.84</mml:mn></mml:math></inline-formula>%) indicates stable performance across data partitions. Tree-based ensembles (GB: 90.57%, RF: 88.52%) substantially outperform basic neural networks (SimpleMLP: 84.34%), but HMF-Net&#x2019;s specialized architecture surpasses even sophisticated ensembles through hierarchical feature learning and adaptive fusion.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Performance comparison using stratified 5-fold cross-validation. Values reported as mean <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Accuracy (%)</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F1-Score (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>HMF-Net</td>
<td>92.47 <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.84</td>
<td>91.83 <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.52</td>
<td>92.12 <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.68</td>
<td>91.97 <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.59</td>
</tr>
<tr>
<td>Gradient Boosting</td>
<td>90.57 <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.23</td>
<td>90.71 <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.39</td>
<td>90.57 <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.23</td>
<td>90.46 <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.25</td>
</tr>
<tr>
<td>Random Forest</td>
<td>88.52 <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.83</td>
<td>88.09 <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.67</td>
<td>88.52 <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.83</td>
<td>88.23 <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.74</td>
</tr>
<tr>
<td>DeepMLP</td>
<td>87.26 <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.09</td>
<td>86.01 <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.14</td>
<td>86.58 <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.26</td>
<td>86.06 <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.16</td>
</tr>
<tr>
<td>SimpleMLP</td>
<td>84.34 <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.64</td>
<td>82.93 <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.78</td>
<td>83.66 <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.06</td>
<td>83.04 <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.97</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-1fn1" fn-type="other">
<p>Note: All metrics macro-averaged. <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> indicates standard deviation across 5 folds.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<sec id="s4_1">
<label>4.1</label>
<title>Ablation Study</title>
<p>Ablation studies quantify the contribution of each model component. Removing HVTE leads to the largest performance drop (&#x2212;4.13%), underscoring the importance of hierarchical semantic representation, as shown in <xref ref-type="table" rid="table-2">Table 2</xref> and <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. This decline indicates that learned embeddings capture behavioral patterns more effectively than alternative representations such as bag-of-words or TF-IDF. TDRA yields a 2.32% improvement, with particularly strong benefits for polymorphic malware, where temporal dynamics provide discriminative signals; its effect is more limited for static malware. Eliminating AFM results in a 2.65-point decrease, demonstrating that attention-based fusion surpasses simple concatenation by adaptively weighting sample-specific modalities. The progressive performance reduction (<inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mn>92.47</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>89.82</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>90.15</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>88.34</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>) highlights the complementary roles of the components, each addressing distinct detection challenges.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Ablation study quantifying the contribution of each architectural component. Accuracy reported as mean <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation across 5 folds.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>Accuracy (%)</th>
<th><inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Full HMF-Net</td>
<td>92.47 <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.84</td>
<td>baseline</td>
</tr>
<tr>
<td>w/o Adaptive Fusion</td>
<td>89.82 <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.31</td>
<td>&#x2212;2.65</td>
</tr>
<tr>
<td>w/o TDRA</td>
<td>90.15 <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 2.08</td>
<td>&#x2212;2.32</td>
</tr>
<tr>
<td>w/o HVTE</td>
<td>88.34 <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.95</td>
<td>&#x2212;4.13</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: Ablation results demonstrate that each component contributes positively to overall performance, with hierarchical embedding providing the largest gain.</p>
</fn>
</table-wrap-foot>
</table-wrap><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Ablation study visualizing component contributions with HVTE removal causing largest performance degradation (&#x2212;4.13%), validating its critical role in semantic behavioral representation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-3.tif"/>
</fig>
<p>Training converges as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> in 50&#x2013;60 epochs with smooth loss decay (1.67 <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> 0.68) and minimal overfitting indicated by &#x003C;1% train-validation gap (93.2% train vs. 92.8% validation). Early stopping triggers at average epoch 52, confirming effective regularization through dropout (<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>) and L<sub>2</sub> weight decay (<inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.0001</mml:mn></mml:math></inline-formula>). Learning rate reductions occur at epochs 35 and 55, enabling the optimizer to escape shallow local minima and refine learned representations. Stable convergence across all folds (low variance in curves) confirms robustness to random initialization and data partitioning.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Training dynamics: (<bold>a</bold>) training loss exhibiting smooth exponential decay from 1.67 to 0.68 without oscillations, (<bold>b</bold>) training and validation accuracy curves showing minimal gap (&#x003C;1%) indicating effective regularization and no significant overfitting.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Confusion Matrix Analysis</title>
<p>The confusion matrix in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows most classes with diagonal accuracy above 90%, indicating strong per-class discrimination. Primary misclassifications include: (1) Trojan&#x2013;Backdoor (5.3% bidirectional) due to similar command-and-control and remote access patterns, (2) Adware&#x2013;PUP (4.8%) from overlapping advertising injection and monitoring behaviors, and (3) Rootkit (15.9%) across multiple classes due to limited training samples (45). Per-class F1-scores correlate strongly with sample size (<inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>0.87</mml:mn></mml:math></inline-formula>): Trojan (850, 94.1%), Adware (620, 91.7%), Backdoor (480, 90.6%), Worm (340, 88.7%), Ransomware (180, 90.3%), Rootkit (45, 84.3%).</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Normalized confusion matrix shows high diagonal accuracy (&#x003E;90% for most classes) with primary confusion between semantically similar families: Trojan-Backdoor (5.3%) due to overlapping C&#x0026;C behaviors, and Adware-PUP (4.8%) reflecting similar advertising mechanisms.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-5.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Attention Analysis</title>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> on attention analysis shows that the mean weight is greatest in HVTE (0.38), which confirms semantic behavioral tags are the most informative across families. This confirms the effectiveness of hierarchical embedding and explains the large ablation degradation (&#x2212;4.13%). The second criterion that is the most attentive (0.29) is MBSE, which indicates the significance of file structural features. It is further analyzed that the MBSE attention is strongly associated with entropy (r &#x003D; 0.67) with packed/encrypted samples having higher MBSE weights, which confirms the model is learning to focus on structural features to obfuscated malware.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Attention weight analysis: (<bold>a</bold>) distribution across modalities showing HVTE highest (0.38 mean), (<bold>b</bold>) sample-specific patterns revealing polymorphic malware emphasizes TDRA (0.34), (<bold>c</bold>) HVTE-TDRA negative correlation (r &#x003D; &#x2212;0.42) indicating complementary information, (<bold>d</bold>) class-specific attention heatmap.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77084-fig-6.tif"/>
</fig>
<p>TDRA shows moderate mean attention (<inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mn>0.21</mml:mn></mml:math></inline-formula>) but the highest variance (<inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.14</mml:mn></mml:math></inline-formula>), indicating selective importance. Sample-level analysis reveals that polymorphic malware exhibits significantly elevated TDRA attention (mean <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mo>=</mml:mo><mml:mn>0.34</mml:mn></mml:math></inline-formula> vs. overall <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mn>0.21</mml:mn></mml:math></inline-formula>), validating the benefit of temporal modeling for this threat category. Static malware shows suppressed TDRA attention (mean <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mo>=</mml:mo><mml:mn>0.12</mml:mn></mml:math></inline-formula>), confirming that the model appropriately emphasizes temporal evolution for threats exhibiting detection-ratio dynamics while down-weighting for static threats.</p>
<p>Behavioral features receive the lowest attention (<inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mn>0.12</mml:mn></mml:math></inline-formula>) with low variance (<inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.07</mml:mn></mml:math></inline-formula>), likely due to correlation with HVTE since both capture behavioral characteristics. The negative HVTE&#x2013;TDRA correlation (<inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>0.42</mml:mn></mml:math></inline-formula>) suggests a trade-off in attention allocation, where high semantic tag importance corresponds to low temporal importance and vice versa, indicating complementary information representation. Class-specific analysis shows that Packers emphasize MBSE (entropy indicators), Polymorphic malware emphasizes TDRA (temporal evolution), and Simple Trojans emphasize HVTE (behavioral tags sufficient).</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Comparison with Existing Studies</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> presents a structured comparison between the proposed HMF-Net and representative malware detection approaches discussed in the literature review. The comparison focuses on detection accuracy, architectural design, interpretability mechanisms, temporal modeling capability, edge deployability, and multi-modal feature integration, all of which are critical requirements for practical IIoT malware detection.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparative analysis of representative malware detection approaches.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Approach</th>
<th>Key Advantages</th>
<th>Key Limitations</th>
</tr>
</thead>
<tbody>
<tr>
<td>Visualization CNN [<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>Captures binary structure; No feature engineering; High accuracy (98%)</td>
<td>Loses behavioral semantics; Vulnerable to packing; High computation; Not edge-suitable</td>
</tr>
<tr>
<td>Static PE DNN [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>Low false positives (0.1%); Scalable; Fast inference</td>
<td>Static-only; No temporal modeling; Misses polymorphic behavior</td>
</tr>
<tr>
<td>EMBER Boosting [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Large benchmark; Strong baseline (91.5%); Feature-level interpretability</td>
<td>No temporal modeling; Limited feature learning; Manual engineering</td>
</tr>
<tr>
<td>Raw Byte RNN [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>End-to-end learning; No feature engineering; Temporal modeling</td>
<td>Very high computation; Poor interpretability; Long training time</td>
</tr>
<tr>
<td>Hybrid CNN&#x2013;RNN [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Spatial&#x2013;temporal modeling; Improved accuracy (3%&#x2013;5%)</td>
<td>No attention; Limited interpretability; High architectural complexity</td>
</tr>
<tr>
<td>Energy-Based IoT Detection [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>Hardware-level features; Ransomware-focused; Evasion-resistant</td>
<td>Family-specific; Requires specialized hardware; Limited generalization</td>
</tr>
<tr>
<td>VAE-Based Detection [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Unsupervised; Zero-day detection capability</td>
<td>Lower accuracy (89.7%); No multimodal fusion; Weak interpretability</td>
</tr>
<tr>
<td>Kitsune [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Lightweight; Very low FP (0.18%); Edge-friendly</td>
<td>Anomaly-only; No family classification; Traffic-centric</td>
</tr>
<tr>
<td>Multimodal Fusion [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Adaptive weighting; Handles class imbalance; Multimodal learning</td>
<td>Fusion complexity; Multiple extractors; High training cost</td>
</tr>
<tr>
<td>IMCFN [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Transfer learning; High accuracy (96.42%)</td>
<td>Image conversion overhead; Loss of behavioral context; High computation</td>
</tr>
<tr>
<td>Lightweight IIoT AI [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>Low latency; Small footprint; IIoT-oriented</td>
<td>Reduced feature richness; Lower accuracy ceiling</td>
</tr>
<tr>
<td><bold>HMF-Net (Proposed)</bold></td>
<td>Hierarchical multimodal fusion; Temporal modeling; Attention-based interpretability; Compact (4.2 MB); Edge-deployable</td>
<td>Slightly lower peak accuracy than single-task models; Requires multimodal feature extraction</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>While several existing approaches demonstrate high detection accuracy, most optimize only a subset of these criteria. Image-based CNN models [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>] achieve strong performance but incur high computational overhead and lose behavioral semantics. Sequence-based and hybrid deep learning methods [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>] incorporate temporal modeling but lack explicit interpretability and are difficult to deploy at the edge. Lightweight and edge-oriented solutions [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>] are efficient but are either family-specific or do not support fine-grained classification.</p>
<p>Recent multimodal and transformer-based methods [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>] improve accuracy and robustness but rely on complex architectures that limit real-time deployment in resource-constrained industrial environments.</p>
<p>Conversely, HMF-Net is the only model that integrates hierarchical multimodal fusion with explicit time modeling and attention-based interpretability with a small and edge-deployable framework. This enables the HMF-Net to achieve a moderate trade-off between accuracy, transparency, efficiency, and generalizability that is one of the primary gaps that have been identified in the earlier IIoT malware detection research.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>HMF-Net is better than Gradient Boosting (92.47% and 90.57%, respectively) based on three reasons that were confirmed by the ablation studies and attention analysis. First, HVTE has the hierarchical semantic representation that can represent subtle behavioral patterns, which cannot be reflected in the shallow features, and removal of HVTE leads to the highest performance decrease (4.13%). Second, TDRA captures variability over time, increasing polymorphic variant recognition as these samples are rated higher in terms of attention (mean value 0.34). Third, AFM places sample-specific weights on each feature modality, which is better than the simple concatenation by enhancing accuracy (2.65%) and interpretability by emphasizing the prevailing modality of each prediction. HMF-Net is a more approachable model in the sense that it does not focus on accuracy alone but on the alternative criterion of detectability and practicality of its operation in comparison with more representative state-of-the-art methods with accuracies of 85% to 99% reported [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>The main strengths are that HMF-Net is interpretable in terms of attention distributions, TDRA-enabled temporal awareness, a compact 4.2 MB architecture, which can be used by resource-constrained edge devices (unlike ensemble baselines of 40 MB and larger), and can be inferred in real time at 3.5 ms per sample (285 samples per second), which satisfies IIoT throughput demands. Nonetheless, the model is also prone to the imbalance in classes, with worse results on the rootkit class (F1 &#x003D; 84.3%), than on the Trojan one (F1 &#x003D; 94.1%). This may be enhanced by techniques such as SMOTE oversampling or focal loss. The testing on a single dataset encourages the wider cross-domain test in different industrial systems and operating systems. The adversarial robustness analysis is also absent in the current work, which is why future testing against gradient-based attacks may be considered [<xref ref-type="bibr" rid="ref-20">20</xref>], and the issue of adversarial training may be explored.</p>
<p>Behavioral tags based on VirusTotal are deployed, but this relies on this dependency, which restricts its use in air-gapped or extremely secure environments. The next step to examine offline options, including fixed API-call extraction and lightweight sandboxing, and synchronized local threat intelligence caches to enhance deployability in sensitive IIoT environments will be developed in the future.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This paper introduced HMF-Net, a hierarchical multi-feature fusion framework for IIoT malware detection that jointly models semantic, temporal, and structural characteristics in an interpretable and computationally efficient manner. By integrating hierarchical VT-Tag embedding, temporal detection ratio analysis, and adaptive attention-based fusion, HMF-Net effectively addresses polymorphic and evolving malware threats. Extensive experiments on a six-family IIoT malware dataset demonstrate statistically significant performance gains, achieving 92.47% accuracy over established baselines. Ablation and attention analyses confirm the dominant contribution of hierarchical behavioral embeddings, while temporal modeling and adaptive fusion further enhance detection performance and interpretability. The compact model size and low inference latency make HMF-Net suitable for real-time deployment on resource-constrained industrial edge devices. Despite these advantages, limitations remain regarding computational overhead, reliance on labeled data, and evaluation on a single dataset. Future work will focus on addressing class imbalance, improving generalization across industrial environments, and enhancing robustness against adversarial threats within the modular HMF-Net framework.</p>
</sec>
</body>
<back>
<ack>
<p>The authors want to acknowledge the fund by Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R346). The authors would also like to acknowledge the support of Prince Sultan University, Riyadh, Saudi Arabia, for the APC of this publication.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R346), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Study conception and design were carried out by Faten S. Alamri, Adil Ali Saleem, Abeer Rashad Mirdad, and Tanzila Saba. Experiments were performed by Faten S. Alamri, Adil Ali Saleem, and Tanzila Saba. Visualization was contributed by Muhammad Amjad Raza, Abeer Rashad Mirdad, and Tanzila Saba. Validation was conducted by Adil Ali Saleem, Abeer Rashad Mirdad, Muhammad Amjad Raza, and Tanzila Saba. Analysis and interpretation of results were carried out by Muhammad Amjad Raza, Adil Ali Saleem, Faten S. Alamri, and Abeer Rashad Mirdad. Draft manuscript preparation was completed by Faten S. Alamri, Adil Ali Saleem, and Tanzila Saba. Funding acquisition was provided by Faten S. Alamri and Tanzila Saba. Supervision was performed by Tanzila Saba and Abeer Rashad Mirdad. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The IIoT malware dataset used in this study is publicly available and can be accessed at: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/lobobobo123/myfirst-malwarecleandata">https://www.kaggle.com/datasets/lobobobo123/myfirst-malwarecleandata</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Maglaras</surname> <given-names>L</given-names></string-name>, <string-name><surname>Moschoyiannis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Janicke</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Deep learning for cyber security intrusion detection: approaches, datasets, and comparative study</article-title>. <source>J Inf Secur Appl</source>. <year>2020</year>;<volume>50</volume>(<issue>1</issue>):<fpage>102419</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jisa.2019.102419</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koroniotis</surname> <given-names>N</given-names></string-name>, <string-name><surname>Moustafa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Sitnikova</surname> <given-names>E</given-names></string-name>, <string-name><surname>Turnbull</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Towards the development of realistic botnet dataset in the internet of things for network forensic analytics: Bot-IoT dataset</article-title>. <source>Future Gener Comput Syst</source>. <year>2019</year>;<volume>100</volume>(<issue>7</issue>):<fpage>779</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2019.05.041</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Service function chain orchestration across multiple domains: a full mesh aggregation approach</article-title>. <source>IEEE Trans Netw Serv Manag</source>. <year>2018</year>;<volume>15</volume>(<issue>3</issue>):<fpage>1175</fpage>&#x2013;<lpage>91</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vinayakumar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alazab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Soman</surname> <given-names>KP</given-names></string-name>, <string-name><surname>Poornachandran</surname> <given-names>P</given-names></string-name>, <string-name><surname>Venkatraman</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Robust intelligent malware detection using deep learning</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>46717</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2019.2906934</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Egele</surname> <given-names>M</given-names></string-name>, <string-name><surname>Scholte</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kirda</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kruegel</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A survey on automated dynamic malware-analysis techniques and tools</article-title>. <source>ACM Comput Surv</source>. <year>2012</year>;<volume>44</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2089125.2089126</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Rasheed</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>A blockchain-based multi-agent security framework for e-commerce systems</article-title>. <source>Int J Theor Appl Computat Intell</source>. <year>2025</year>;<volume>2025</volume>:<fpage>227</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.65278/IJTACI.2025.15</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Application of deep learning in malware detection: a review</article-title>. <source>J Big Data</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>99</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-025-01157-y</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bhatti</surname> <given-names>UA</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Enhanced ransomware attacks detection using feature selection, sensitivity analysis, and optimized hybrid model</article-title>. <source>J Big Data</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>245</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-025-01289-1</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Adjeroh</surname> <given-names>D</given-names></string-name>, <string-name><surname>Iyengar</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>A survey on malware detection using data mining techniques</article-title>. <source>ACM Comput Surv</source>. <year>2017</year>;<volume>50</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3073559</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gibert</surname> <given-names>D</given-names></string-name>, <string-name><surname>Mateu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Planes</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The rise of machine learning for detection and classification of malware: research developments, trends and challenges</article-title>. <source>J Netw Comput Appl</source>. <year>2020</year>;<volume>153</volume>(<issue>4</issue>):<fpage>102526</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jnca.2019.102526</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nataraj</surname> <given-names>L</given-names></string-name>, <string-name><surname>Karthikeyan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jacob</surname> <given-names>G</given-names></string-name>, <string-name><surname>Manjunath</surname> <given-names>BS</given-names></string-name></person-group>. <article-title>Malware images: visualization and automatic classification</article-title>. In: <conf-name>Proceedings of the 8th International Symposium on Visualization for Cyber Security; 2011 Jul 20</conf-name>; <publisher-loc>Pittsburgh, PA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Saxe</surname> <given-names>J</given-names></string-name>, <string-name><surname>Berlin</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Deep neural network based malware detection using two dimensional binary program features</article-title>. In: <conf-name>Proceedings of the 2015 10th International Conference on Malicious and Unwanted Software (MALWARE); 2015 Oct 20&#x2013;22</conf-name>; <publisher-loc>Fajardo, PR, USA</publisher-loc>. p. <fpage>11</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Anderson</surname> <given-names>HS</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>P</given-names></string-name></person-group>. <article-title>EMBER: an open dataset for training static PE malware machine learning models</article-title>. <comment>arXiv:1804.04637. 2018</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Raff</surname> <given-names>E</given-names></string-name>, <string-name><surname>Barker</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sylvester</surname> <given-names>J</given-names></string-name>, <string-name><surname>Brandon</surname> <given-names>R</given-names></string-name>, <string-name><surname>Catanzaro</surname> <given-names>B</given-names></string-name>, <string-name><surname>Nicholas</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Malware detection by eating a whole EXE</article-title>. In: <conf-name>Proceedings of the 2018 AAAI Workshop on AI for Cyber Security</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. <year>2018</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vinayakumar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alazab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Soman</surname> <given-names>KP</given-names></string-name>, <string-name><surname>Poornachandran</surname> <given-names>P</given-names></string-name>, <string-name><surname>Al-Nemrat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Venkatraman</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Deep learning approach for intelligent intrusion detection system</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>41525</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2895334</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Adaptively diagnosing system faults in microservice architecture: an autonomous predictive model construction framework</article-title>. <source>Future Gener Comput Syst</source>. <year>2025</year>;<volume>17</volume>(<issue>5</issue>):<fpage>108256</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2025.108256</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Azmoodeh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dehghantanha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Conti</surname> <given-names>M</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>K-KR</given-names></string-name></person-group>. <article-title>Detecting crypto-ransomware in IoT networks based on energy consumption footprint</article-title>. <source>J Ambient Intell Humaniz Comput</source>. <year>2018</year>;<volume>9</volume>(<issue>4</issue>):<fpage>1141</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12652-017-0558-5</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cui</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G-G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Detection of malicious code variants based on deep learning</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2018</year>;<volume>14</volume>(<issue>7</issue>):<fpage>3187</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2018.2822680</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mirsky</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Doitshman</surname> <given-names>T</given-names></string-name>, <string-name><surname>Elovici</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shabtai</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Kitsune: an ensemble of autoencoders for online network intrusion detection</article-title>. In: <conf-name>Proceedings of the Network and Distributed Systems Security (NDSS) Symposium 2018; 2018 Feb 18&#x2013;21</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Demetrio</surname> <given-names>L</given-names></string-name>, <string-name><surname>Biggio</surname> <given-names>B</given-names></string-name>, <string-name><surname>Lagorio</surname> <given-names>G</given-names></string-name>, <string-name><surname>Roli</surname> <given-names>F</given-names></string-name>, <string-name><surname>Armando</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Explaining vulnerabilities of deep learning to adversarial malware binaries</article-title>. <comment>arXiv:1901.03583. 2019</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ta&#x015F;c&#x0131;</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Deep-learning-based approach for IoT attack and malware detection</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>18</issue>):<fpage>8505</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app14188505</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Maddali</surname> <given-names>D</given-names></string-name></person-group>. <article-title>ConvNeXt-EESNN: an effective deep learning based malware detection in edge based IIoT</article-title>. <source>J Intell Fuzzy Syst</source>. <year>2024</year>;<volume>46</volume>(<issue>4</issue>):<fpage>8883</fpage>&#x2013;<lpage>903</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Otaibi</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Imbalanced malware family classification using multimodal fusion and weight self-learning</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2023</year>;<volume>24</volume>(<issue>7</issue>):<fpage>7642</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2022.3208891</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>YJ</given-names></string-name>, <string-name><surname>Park</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>M</given-names></string-name></person-group>. <article-title>IIoT malware detection using edge computing and deep learning for cybersecurity in smart factories</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>15</issue>):<fpage>7679</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app12157679</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Imran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Siddiqui</surname> <given-names>HUR</given-names></string-name>, <string-name><surname>Raza</surname> <given-names>A</given-names></string-name>, <string-name><surname>Raza</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Rustam</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ashraf</surname> <given-names>I</given-names></string-name></person-group>. <article-title>A performance overview of machine learning-based defense strategies for advanced persistent threats in industrial control systems</article-title>. <source>Comput Secur</source>. <year>2023</year>;<volume>134</volume>(<issue>4</issue>):<fpage>103445</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2023.103445</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Darwish</surname> <given-names>R</given-names></string-name>, <string-name><surname>Abdelsalam</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khorsandroo</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Deep learning based XIoT malware analysis: a comprehensive survey, taxonomy, and research challenges</article-title>. <source>J Netw Comput Appl</source>. <year>2025</year>;<volume>242</volume>:<fpage>104258</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jnca.2025.104258</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nankya</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chataut</surname> <given-names>R</given-names></string-name>, <string-name><surname>Akl</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Securing industrial control systems: components, cyber threats, and machine learning-driven defense strategies</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>21</issue>):<fpage>8840</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23218840</pub-id>; <pub-id pub-id-type="pmid">37960539</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Redhu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Choudhary</surname> <given-names>P</given-names></string-name>, <string-name><surname>Srinivasan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kumar Das</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Deep learning-powered malware detection in cyberspace: a contemporary review</article-title>. <source>Front Phys</source>. <year>2024</year>;<volume>12</volume>:<fpage>1349463</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fphy.2024.1349463</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gaber</surname> <given-names>MG</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Janicke</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A survey of recent advances in deep learning models for detecting malware in desktop and mobile platforms</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3638240</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>IIoT malware detection dataset</collab></person-group>. <article-title>Kaggle</article-title>; <year>2024</year> <comment>[cited 2026 Mar 2]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/lobobobo123/myfirst-malwarecleandata">https://www.kaggle.com/datasets/lobobobo123/myfirst-malwarecleandata</ext-link>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A malware detection and extraction method for the related information using the ViT attention mechanism on android operating system</article-title>. <source>Appl Sci</source>. <year>2023</year>;<volume>13</volume>(<issue>11</issue>):<fpage>6839</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app13116839</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vasan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Alazab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wassan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Naeem</surname> <given-names>H</given-names></string-name>, <string-name><surname>Safaei</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>IMCFN: image-based malware classification using fine-tuned convolutional neural network architecture</article-title>. <source>Comput Netw</source>. <year>2020</year>;<volume>171</volume>:<fpage>107138</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comnet.2020.107138</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdelkhalek</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mashaly</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Addressing the class imbalance problem in network intrusion detection systems using data resampling and deep learning</article-title>. <source>J Supercomput</source>. <year>2023</year>;<volume>79</volume>:<fpage>10611</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11227-023-05073-x</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Smmarwar</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>GP</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>AI-empowered malware detection system for industrial internet of things</article-title>. <source>Comput Elect Eng</source>. <year>2023</year>;<volume>108</volume>:<fpage>108731</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2023.108731</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nazim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alam</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Rizvi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mustapha</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Su&#x2019;ud</surname> <given-names>MM</given-names></string-name></person-group>. <article-title>Multimodal malware classification using proposed ensemble deep neural network framework</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>18006</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-96203-3</pub-id>; <pub-id pub-id-type="pmid">40410526</pub-id></mixed-citation></ref>
</ref-list>
</back></article>