<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">78282</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.078282</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Method for Detecting Spatio-Temporal Correlation Anomalies of WSN Nodes Based on Topological Information Enhancement and Time-Frequency Feature Extraction</article-title>
<alt-title alt-title-type="left-running-head">A Method for Detecting Spatio-Temporal Correlation Anomalies of WSN Nodes Based on Topological Information Enhancement and Time-Frequency Feature Extraction</alt-title>
<alt-title alt-title-type="right-running-head">A Method for Detecting Spatio-Temporal Correlation Anomalies of WSN Nodes Based on Topological Information Enhancement and Time-Frequency Feature Extraction</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Ye</surname><given-names>Miao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Ziheng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Jiang</surname><given-names>Qiuxiang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Xue</surname><given-names>Xingsi</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Wenxi</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Ning</surname><given-names>Yu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-7" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zhu</surname><given-names>Cheng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-4">4</xref><email>zhucheng@glmu.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information and Communication, Guilin University of Electronic Technology</institution>, <addr-line>Guilin</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Fujian Provincial Key Laboratory of Big Data Mining and Applications, Fujian University of Technology</institution>, <addr-line>Fuzhou</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>College of Computer and Data Science, Fuzhou University</institution>, <addr-line>Fuzhou</addr-line>, <country>China</country></aff>
<aff id="aff-4"><label>4</label><institution>Information Center, Guilin Medical University</institution>, <addr-line>Guilin</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Cheng Zhu. Email: <email>zhucheng@glmu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>77</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>12</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_78282.pdf"></self-uri>
<abstract>
<p>In recent years, anomaly detection in Wireless Sensor Networks (WSNs) has been widely studied using Graph Neural Networks and Transformer-based methods. However, in multi-node and multi-modal data scenarios, these approaches still face challenges such as insufficient extraction of spatiotemporal correlation features, limited modeling capabilities when relying solely on either time-domain or frequency-domain information, and high computational overhead. To address these issues, this work aims to develop an anomaly detection model that balances detection performance with computational efficiency, enabling effective identification of complex anomaly patterns. Specifically, we propose a time&#x2013;frequency feature extraction method with topological information enhancement, topology-enhanced multi-modal spatio-temporal anomaly detection (TE-MSTAD). Building upon the Receptance Weighted Key Value (RWKV) model with linear complexity, a cross-modal feature extraction module is introduced to strengthen the modeling of multi-modal correlations. Meanwhile, adaptive adjacency matrices are constructed by integrating time&#x2013;frequency features and combining outputs from different Graph Neural Networks, thereby enhancing topological information. Furthermore, a dual-branch structure is designed to jointly model time-domain and frequency-domain features, improving the extraction of complex anomaly characteristics. Experiments on both publicly available datasets and real-world collected data demonstrate that the proposed method achieves F1-scores of 92.52% and 93.28%, respectively, outperforming existing methods in detection performance and generalization capability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Wireless sensor networks</kwd>
<kwd>anomaly detection</kwd>
<kwd>time-frequency domain fusion</kwd>
<kwd>graph neural networks</kwd>
<kwd>information enhancement</kwd>
</kwd-group><funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62161006</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Guangxi Science and Technology Program under Grant</funding-source>
<award-id>FN2504240022</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Innovation Project of GUET Graduate Education</funding-source>
<award-id>2025YCXS078</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Wireless sensor networks (WSNs) are self-organizing networks composed of numerous distributed sensor nodes, typically employing multi-hop routing for data transmission. WSN nodes can sense and transmit environmental physical data such as temperature, humidity, carbon dioxide concentration, and light intensity. Due to their convenient deployment and flexible network topology, WSNs are widely applied in various fields including defense and military [<xref ref-type="bibr" rid="ref-1">1</xref>], industrial environmental monitoring [<xref ref-type="bibr" rid="ref-2">2</xref>], medical monitoring [<xref ref-type="bibr" rid="ref-3">3</xref>], smart agriculture [<xref ref-type="bibr" rid="ref-4">4</xref>], and smart city transportation [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>However, WSN deployments frequently encounter external interference from complex natural environments or mutual interference between indoor nodes [<xref ref-type="bibr" rid="ref-6">6</xref>]. Furthermore, WSN nodes themselves face limitations from internal factors such as insufficient power supply, software programming defects, long-term hardware aging, and unstable signal transmission/reception [<xref ref-type="bibr" rid="ref-7">7</xref>], leading to data collection and transmission distortion and anomalies. Designing efficient and practical WSN anomaly node detection algorithms is crucial for ensuring stable operation and reliable application of WSNs.</p>
<p>Typically, data collected individually for a single physical quantity can be regarded as a single-time-series data point [<xref ref-type="bibr" rid="ref-8">8</xref>]. Data collected simultaneously for multiple physical quantities is considered multiple time-series data, referred to as multimodal time-series data [<xref ref-type="bibr" rid="ref-9">9</xref>]. Such multi-time-series data may consist of different physical quantities collected from the same node or physical quantities collected separately from different nodes. Anomalies in single-time-series data from WSNs encompass three types of temporal correlations: point anomalies, context anomalies, and collective anomalies [<xref ref-type="bibr" rid="ref-10">10</xref>]. A point anomaly occurs when data at a specific time point significantly deviates from the normal data at other time points within that time series. A context anomaly arises when, within the specific contextual scenario of that time series, some data points diverge from the majority of normal data points. Context anomalies are localized and context-dependent; they may be considered normal in other contextual scenarios. For example, a set of temperature readings reaching 30&#x00B0;C collected by a WSN during winter in a subtropical region would be deemed a context anomaly. However, the same set of readings might be considered normal during midday in summer in that region. Collective anomalies occur when individual data points may not be anomalous on their own, but a group of data points collectively exhibits abnormal behavior. Within a single time series, collective anomalies typically manifest as repeated fluctuations across multiple consecutive data points. For instance, if the temperature in a region should gradually change around 22&#x00B0;C over two minutes, but instead fluctuates repeatedly&#x2014;even dozens of times&#x2014;between 20&#x00B0;C and 25&#x00B0;C. Anomalies in multi-temporal data within WSNs refer to spatio-temporal correlation anomalies among these time series. This arises because multi-modal time series data collected by the same sensor node typically exhibit spatio-temporal correlations [<xref ref-type="bibr" rid="ref-8">8</xref>], and modal data collected by different sensor nodes also often display such correlations [<xref ref-type="bibr" rid="ref-10">10</xref>]. For example, when ambient temperature rises, humidity data collected by sensors should decrease, while the voltage of their power supply batteries should slightly increase. This indicates temporal negative correlation between different time-series data collected by the same node. When a fire occurs in the environment, temperature data collected by sensors around the fire center should all increase, indicating spatial positive correlation between different time-series data collected by different nodes. When such temporal or spatial correlations in collected multi-temporal data are disrupted, this phenomenon is termed a correlation anomaly [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Regarding the aforementioned spatio-temporal correlation anomalies in WSNs, researchers have progressively developed a series of anomaly detection methods for WSN correlations. These WSN anomaly detection approaches have evolved from traditional statistical methods to deep learning, driven by application requirements and technological advancements. Early anomaly detection methods included statistical-based approaches and traditional machine learning techniques, such as threshold detection [<xref ref-type="bibr" rid="ref-11">11</xref>] and clustering methods [<xref ref-type="bibr" rid="ref-2">2</xref>]. However, these approaches have shown limitations in handling applications featuring complex topological structures, high-dimensional multimodal features [<xref ref-type="bibr" rid="ref-12">12</xref>], and scenarios involving both long-term dependencies and spatiotemporal correlations [<xref ref-type="bibr" rid="ref-13">13</xref>]. With the rise of deep learning technologies, convolutional neural networks (CNN) [<xref ref-type="bibr" rid="ref-14">14</xref>], recurrent neural networks (RNN) [<xref ref-type="bibr" rid="ref-15">15</xref>], long short-term memory (LSTM) networks [<xref ref-type="bibr" rid="ref-16">16</xref>], and gated recurrent units (GRU) [<xref ref-type="bibr" rid="ref-17">17</xref>] have emerged. These methods can capture nonlinear relationships in data and automatically extract features using deep networks, thereby enhancing the ability to extract high-dimensional temporal features to some extent. Subsequently, the Transformer [<xref ref-type="bibr" rid="ref-18">18</xref>] enhanced global feature extraction and the capture of long-range dependencies. It mitigated gradient explosion and vanishing gradient issues when processing long-range dependent time series data, enabling better recognition of complex multivariate time series anomaly patterns and thus improving detection performance. It has now become a crucial tool in anomaly detection methods. Furthermore, to effectively mine complex spatial correlation features among nodes in WSNs, graph neural networks (GNN) [<xref ref-type="bibr" rid="ref-13">13</xref>] have been progressively integrated into the spatial feature extraction process for WSN-collected data. However, current mainstream deep learning-based WSN time-series anomaly detection methods, such as those based on Transformers and GNNs, still exhibit the following shortcomings:</p>
<p>First, existing WSN anomaly detection methods lack sufficient capability to extract spatio-temporal correlation features across multiple nodes and modalities. Current deep learning-based WSN anomaly detection approaches typically focus solely on detecting anomalies in correlations between different modalities within the same sensor node or within the same modality across different sensor nodes. They fail to address the detection of spatio-temporal correlation anomalies across different modalities and sensor nodes. Second, due to the quadratic growth of computational complexity with sequence length in self-attention mechanisms, Transformers incur significantly increased computational overhead and memory consumption when processing long sequences. This leads to prolonged training times and excessive resource consumption [<xref ref-type="bibr" rid="ref-19">19</xref>]. Furthermore, since WSN-collected data contains not only sensor node attribute features but also spatial topology information between nodes [<xref ref-type="bibr" rid="ref-1">1</xref>], existing GNN-based anomaly detection methods suffer from weak generalization capabilities. This is due to their monolithic model structures and lack of effective topology information augmentation mechanisms, making it difficult to fully capture the complex spatial topology features in WSNs and thereby reducing anomaly detection performance. Finally, based on the uncertainty principle in time-series representation [<xref ref-type="bibr" rid="ref-20">20</xref>] demonstrates that data exhibits significant differences in temporal and frequency domains for different anomaly types. When a certain type of anomaly is difficult to detect in the time domain, detection performance in the frequency domain significantly improves, and <italic>vice versa</italic> [<xref ref-type="bibr" rid="ref-21">21</xref>]. However, most existing WSN anomaly detection methods perform feature analysis and extraction solely on signals in either the time domain or frequency domain, limiting the performance of WSN anomaly detection. For example, traditional time-domain analysis has limited capability in extracting periodic features from data, reducing detection performance for anomalies exhibiting periodic variations. In contrast, the frequency domain can reveal the periodic structure and energy distribution characteristics of signals, enabling more effective identification.</p>
<p>To address the aforementioned issues, this paper proposes a topology-enhanced multi-modal spatio-temporal anomaly detection (TE-MSTAD) method. First, to overcome the limitations in WSN anomaly detection accuracy caused by insufficient single-domain feature extraction due to uncertainty in time series representation, this paper designs a WSN anomaly detection framework based on a spatio-temporal domain fusion reconstruction mechanism. This framework employs the Fourier transform to map raw time series to the frequency domain, decomposing them into phase and amplitude features to construct a frequency domain matrix. Simultaneously, it preprocesses the time series and embeds them to form a time domain matrix. By learning features from the combined spatio-temporal matrix, the method enhances its ability to recognize different types of anomalies. Second, to address the limitations of existing WSN anomaly detection methods&#x2014;inadequate extraction of multi-node spatial correlation features and insufficient detection capability due to reliance on single models&#x2014;this paper introduces an information enhancement mechanism in three aspects. This is achieved by jointly computing the spatial correlation of spatiotemporal feature vectors and employing an ensemble strategy to fuse outputs from different base models. Before feature extraction by the backbone network, time-series data is first transformed into the frequency domain. The spatial correlation between time-domain and frequency-domain variables is calculated using the Spearman correlation coefficient, constructing an adjacency graph structure incorporating spatiotemporal correlation. This provides the model with node distribution association information across both time and frequency dimensions. When extracting spatially correlated features in both temporal and frequency domains, this paper constructs graph neural network submodels incorporating graph convolutional network (GCN) [<xref ref-type="bibr" rid="ref-22">22</xref>], graph attention network (GAT) [<xref ref-type="bibr" rid="ref-23">23</xref>], and predict-then-propagate graph neural network (PPNP) [<xref ref-type="bibr" rid="ref-24">24</xref>], and fuses the outputs of each submodel through an adaptive ensemble strategy. This strategy combines the advantages of GCN&#x2019;s local smoothing aggregation capability, GAT&#x2019;s neighbor-adaptive weighting capability, and PPNP&#x2019;s global information propagation capability. The fused features serve as the final output, enhancing the model&#x2019;s generalization ability. To fully exploit the temporal correlations among multiple modalities in WSN time-series data, this paper designs a Cross-modal Feature Extraction (CFE) module based on the Receptance Weighted Key Value (RWKV) model. By performing cross-modal modeling of the temporal features from different modalities within a single node, the module enhances the model&#x2019;s ability to represent multi-modal correlations. Finally, to address the high computational complexity and memory consumption of Transformers, this paper implements the aforementioned improvements on the linearly complex RWKV model. Compared to Transformers, RWKV introduces temporal and channel mixing mechanisms, replaces global computations of self-attention with recursive state updates, and reduces computational resource overhead while maintaining highly efficient parallel training [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<p>In summary, the main contributions of this work are as follows:<list list-type="order">
<list-item>
<p>To address the issue that existing methods are difficult to fully extract the temporal correlation features of multiple temporal modalities, this paper adds a designed CFE module to the RWKV model. While maintaining the parallel training of RWKV to handle long-distance dependent tasks and reducing the computational complexity, it can fully extract the temporal correlation features between different temporal modalities. It is suitable for anomaly detection in multi-temporal modal scenarios of WSN.</p></list-item>
<list-item>
<p>To address the issue that existing methods are difficult to fully extract the spatial correlation features of multiple nodes, this paper proposes two topological information enhancement strategies. Firstly, this paper proposes a method for calculating the similarity between nodes through time-domain and frequency-domain features, and builds a graph adjacency matrix containing time-frequency domain information. Secondly, this paper proposes a method for enhancing spatial feature extraction based on graph neural network ensemble. By constructing a series of different graph neural network complex models, the spatial features that different models focus on outputting are obtained. Therefore, this design can fully extract the spatial correlation features among multiple nodes and is suitable for anomaly detection in the multi-node scenario of WSN.</p></list-item>
<list-item>
<p>Aiming at the problem that the existing methods only rely on a single time domain or frequency domain for feature extraction, and the anomaly detection performance is limited due to the uncertainty principle of time series representation, this paper designs a dual-branch network TE-MSTAD based on feature extraction in the time and frequency domains. This network can combine information in the time domain and frequency domain for WSN anomaly detection, making up for the limitations of relying on a single time domain or frequency domain, thereby improving the performance of WSN anomaly detection.</p></list-item>
</list></p>
<p>The remainder of this paper is structured as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews relevant research progress in the field of WSN anomaly detection; <xref ref-type="sec" rid="s3">Section 3</xref> formally defines the research problem and briefly introduces the fundamental principles of the RWKV model; <xref ref-type="sec" rid="s4">Section 4</xref> elaborates on the overall framework and constituent modules of the proposed TE-MSTAD model; <xref ref-type="sec" rid="s5">Section 5</xref> demonstrates the performance of the proposed method on publicly available indoor datasets and real-world outdoor datasets, along with comparative analysis against existing approaches; finally, <xref ref-type="sec" rid="s6">Section 6</xref> summarizes the work and outlines future research directions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>In recent years, with the widespread application of WSNs, accurately and efficiently detecting anomalous data has become a key concern for both academia and industry. Extensive research has been conducted by scholars worldwide on anomaly detection for WSN data. This section reviews existing studies, focusing on the development and evolution of WSN anomaly detection methods and frequency domain analysis approaches, providing theoretical foundations and technical references for subsequent work.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Deep-Learning Model for Anomaly Detection in WSNs</title>
<p>Historically, WSN anomaly detection primarily relied on traditional statistical methods and machine learning-based approaches. Statistical methods include Kalman filtering, autoregressive integrated moving average (ARIMA) models, and principal component analysis (PCA), while machine learning-based methods encompass clustering-based anomaly detection and dimension reduction techniques. Ref. [<xref ref-type="bibr" rid="ref-11">11</xref>] proposed a novel online adaptive Kalman filtering method specifically for real-time anomaly detection in WSNs by dynamically adjusting filtering parameters and anomaly detection thresholds in response to real-time data. Ref. [<xref ref-type="bibr" rid="ref-2">2</xref>] proposes a monotonic split-and-conquer scheme for detecting anomalous sensor data by leveraging spatio-temporal correlations between neighboring sensors through principal component analysis. Ref. [<xref ref-type="bibr" rid="ref-26">26</xref>] presents a weighted k-means spectral and hierarchical clustering ensemble scheme for graph anomaly detection, based on weighted Euclidean distance computation and weighting. Traditional statistical methods typically assume data conform to specific distribution models, which may not always hold in practical applications, limiting detection accuracy. Furthermore, these methods often exhibit high computational complexity when handling complex, high-dimensional data, making them ill-suited for real-time detection demands on large-scale datasets [<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>In contrast, deep learning methods can automatically learn complex features and patterns in data, demonstrating greater adaptability and accuracy when handling high-dimensional, nonlinear data. Furthermore, deep learning approaches can capture long-term dependencies within data, making them highly suitable for WSN data with time-series characteristics. Ref. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposes a data-driven anomaly detection method termed Median Filter-Stacked Long Short-Term Memory-Exponentially Weighted Moving Average for anomaly identification. Ref. [<xref ref-type="bibr" rid="ref-29">29</xref>] integrates inductive bias and convolutional operations into Transformers, leveraging multi-layer pyramid structures and multi-level skip connections to extract multi-scale features from data. By incorporating anomaly detection into the feature space, it achieves more accurate industrial anomaly detection and localization results. Ref. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposes a masked network Swin Transformer Unet for anomaly detection. It generates simulated anomalies by applying anomaly simulation and masking strategies to non-anomalous samples, leveraging the Swin Transformer&#x2019;s robust global learning capability to repair masked regions. Building upon the original RWKV model, Ref. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposes a novel detection scheme for harsh environments using an ensemble of autoencoders, Gaussian mixture models, and K-means, focusing on analyzing single-round forwarding rate time series of nodes. These methods demonstrate superior performance compared to traditional and machine learning approaches when handling complex, high-dimensional, long-term sequence data. However, such approaches focus solely on the temporal correlation features of WSN time-series data, making it difficult to extract latent spatial correlation information from graph-structured WSN data. Consequently, they cannot achieve spatial correlation anomaly detection for WSN time-series data.</p>
<p>With the emergence of graph neural networks (GNNs), anomaly detection methods leveraging these models can aggregate node information via adjacency matrices, effectively extracting spatial correlation features between nodes. Ref. [<xref ref-type="bibr" rid="ref-31">31</xref>] proposes a collaborative approach where pattern mining guides GNN algorithms to aggregate local information through connections, thereby capturing global patterns. This method employs a GNN encoder for feature aggregation, while the pattern mining algorithm supervises the GNN training process through a novel loss function. Ref. [<xref ref-type="bibr" rid="ref-32">32</xref>] proposes an interpretable spatio-temporal graph convolutional network. By integrating temporal and event similarity perspectives, IST-GCN leverages both directed and undirected graphs to capture system features, providing temporal and spatial interpretability. Ref. [<xref ref-type="bibr" rid="ref-33">33</xref>] adopts a semi-supervised learning approach relying solely on normal data for effective anomaly pattern detection, selecting the GCN-VAE model. By combining the spatial feature extraction capability of graph convolutional networks with the latent temporal feature modeling of variational autoencoders, this approach effectively detects anomalous signs in data. Ref. [<xref ref-type="bibr" rid="ref-34">34</xref>] proposes an event-aware graph attention network that detects and tracks sensors and their spatial correlations within cyber-physical systems. It graphically analyzes and models relationships between components during labeled time periods, identifying anomalies through the constructed graph model. Ref. [<xref ref-type="bibr" rid="ref-35">35</xref>] presents an anomaly detection scheme based on GAT and Informer. GAT effectively learns sequence features, while Informer excels in long-term sequence prediction. Their combined approach utilizes long-term prediction loss and short-term prediction loss to detect anomalies in multivariate time series. Short-term prediction forecasts the next time point&#x2019;s value, while long-term prediction assists short-term forecasting.</p>
<p>However, these methods primarily focus on extracting spatio-temporal features between multiple nodes and a single temporal modality in WSNs, making it challenging to comprehensively capture the intrinsic correlations among multiple nodes and multiple temporal modalities in complex environments. At the same time, Transformer-based approaches exhibit quadratic time complexity and high memory consumption. Moreover, most existing anomaly detection methods concentrate on the acquisition and analysis of time-domain information, typically relying on statistical characteristics, trend variations, or pattern matching of time-series data to identify anomalies [<xref ref-type="bibr" rid="ref-27">27</xref>] Such approaches may pay insufficient attention to periodic features and latent anomaly patterns embedded in the frequency domain. Consequently, these methods fail to fully exploit frequency-domain characteristics, potentially limiting further improvements in detection performance.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Frequency-Domain Analysis Approach for Anomaly Detection in WSNs</title>
<p>In the field of anomaly detection, frequency domain analysis has gained increasing attention as a crucial complementary approach. By transforming time-series data into the frequency domain, it reveals periodicity, frequency distribution, and the variation patterns of different frequency components. This information plays an indispensable role in understanding data correlation features and improving detection performance [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>Currently, for WSN anomaly detection, existing research has extensively employed various typical frequency-domain analysis methods to reveal hidden periodic or frequency-based anomaly features. Among these, the Fourier Transform (FT), as the most commonly used frequency-domain tool, can rapidly map time-domain signals to the frequency domain, uncovering stable periodic patterns. It is frequently utilized to detect anomalies caused by periodic drift or noise interference. For instance, Ref. [<xref ref-type="bibr" rid="ref-36">36</xref>] addresses sensor electrical signal acquisition in indoor environments. It employs the Fourier Transform to convert sensor-perceived signals from the time domain to the frequency domain, generating spatiotemporal image datasets. Combined with Generative Adversarial Networks (GANs), this approach detects anomalous behaviors within electrical signals. This method not only validates the effectiveness of frequency-domain information in anomaly detection for sensor-perceived signals but also demonstrates application potential in areas such as private space monitoring support and human activity perception. For multi-sensor signals of industrial gears, ref. [<xref ref-type="bibr" rid="ref-37">37</xref>] extracts frequency-domain features using the Fast Fourier Transform (FFT) and combines graph neural networks with adversarial autoencoders to achieve unsupervised anomaly detection. By mining multi-scale features within the frequency domain, this approach significantly enhances anomaly detection accuracy, validating the effectiveness of frequency-domain analysis in industrial multi-sensor scenarios. Ref. [<xref ref-type="bibr" rid="ref-38">38</xref>] proposes a frequency-domain-based anomaly detection method. It models the background in the frequency domain using the Fast Fourier Transform (FFT), detects anomalies through peak features in the amplitude spectrum and Gaussian low-pass filtering, and further suppresses non-anomalous high-frequency details using phase spectrum reconstruction. This significantly enhances background suppression capability and detection accuracy.</p>
<p>Additionally, wavelet transform (WT) performs localized analysis of non-stationary signals in the time-frequency domain through multiscale decomposition, widely applied for detecting sudden anomalies and local pattern changes. Ref. [<xref ref-type="bibr" rid="ref-39">39</xref>] combines continuous wavelet transform with support vector clustering to construct a lightweight unsupervised anomaly detection framework capable of effectively identifying drift anomalies in sensor data. Experimental results on the IBRL dataset validate the method&#x2019;s robustness and high detection accuracy when processing non-stationary sensor signals, further demonstrating the applicability and advantages of wavelet transform in dynamic data stream scenarios. Ref. [<xref ref-type="bibr" rid="ref-40">40</xref>] proposes a variability profile anomaly detection scheme for continuous IoT stream data by integrating discrete wavelet transform with K-means clustering. By rapidly constructing and dynamically updating sensor variance profiles, it significantly enhances the accuracy and real-time performance of both short-term and long-term anomaly detection, validating the application potential of wavelet transform in large-scale online data stream scenarios. Addressing computational resource constraints at edge nodes in industrial IoT environments, ref. [<xref ref-type="bibr" rid="ref-41">41</xref>] proposes a parallel discrete wavelet transform method. It effectively compresses acoustic signals and extracts features while reducing memory consumption and computational overhead. Based on this, a lightweight anomaly detection model is developed, demonstrating its feasibility and real-time advantages in practical industrial equipment monitoring.</p>
<p>In recent years, adaptive decomposition methods such as Empirical Mode Decomposition (EMD) have been introduced to decompose complex nonlinear non-stationary sequences into intrinsic mode functions. This approach separates anomaly features across different frequency bands and enables anomaly point localization. These frequency-domain methods partially overcome the limitations of single-time-domain analysis, which is sensitive to noise and struggles to capture periodic variations, providing an effective complementary approach for WSN anomaly detection. By integrating multivariate empirical mode decomposition with wavelet transform, Ref. [<xref ref-type="bibr" rid="ref-42">42</xref>] addresses feature extraction challenges in vibration response data for environmental monitoring. Through signal decomposition followed by input into multiple deep learning models, it achieves efficient identification of damage types and locations within non-stationary signals, validating the practicality and high accuracy of combined EMD and time-frequency domain features for complex structural anomaly detection. Addressing security monitoring challenges for distributed sparse sensor data in industrial IoT, Ref. [<xref ref-type="bibr" rid="ref-43">43</xref>] designed an adaptive noise and energy entropy feature extraction method based on adaptive fully integrated empirical mode decomposition. Combined with a swarm-optimized classifier, it effectively extracts intrinsic modal features, enabling accurate detection of multiple anomaly perception patterns in industrial production. This validates the robustness and superiority of improved EMD and adaptive selection in sparse data scenarios.</p>
<p>However, the aforementioned anomaly detection methods rely solely on frequency-domain processing and fail to fully leverage time-domain information, which limits further improvements in detection performance. Similarly, these methods focus only on spatio-temporal feature extraction between single nodes and multiple modalities or between single modalities and multiple nodes, without adequately considering spatio-temporal correlations in scenarios involving multiple nodes and multiple modalities, thereby affecting performance in complex anomaly cases.</p>
<p>Based on the analysis of the aforementioned related work, this paper proposes a method called TE-MSTAD to address these limitations in existing WSN anomaly detection approaches. TE-MSTAD jointly analyzes both time-domain and frequency-domain information and employs multiple information-enhancement strategies to more comprehensively capture spatial correlations among multiple nodes. Moreover, based on an improved RWKV model, a cross-modal feature extraction module is incorporated to identify anomalies across different temporal modalities, thereby enhancing overall anomaly detection performance.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Problem Description</title>
<p>This paper employs graph neural networks to model multi-temporal data (also termed multimodal data in literature) collected by WSNs as dynamic attribute graphs, expressing their spatial correlations. Sensor nodes correspond to vertices in the graph, multi-temporal data collected by nodes correspond to attribute matrices, and sensor network topology connections correspond to edges in the attribute graph. Data collected by a wireless sensor network at timestamp <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula> can be modeled as the attribute graph <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula> at time <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Here, the attribute matrix <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>M</mml:mi></mml:math></inline-formula> distinct modalities of data collected by <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>N</mml:mi></mml:math></inline-formula> distinct WSN nodes. The value of an element <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> in the adjacency matrix <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> depends on the connection status between node <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>i</mml:mi></mml:math></inline-formula> and node <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>j</mml:mi></mml:math></inline-formula>. If an edge exists between node <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>i</mml:mi></mml:math></inline-formula> and node <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>j</mml:mi></mml:math></inline-formula>, then <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>; otherwise, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mtext>ij</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>.</p>
<p>Considering a sequence of property graphs <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x003A;</mml:mo><mml:mi>T</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> over <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>T</mml:mi></mml:math></inline-formula> time steps, where the property matrix is <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> and the adjacency matrix is <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, each <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> corresponds to the property graph at timestamp <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula>. The anomaly detection problem for WSN time-series data can then be formulated as a classification task for property graphs. Designing an appropriate neural network architecture and its weight parameters <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, we define the corresponding mapping function <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>f</mml:mi></mml:math></inline-formula> as: <disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi mathvariant="bold-italic">Y</mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">G</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003A;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>W</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>A collection of temporal attribute graphs <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>G</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003A;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>W</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> formed by data collected by WSN within the time window from <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> to <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi mathvariant="bold-italic">Y</mml:mi></mml:math></inline-formula> is a label matrix with shape <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>, corresponding to the output matrix of <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>N</mml:mi></mml:math></inline-formula> nodes across <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>M</mml:mi></mml:math></inline-formula> modalities within a time window of length <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>W</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the label for the <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>j</mml:mi></mml:math></inline-formula>th modality at node <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>i</mml:mi></mml:math></inline-formula> at time <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>t</mml:mi></mml:math></inline-formula>. When <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is 1, the target modality at the target node exhibits an anomaly at this time; when <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is 0, the target modality at the target node behaves normally at this time.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Model Design</title>
<p>The proposed TE-MSTAD anomaly detection model structure is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, comprising a data preprocessing module, a topology information learning module, a time-domain feature extraction module, and a frequency-domain feature extraction module. This model adopts a dual-branch reconstruction architecture to perform anomaly detection on both the time-domain and frequency-domain features of signals. On one hand, the time-domain branch captures temporal characteristics within the time series. On the other hand, the frequency-domain branch utilizes spectral analysis to uncover periodic and latent frequency features. These two branches work collaboratively to achieve more comprehensive anomaly detection performance.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Block diagram of the TE-MSTAD model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-1.tif"/>
</fig>
<p>After preprocessing the WSN-collected dataset, the topology learning module acquires the topological relationships between nodes, deriving their adjacency matrix and constructing an attribute graph. Subsequently, the time series data undergoes wavelet transformation to extract frequency domain information, including amplitude and phase, forming a frequency matrix. The processed data is then input into the model&#x2019;s two branches. In the time-domain branch, the original time series undergoes encoding to extract temporal features and spatio-temporal correlations. These are then reconstructed via a decoder to produce a time-reconstructed sequence. Simultaneously, the frequency matrix enters the frequency-domain branch, where a similar encoding-decoding process extracts its frequency-domain features and completes reconstruction. Ultimately, the model achieves comprehensive anomaly detection across both time and frequency domains through its dual-branch reconstruction mechanism.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Data Preprocessing</title>
<p>The data preprocessing module transforms raw data collected by WSNs into training samples suitable for the model, primarily involving three steps: filtering, downsampling, and normalization.</p>
<p>First, WSN-collected data typically contains substantial high-frequency natural noise originating from environmental interference, hardware components, and other factors. This noise includes large amounts of useless or even misleading information that can interfere with the model&#x2019;s extraction of key features, thereby affecting anomaly detection performance. To effectively suppress high-frequency noise, this module applies a Gaussian filter in the frequency domain. The processing steps are as follows: 
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext>G</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>w</mml:mi><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mover><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the frequency-domain Gaussian filter, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the standard deviation of the Gaussian filter; <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msubsup><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the Fast Fourier Transform (FFT) and Inverse Fourier Transform (IFT), respectively; <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the original time-domain signal, and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the smoothed time-domain signal obtained after filtering.</p>
<p>Subsequently, to reduce redundancy in the raw data, this paper employs a sliding window mechanism for time series downsampling after completing Gaussian filtering noise reduction in the frequency domain. This method segments continuous time series data by setting a sliding step size <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>k</mml:mi></mml:math></inline-formula> and window length <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mrow><mml:mtext>kW</mml:mtext></mml:mrow></mml:math></inline-formula>. Let the current time be <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula>. Starting from this point, <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mrow><mml:mtext>kW</mml:mtext></mml:mrow></mml:math></inline-formula> data points form a time series sample of length <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mrow><mml:mtext>kW</mml:mtext></mml:mrow></mml:math></inline-formula>, denoted as <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>kW</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>. After sliding window subsampling, the data <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is divided into <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>k</mml:mi></mml:math></inline-formula> subsamples. Each subsample can be regarded as an observation sequence for a node: <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Finally, to eliminate differences in dimensionality and numerical ranges between sensor data and enhance model training stability and convergence speed, this paper applies standardization to the data after downsampling. The commonly used Z-Score standardization method is employed. By subtracting the mean and dividing by the standard deviation for each data point, it is transformed into a standard normal distribution with mean 0 and standard deviation 1, thereby achieving a unified numerical scale. The specific procedure is as follows:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the input data after downsampling of the <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>j</mml:mi></mml:math></inline-formula>th mode of the <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>i</mml:mi></mml:math></inline-formula>th node; <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the mean value of <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> over the entire time series; <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the standard deviation of <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>; and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the processed standardized result.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Spatial Correlation Feature Learning Enhancement Module</title>
<p>Following the data preprocessing module, this paper designs and introduces a joint time-frequency domain feature correlation learning enhancement algorithm. This algorithm aims to fully explore the spatial correlations among nodes in wireless sensor networks within the time-frequency domain, thereby constructing a more reasonable graph structure to enhance the model&#x2019;s ability to represent topological information.</p>
<p>Specifically, this module first maps the time-domain observation data <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of sensor nodes to the frequency domain via Fourier transform, yielding the frequency-domain feature matrix <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>Next, to measure spatial correlation between nodes, this paper employs the Spearman correlation coefficient. The average correlation value across three modalities serves as the basis for constructing the adjacency matrix for each node. For any two nodes <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow></mml:math></inline-formula>, their correlation in the <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>j</mml:mi></mml:math></inline-formula>th modality is calculated as follows:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>cov</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula>where <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the Spearman rank correlation coefficient between nodes <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow></mml:math></inline-formula>, in the <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>j</mml:mi></mml:math></inline-formula>th mode; <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the rank transformation applied to the sequence; <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mrow><mml:mtext>cov</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote the covariance and standard deviation operations, respectively.</p>
<p>Subsequently, the correlations within each modality in the time domain and frequency domain are averaged separately. The correlation values from the three modalities are then aggregated to obtain the node&#x2019;s overall correlation score:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the joint correlation score between node <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>a</mml:mi></mml:math></inline-formula> and node <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>b</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>N</mml:mi></mml:math></inline-formula> denotes the number of modalities within a single node; <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the data correlation in the time domain and frequency domain, respectively, for the corresponding node in the <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>n</mml:mi></mml:math></inline-formula>th modality.</p>
<p>Finally, to construct a sparse and discriminative adjacency matrix <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, this paper employs a <italic>Top K</italic> strategy to select the top k most relevant nodes for each node. Corresponding elements are set to 1, while all others are set to 0, thereby building an adaptive graph structure:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the top <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>K</mml:mi></mml:math></inline-formula> nodes most correlated with node <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow></mml:math></inline-formula>; <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the elements in the adjacency matrix <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi mathvariant="bold-italic">A</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Time-Domain Feature Extraction Network Branch</title>
<p>In the temporal feature extraction branch, the model takes the attribute matrix <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and adjacency matrix <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>&#x2014;both processed into batches of size <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>B</mml:mi></mml:math></inline-formula>&#x2014;as inputs. The temporal encoder extracts temporal and spatial correlations from the input data, compressing them into a low-dimensional representation. Subsequently, the temporal decoder decodes and reconstructs the encoded representation, aiming to restore the original input sequence as faithfully as possible. The encoder and decoder work in tandem, enabling the model to capture critical temporal features. Next, we will provide a detailed introduction to the specific structure and functions of the temporal encoder module and the temporal decoder module.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Time-Domain Encoder Module</title>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, within the temporal domain encoder module, the input temporal domain data sequentially passes through the temporal feature extraction component and the spatial feature extraction component. The temporal feature extraction component is responsible for uncovering temporal correlations in node attributes over time while capturing intermodal correlations among different physical quantities. The spatial feature extraction component, meanwhile, extracts spatial correlations within the network topology.</p>
<p>The encoder module employs the RWKV model, as illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The RWKV model is a sequence modeling framework that combines the strengths of RNNs and Transformers, featuring linear time complexity and suitability for long text processing [<xref ref-type="bibr" rid="ref-19">19</xref>]. Composed of stacked residual blocks, it incorporates temporal mixing sub-blocks and channel mixing sub-blocks. By introducing a recursive structure, it fully leverages past information. Capitalizing on this advantage, this paper integrates RWKV into the temporal encoding module of the TE-MSTAD network to more efficiently extract relevant features from multimodal temporal data.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>RWKV model block diagram.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-2.tif"/>
</fig>
<p>The temporal mixing sub-block of the RWKV model primarily handles temporal dependencies in sequence data. It efficiently fuses information from the current time step with states from historical time steps through a recursive update mechanism, enabling feature extraction from long sequences:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>wkv</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>u</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>u</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mrow><mml:mtext>wkv</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the input is the sequence <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msup><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is a learnable vector; <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the linear weight; <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the K-vector and V-vector at the <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>i</mml:mi></mml:math></inline-formula>th time step; <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>u</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a learnable log decay parameter controlling the degree of historical retention; <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mrow><mml:mtext>wkv</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a weighted sum similar to attention but without Q or quadratic matrix multiplication, resulting in low computational cost. <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a learnable linear projection matrix; <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the Sigmoid function controlling the information pass rate; <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> represents element-wise multiplication; <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the tensor output from the temporal mixing module.</p>
<p>The channel mixing subblock in the RWKV model primarily enhances the feature expression capability of sequence representations. It models inter-channel dependencies by applying nonlinear transformations and combining different feature channels of the input sequence, thereby extracting richer, higher-dimensional semantic information:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>channel</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>channel</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represent inputs at the current and previous time steps; <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>channel</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes a learnable vector; <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msubsup><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the new input representation; <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are linear weights; <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msubsup><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msubsup><mml:mi mathvariant="bold-italic">k</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msubsup><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> correspond to vector outputs; and <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> constitutes the final output.</p>
<p>In the TE-MSTAD network model, the time-domain embedding vector <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is first fed into the enhanced RWKV network for processing. To effectively capture the correlations among multiple modalities, this study introduces the CFE based on the temporal mixing component of the RWKV model. In this module, the Cross-modal Feature Extraction (CFE) mechanism employs a structure similar to cross-attention. For the <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>i</mml:mi></mml:math></inline-formula>th modal feature, its own V vector undergoes cross-computations with the K vectors generated by other modalities, producing WKV operators <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mrow><mml:mtext mathvariant="bold">wk</mml:mtext></mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> containing cross-modal information. Subsequently, the R vector of this feature is computed with the operator <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mrow><mml:mtext mathvariant="bold">wk</mml:mtext></mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, yielding an output that fuses information across different modalities:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext mathvariant="bold">wk</mml:mtext></mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>u</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow></mml:msup><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mtext>u</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mtext>o</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">wk</mml:mtext></mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The above formula represents an improvement over <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. Here, the superscript <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>M</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>j</mml:mi></mml:math></inline-formula> denotes the sequence numbers of two distinct modalities. <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> corresponds to the data from modality <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>i</mml:mi></mml:math></inline-formula> within the input tensor <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> to the CFE block. <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the feature dimensions of modality <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>i</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the K-vector obtained by cross-mapping <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> through the learnable matrix <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> corresponding to the remaining modality features. The output <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is derived by averaging the product of the temporal mixed input tensor <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>mix</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> and the mapping matrix <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> corresponding to other distinct modalities. This cross-modal operation enables information fusion across modalities and establishes intermodal correlations. <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mrow><mml:mtext mathvariant="bold">wk</mml:mtext></mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the WKV operator generated using <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msubsup><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cross</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>; <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msub><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the output of the temporal mixing component, which is subsequently fed into the channel mixing part of RWKV and concatenated to produce the output vector <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:msubsup><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>rwkv</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where the subscript denotes the improved RWKV model.</p>
<p>In the design of the CFE module, a simple averaging strategy is adopted to fuse the mapped features from other modalities instead of employing complex attention mechanisms. This design is motivated by several considerations. First, in WSN scenarios, different modalities (e.g., temperature, humidity, and voltage) are typically collected from the same physical environment and exhibit strong statistical consistency; thus, averaging helps capture stable shared cross-modal information. Second, the CFE module is embedded within the RWKV-based temporal mixing framework, where RWKV already provides strong sequential modeling capability, and a lightweight fusion strategy helps maintain a balance between representation capacity and model complexity. Finally, the averaging operation introduces minimal computational overhead and offers better scalability, making it more suitable for resource-constrained WSN deployments. Overall, this design achieves an effective trade-off between performance, robustness, and computational efficiency.</p>
<p>To further extract temporal features and perform dimensionality reduction, this model explored multiple network architectures, ultimately incorporating a Temporal Convolutional Network (TCN) and Multi-Layer Perceptron (MLP) structure. The TCN consists of a series of causal convolutional layers with dilation rates, enabling the capture of temporal dependencies across different scales. The MLP facilitates dimension reduction during encoding, feeding low-dimensional embedding vectors to the decoder. Its primary computation can be represented as:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TCN</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Re</mml:mtext></mml:mrow><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>MLP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TCN</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where, in TCN, <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msubsup><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the output of the dilated convolution operation at time step <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mi>t</mml:mi></mml:math></inline-formula> in layer <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>l</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>i</mml:mi></mml:math></inline-formula>th input to layer <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> with dilation stride <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi>d</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mi>d</mml:mi></mml:math></inline-formula> is the dilation factor; <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi>k</mml:mi></mml:math></inline-formula> is the convolution kernel size; <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:msubsup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the weight matrix for the <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mi>i</mml:mi></mml:math></inline-formula>th convolution kernel in layer <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mi>l</mml:mi></mml:math></inline-formula>; and <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is the bias term. In the MLP, <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>MLP</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>C</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the final output tensors of the TCN and MLP networks, respectively; <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denote the weight matrices for the two fully connected layers; <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> correspond to the ReLU activation function; <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represent the bias terms.</p>
<p>After completing feature extraction and dimensionality reduction for temporal features, the attribute matrix output <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>MLP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>MLP</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> by the MLP serves as the node feature input. This is combined with the adjacency matrix <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> generated by the topology information learning module and jointly fed into the spatial feature extraction module. This module integrates two types of graph neural network submodels: GCN and GAT. GCN performs mean aggregation of neighbor node features using a static adjacency matrix, offering high modeling efficiency and strong global structural modeling capabilities. GAT, on the other hand, introduces an attention mechanism that assigns different weights to distinct adjacent nodes, enabling more refined feature representation. Combining GCN and GAT as an information enhancement approach balances global structural stability with local feature adaptability, enhancing the model&#x2019;s ability to capture complex spatial dependencies.</p>
<p>Specifically, GCN achieves local updates to node representations through weighted aggregation of neighboring node information, following this fundamental propagation rule:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>GCN</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">D</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:msup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:msup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">D</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:msup><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gcn</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mrow><mml:mover><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow></mml:math></inline-formula> denotes the adjacency matrix with self-loops, <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">D</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> represents the corresponding degree matrix, and <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gcn</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the learnable weight matrix.</p>
<p>In contrast, GAT introduces an attention mechanism that adaptively learns the importance of different neighboring nodes to the current node. Its basic operation is:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gat</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gat</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gat</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>gat</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the weight matrix, <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow></mml:math></inline-formula> represents the attention weight vector, <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the normalized attention coefficient, <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mo>&#x2225;</mml:mo></mml:math></inline-formula> indicates vector concatenation, and <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mrow><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the set of neighbors for node <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mi>i</mml:mi></mml:math></inline-formula>.</p>
<p>Subsequently, this paper performs adaptive integration of the outputs from GCN and GAT, ultimately expressed as:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>spatial</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>GAT</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>GCN</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where the fusion coefficient <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is a trainable parameter. It is adaptively adjusted through model training to determine the contribution ratio of the two submodels to the final representation.</p>
<p>Finally, the fused spatial feature matrix <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>spatial</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> serves as the output of the temporal encoder module, enabling subsequent decoding reconstruction and anomaly detection.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Temporal Decoder Module</title>
<p>In the network architecture designed herein, the temporal decoder sequentially incorporates a Dense Layer, Batch Normalization, a Dropout Layer, and a Leaky ReLU activation function. These modules not only perform dimensionality-increasing decoding on the feature matrix but also enhance training stability and mitigate overfitting risks, thereby accomplishing feature reconstruction during the decoding phase. First, the Dense Layer performs a linear mapping on the temporal-encoded features to achieve dimensionality expansion. Subsequently, Batch Normalization is introduced to mitigate internal covariate shifts, accelerate model convergence, and enhance stability during mini-batch training. To improve generalization and prevent overfitting, a Dropout Layer is then incorporated into the decoder. During each training iteration, the dropout layer randomly masks the activation values of some neurons with probability <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mi>p</mml:mi></mml:math></inline-formula>. Finally, Leaky ReLU is selected as the nonlinear activation function to avoid the &#x201C;neuron death&#x201D; phenomenon that may occur in the negative range with standard ReLU. The specific operations described above can be expressed as:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>WH</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>spatial</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>BN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>z</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:msqrt><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mo>&#x03B5;</mml:mo></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LeakyReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the input vector <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>spatial</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> from the temporal encoder. <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:math></inline-formula> denote learnable weights and bias terms of the fully connected layer. <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> denote the mean and standard deviation of the current mini-batch data. <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mi>&#x03B3;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> denote the learnable scaling and offset parameters in batch normalization. <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mo>&#x03B5;</mml:mo></mml:math></inline-formula> is a small constant to prevent division by zero. <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow></mml:math></inline-formula> is the dropout mask sampled from a Bernoulli distribution, satisfying <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mtext>Bernoulli</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> denotes element-wise multiplication. <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> denotes Slope coefficient for the negative region in Leaky ReLU. <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:msubsup><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the feature vector finally output by the time-domain decoder.</p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Frequency Domain Feature Extraction Module</title>
<p>In WSN anomaly detection, frequency-domain information often reveals features that temporal analysis cannot fully capture. To further illustrate the importance of frequency-domain information, a specific example is analyzed below. For instance, <xref ref-type="fig" rid="fig-3">Fig. 3</xref> depicts the differences between point anomalies, collective anomalies, and contextual anomalies in the time-frequency domain. In <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the blue curve represents normal data, while the red curve indicates abnormal data. Comparing <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>,<xref ref-type="fig" rid="fig-3">c</xref> shows that collective anomalies exhibit more pronounced changes in the frequency domain than in the time domain, making frequency-domain detection more suitable. Conversely, comparing <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>,<xref ref-type="fig" rid="fig-3">d</xref> demonstrates that time-domain analysis is more advantageous for identifying contextual anomalies.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Comparison chart of normal data and abnormal data in the time-frequency domain: (<bold>a</bold>) Normal data graph; (<bold>b</bold>) point abnormal data graph; (<bold>c</bold>) collective anomaly data graph; (<bold>d</bold>) context abnormal data graph.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-3.tif"/>
</fig>
<p>In the feature extraction module of the frequency domain, the input data <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> is first mapped to the frequency domain space via Fast Fourier Transform (FFT), yielding the frequency domain tensor <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. The time-frequency transformation process is as follows:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mi>T</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>i</mml:mi><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mi>T</mml:mi></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the imaginary part, <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:msup><mml:mi>i</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>; <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the time series corresponding to the mode <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> of node <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> in <inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, with the superscript <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>; similarly, <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> corresponds to the frequency domain sequence mapped from <inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>. However, since <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is a complex sequence, it cannot be directly used for training neural network models. To address this issue, the frequency domain information is typically decomposed into amplitude and phase components for separate representation and storage. Specifically, for the frequency domain sequence <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, the amplitude can be understood as the Euclidean distance of this complex number from the origin in the complex plane, while the phase represents the angle formed with the positive real axis in the complex plane. Assuming <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mi>m</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mi>n</mml:mi></mml:math></inline-formula> denote the real and imaginary parts respectively, the amplitude <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mi>a</mml:mi></mml:math></inline-formula> and phase <inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mi>p</mml:mi></mml:math></inline-formula> are calculated as follows:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi>arctan</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>By calculating the magnitude and phase of each complex element in the frequency domain sequence <inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, the original complex sequence can be transformed into two real-valued sequences suitable for model training: the magnitude sequence <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the phase sequence <inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:msub><mml:mi mathvariant="bold-italic">p</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Similarly, for each input sample set <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, after undergoing the time-frequency domain transformation, the magnitude matrix <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and the phase matrix <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:mi mathvariant="bold-italic">P</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> can be obtained. These two matrices are concatenated and divided into batches to form the frequency-domain matrix <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which serves as input to the frequency-domain encoding network. This enables the encoder to effectively extract sample features based on frequency-domain information.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Frequency Domain Encoder Modulesz</title>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, similarly to the temporal domain encoder module, the input temporal domain data sequentially passes through the temporal feature extraction and spatial feature extraction components. The temporal embedding vector <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">R</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> first traverses a basic RWKV network layer to capture long-range temporal dependencies within the frequency domain. The RWKV model effectively mitigates the impact of sequence length on computational complexity through gating mechanisms and linear attention structures, thereby enhancing its modeling capability for long-term sequence dependencies. Subsequently, to extract intermodal correlations, the output feature <inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is further fed into an enhanced RWKV model. This module incorporates a cross-modal feature extraction (CFE) block. After thorough extraction of temporal dependency features, the output feature <inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is input into a Multi-Layer Perceptron (MLP) network for dimensionality reduction and nonlinear feature mapping. As the detailed process has been described earlier, it is simply represented as:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>RWKV</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>W</mml:mi><mml:mi>K</mml:mi><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the superscript denotes the reduced dimension after dimension reduction <inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>; <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> represents the output tensor after processing through each network; and <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mi>R</mml:mi><mml:mi>W</mml:mi><mml:mi>K</mml:mi><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> denotes the improved RWKV model.</p>
<p>After completing the temporal feature extraction and dimensionality reduction of the frequency-domain embedding vectors, the frequency-domain attribute matrix <inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:msup><mml:mi mathvariant="bold-italic">H</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> output by the MLP serves as the node feature representation. Combined with the adjacency matrix <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> generated by the topology learning module, these are jointly input into the frequency-domain spatial feature extraction module to further learn the spatial dependencies between nodes in the frequency domain. To fully exploit both local and global spatial structural information, this paper introduces two complementary graph neural network submodels in the frequency-domain spatial feature extraction component: the propagatable approximation Personalized Propagation of Neural Predictions (PPNP) and the Graph Attention Network (GAT). Each focuses on learning spatial features at distinct levels, and through subsequent adaptive weighted fusion, they jointly generate frequency-domain feature representations enriched with spatial information.</p>
<p>First, the PPNP subnetwork employs the Personalized PageRank (PPR) propagation mechanism to smoothly transmit node information across the entire graph. This balances the need for both local neighborhood information and broader neighborhood insights, enabling the model to capture more stable spatial structural dependencies even in sparse or noisy networks. Its core propagation process can be formalized as:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>PPNP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mover><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>PPNP</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mtext>AD</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is an adjustable propagation coefficient, <inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is the symmetric normalized adjacency matrix, <inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mi mathvariant="bold-italic">D</mml:mi></mml:math></inline-formula> is the node degree matrix, and <inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mrow><mml:mtext>PPNP</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the learnable weight matrix. This propagation process effectively suppresses local noise interference in node representations and enhances global consistency of features.</p>
<p>In contrast, the GAT subnetwork introduces an adaptive attention mechanism to assign different weights to distinct neighboring nodes, focusing on modeling dependencies among local neighbors. Since the detailed process has been described earlier, it is simply represented as:<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mi mathvariant="bold-italic">H</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">H</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Finally, the spatial features extracted by the PPNP and GAT subnetworks undergo adaptive weight fusion in the channel dimension. This approach not only effectively captures the complex spatial relationships between nodes in the frequency domain but also enhances the accuracy and generalization capability of anomaly detection in complex network environments. represents the adaptive learning fusion weight, with the process expressed as:<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msubsup><mml:mi mathvariant="bold-italic">H</mml:mi><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>spatial</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>PPNP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>GAT</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Frequency Domain Decoder Module</title>
<p>In the network architecture designed herein, the frequency-domain decoder section first introduces an MLP to perform nonlinear mapping and dimension recovery on the results extracted from the frequency-domain spatial features. This process primarily integrates multimodal information captured by nodes at the frequency-domain spatial level through a combination of multi-layer linear transformations and activation functions, providing a more expressive high-dimensional feature representation for subsequent temporal dependency reconstruction of sequences. Simultaneously, the introduction of the MLP helps mitigate overfitting risks caused by information redundancy, enhancing the stability and generalization capability of the reconstruction stage.</p>
<p>Following the initial feature dimensionality expansion by the MLP, the frequency-domain decoder further incorporates an RWKV module to leverage its strengths in modeling long-range temporal dependencies. RWKV enhances the decoder&#x2019;s reconstruction quality while preserving the periodicity and phase characteristics of the frequency-domain information. Finally, to effectively map the RWKV output tensor back to the feature space consistent with the original input, the decoder applies another MLP module for output refinement. Overall, the frequency-domain decoder employs an &#x201C;MLP-RWKV-MLP&#x201D; structural combination to ensure thorough restoration and reconstruction of frequency-domain features during decoding, providing an accurate and reliable reconstruction foundation for anomaly detection results.</p>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Anomaly Score</title>
<p>The anomaly detection method proposed in this study employs reconstruction error as the anomaly discrimination criterion, adhering to the design philosophy of autoencoder-like frameworks. During training, the model aims to minimize the discrepancy between reconstructed outputs and observed input data, leveraging this discrepancy to quantify whether anomalies exist in the input data. Considering the reconstructed outputs generated in both the time domain and frequency domain after the input data passes through the network model, they are denoted respectively as:</p>
<p>The reconstructed outputs generated in the time domain and frequency domain after the input data passes through the network model are denoted as <inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, respectively, with their corresponding original inputs being <inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. Additionally, the frequency-domain reconstruction result <inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> can be restored to time-domain form <inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> via inverse Fourier transform <inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, enabling comparison with <inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:msub><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. Therefore, to comprehensively consider these three reconstruction errors, three loss functions are defined as follows:<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>W</mml:mi><mml:mi>N</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>W</mml:mi><mml:mi>N</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>W</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>W</mml:mi><mml:mi>N</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>To fully leverage both time-domain and frequency-domain information while adaptively weighting the influence of different loss terms, this paper introduces an attention-based weighting mechanism. Each loss term is assigned distinct weight coefficients via a softmax operation. Let the weight coefficients be denoted as <inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, satisfying <inline-formula id="ieqn-224"><mml:math id="mml-ieqn-224"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. The final total loss function and anomaly score are then expressed as:<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msub><mml:mrow><mml:mtext>score</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>A higher anomaly score indicates greater deviation between the input data and the model prediction at that time point, making it more likely to be classified as an anomaly. By setting a threshold <inline-formula id="ieqn-225"><mml:math id="mml-ieqn-225"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> based on the final loss entropy value during model training, we can determine whether a specific time point is anomalous. When <inline-formula id="ieqn-226"><mml:math id="mml-ieqn-226"><mml:msub><mml:mrow><mml:mtext>score</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> exceeds the threshold <inline-formula id="ieqn-227"><mml:math id="mml-ieqn-227"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, the data is classified as anomalous and labeled as 1; otherwise, it is classified as non-anomalous and labeled as 0.
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Analysis</title>
<sec id="s5_1">
<label>5.1</label>
<title>Dataset and Parameter Settings</title>
<p>To comprehensively evaluate the performance and adaptability of the anomaly detection model proposed in this paper, experimental validation was conducted on both public datasets and real-world collected datasets. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the public dataset consists of WSN data collected and compiled by Intel Berkeley Research Lab through deploying multiple wireless sensor nodes (Intel Berkeley Research Lab datasets, IBRL). This dataset comprises 54 Mica2Dot sensor nodes, each collecting environmental data such as temperature, humidity, light intensity, and voltage every 31 s. As illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, the real-world dataset was generated by constructing a WSN data collection system suitable for outdoor scenarios based on the LoRa communication protocol. This system consists of 14 distributed sensor nodes and one aggregation node. All sensor nodes were deployed in an open outdoor area, arranged sequentially along a perimeter wall at a uniform height of 35 cm above ground level. The aggregation node centrally receives data from all sensors and transmits it via WiFi to a cloud-based data center for remote data transmission and management. Each sensor node collects temperature, humidity, and voltage parameters every 30 s. The resulting dataset, named LoRA-OSD, covers the time period from 26 November 2024, to 07 January 2025, documenting environmental monitoring information during this phase.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Spatial position distribution map of sensor nodes in the IBRL dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Deployment of sensor nodes in the WSN abnormal node detection system: (<bold>a</bold>) Node distribution map; (<bold>b</bold>) physical diagram of the sensor node.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-5.tif"/>
</fig>
<p>The experimental platform configuration used in this study is as follows: the processor is an Intel<sup>&#x00AE;</sup> Xeon<sup>&#x00AE;</sup> Gold5218CPU@2.30 GHz, the graphics processing unit is an NVIDIA GeForce RTX 3090, and the operating system is Ubuntu 18.04.2 LTS. The experimental environment was developed using Python 3.6.9, employing the PyTorch 2.0.0 deep learning framework and integrating CUDA 11.7 for GPU-accelerated computation. During subsequent debugging experiments, the learning rate was set to lr &#x003D; 0.0002, training epochs to epoch &#x003D; 120, sliding window size to W &#x003D; 200, sliding stride to L &#x003D; 150, and the Adam optimizer was employed.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Evaluation Metrics</title>
<p>To comprehensively evaluate the performance of the proposed anomaly detection model, this paper selects Precision, Recall, and F1-score as the primary evaluation metrics. Let TP (True Positives) denote the number of samples correctly identified as anomalies, FP (False Positives) denote the number of normal samples incorrectly identified as anomalies, FN (False Negatives) denote the number of anomalies not detected, and TN (True Negatives) denote the number of samples correctly identified as normal. The confusion matrix is then represented as shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Classification confusion matrix.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th></th>
<th>Predicted Anomalies</th>
<th>Predicted Normal</th>
</tr>
</thead>
<tbody>
<tr>
<td>True Anomalies</td>
<td>TP</td>
<td>FN</td>
</tr>
<tr>
<td>True Normal</td>
<td>FP</td>
<td>TN</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Precision measures the proportion of samples correctly classified as anomalies (True Positives, TP) among all samples predicted as anomalies (True Positives &#x002B; False Positives, TP &#x002B; FP). A high precision indicates fewer false positives and higher predictive reliability. Its formula is:<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
<p>Recall measures the proportion of actual anomalies (TP &#x002B; FN) that are correctly identified by the model. A high recall indicates fewer false negatives and stronger detection capability. Its formula is:<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
<p>The F1-score is the harmonic mean of precision and recall, serving as a balanced, comprehensive evaluation metric between these two indicators. The F1-score unifies precision and recall into a single metric. Unlike simple arithmetic averaging, F1 employs harmonic averaging. This ensures that when either precision or recall is too low, the F1-score drops significantly. This prevents an excessively high value in one dimension from masking a low value in the other. Consequently, the F1-score provides a more comprehensive reflection of the model&#x2019;s overall performance in practical detection tasks. Its formula is: 
<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Ablation Experiments</title>
<p>To evaluate the role and performance of each module in the proposed network model, corresponding ablation experiments were designed and conducted. In these experiments, multiple ablation schemes were established, removing or modifying specific key modules within the network model to assess their actual impact on overall performance. The following ablation schemes were designed:</p>
<p>Scheme 1: Remove the CE module that extracts intermodal correlations, reverting to the original RWVK model.</p>
<p>Scheme 2: Modify the topology information enhancement method to use only GAT networks in both the time and frequency domains.</p>
<p>Scheme 3: Modify the topology learning enhancement method to compute correlations only in the temporal domain to obtain the adjacency matrix.</p>
<p>Scheme 4: Simultaneously modify both the topology information enhancement method and the topology learning enhancement method. Use only the GAT network on both the time-frequency domain branch and the time domain branch, and employ the adjacency matrix computed solely in the time domain.</p>
<p>Scheme 5: Remove the frequency domain branch after attribute matrix input, perform feature extraction solely in the time domain, and complete both reconstruction and anomaly detection.</p>
<p>Scheme 6: Remove the time-domain branch after attribute matrix input, perform feature extraction solely in the frequency domain, and complete both the reconstruction and anomaly detection tasks.</p>
<p>Scheme 7: Use the complete model designed in this paper without any removal or modification.</p>
<p>The experimental results for different ablation schemes are shown in <xref ref-type="table" rid="table-2">Table 2</xref>. Since scheme 1 removed the CFE module, the RWKV model in the experiment could not effectively mine potential correlation features between multimodal data and lost its ability to extract multimodal correlations. Therefore, compared to the baseline model (Scheme 7), the model in this scheme overlooked a significant number of intermodal correlation anomalies during detection, resulting in decreases of 4.51%, 15.74%, and 10.28% in precision, recall, and F1 score, respectively. In scheme 2, both the temporal and spatial domains of the model employ a single GAT for spatial feature extraction, failing to fully leverage the strengths and complementarity of different submodels. Compared to the baseline model, the F1 score decreased by 5.34%, while precision and recall dropped by 0.99% and 9.56%, respectively. This indicates that a single graph neural network architecture has limitations in extracting topological features and fails to adequately detect anomalies between nodes. Scheme 3 restricted the construction of the adjacency matrix to time-domain information only, neglecting hidden features in the frequency domain. The F1 score decreased by approximately 2.71%, indicating that phase, frequency, and other information in the frequency domain can significantly enhance the node association patterns captured by the adjacency matrix, thereby improving the quality of spatial structure learning. Scheme 4 simultaneously modifies both information enhancement methods of the model: it removes the ensemble approach for multiple submodels in the topological information enhancement and simplifies the adjacency matrix construction to rely solely on temporal correlations. Experimental results show that compared to the baseline model, the F1 score decreased by 6.71%. This indicates a noticeable decline in overall model performance, validating the synergistic role of both topological information enhancement approaches in improving anomaly detection capabilities. Schemes 5 and 6, which removed the frequency domain or time domain branch respectively, exhibited significant deficiencies in information capture, resulting in F1 score decreases of 8.11% and 10.25%. This demonstrates that a single-branch structure struggles to comprehensively capture the latent time-frequency feature information within the data. As the complete model, scheme 7 integrates multimodal modeling, topological information enhancement, graph structure learning, and dual-branch time-frequency feature extraction, demonstrating optimal performance.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Ablation experiments protocols and results.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Scheme</th>
<th>Precision</th>
<th>Recall</th>
<th>F1 Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>86.53%</td>
<td>78.35%</td>
<td>82.24%</td>
</tr>
<tr>
<td>2</td>
<td>90.05%</td>
<td>84.53%</td>
<td>87.18%</td>
</tr>
<tr>
<td>3</td>
<td>91.12%</td>
<td>86.61%</td>
<td>89.81%</td>
</tr>
<tr>
<td>4</td>
<td>89.66%</td>
<td>82.27%</td>
<td>85.81%</td>
</tr>
<tr>
<td>5</td>
<td>89.94%</td>
<td>79.51%</td>
<td>84.41%</td>
</tr>
<tr>
<td>6</td>
<td>89.10%</td>
<td>76.41%</td>
<td>82.27%</td>
</tr>
<tr>
<td>7</td>
<td>91.04%</td>
<td>94.09%</td>
<td>92.52%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the model architecture, an asymmetric design is adopted for the time-domain and frequency-domain branches, where GCN and GAT are employed in the time-domain branch, while PPNP and GAT are introduced in the frequency-domain branch. To validate the effectiveness of this design, a systematic comparison of different graph neural network combinations is conducted, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>, including configurations that incorporate PPNP in the time domain or use only GCN/GAT in the frequency domain. The experimental results indicate that model performance varies across domains, with PPNP achieving superior performance in the frequency-domain branch (F1-score of 92.52%), while its inclusion in the time domain does not yield significant improvements. From a theoretical perspective, PPNP can be regarded as a spectral-domain filtering method with global smoothing and information propagation properties, which is beneficial for modeling low-frequency structural information in the frequency domain. Therefore, incorporating PPNP in the frequency-domain branch enables more effective feature extraction, whereas GCN and GAT are sufficient for capturing local temporal dependencies in the time domain. Based on both experimental observations and theoretical insights, the proposed asymmetric architecture achieves a favorable balance between detection performance and model complexity.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance comparison under different time&#x2013;frequency branch combinations.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Scheme</th>
<th>Time-Domain Branch</th>
<th>Frequency-Domain Branch</th>
<th>Precision (%)</th>
<th>Rec. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>A</td>
<td>GCN&#x002B;GAT</td>
<td>GCN&#x002B;GAT</td>
<td>89.90</td>
<td>93.83</td>
<td>91.81</td>
</tr>
<tr>
<td>B</td>
<td>GCN&#x002B;PPNP</td>
<td>GCN&#x002B;GAT</td>
<td>88.30</td>
<td>93.54</td>
<td>90.85</td>
</tr>
<tr>
<td>C</td>
<td>PPNP&#x002B;GAT</td>
<td>GCN&#x002B;GAT</td>
<td>89.24</td>
<td>93.71</td>
<td>91.42</td>
</tr>
<tr>
<td>D</td>
<td>GCN&#x002B;GAT</td>
<td>GCN&#x002B;PPNP</td>
<td>90.20</td>
<td>93.89</td>
<td>92.02</td>
</tr>
<tr>
<td>E</td>
<td>GCN&#x002B;GAT</td>
<td>PPNP&#x002B;GAT</td>
<td>91.04</td>
<td>94.09</td>
<td>92.52</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Comparative Experiments</title>
<p>To comprehensively validate the effectiveness and advantages of the proposed method, comparative experiments were conducted against the following four existing approaches:</p>
<p>MTAD-GAT [<xref ref-type="bibr" rid="ref-31">31</xref>] integrates a one-dimensional convolutional neural network (1D CNN), GAT and gated recurrent unit (GRU) to build a multivariate time-series anomaly detection framework that combines reconstruction and prediction, where 1D CNN extracts features, GAT models temporal and feature dependencies, and GRU performs reconstruction and trend prediction. GAT-GRU [<xref ref-type="bibr" rid="ref-8">8</xref>] focuses on multi-dimensional feature extraction across temporal, modal, and spatial domains by first applying GAT to capture local and global dynamic relationships, followed by GRU to model long-term dependencies and perform dimensionality reduction, and finally reapplying GAT for spatial feature extraction and anomaly detection via reconstruction error. GLSL [<xref ref-type="bibr" rid="ref-10">10</xref>] constructs multiple graph structures from a modal perspective for WSN data, utilizes GAT to extract spatio-temporal features across modalities, employs GRU to capture long-term dependencies, and integrates reconstruction and prediction mechanisms to enhance detection performance and adaptability in multi-modal scenarios.</p>
<p>The experimental results are shown in <xref ref-type="table" rid="table-4">Table 4</xref>. Our proposed method achieves optimal or near-optimal performance in terms of Precision, Recall, and F1-score, with an F1-score of 92.52%, surpassing all comparison methods and demonstrating outstanding capability in WSN anomaly detection tasks. Furthermore, the method achieves an AUC of 0.95, indicating robust discrimination ability and stability across different decision thresholds. In terms of model complexity, the parameter count is 0.80 M, which is significantly lower than that of GAT-GRU (36.5 M) and comparable to GLSL (0.6 M), suggesting that the proposed method maintains high performance while remaining computationally efficient and suitable for resource-constrained WSN scenarios.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of experimental results.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Precision (%)</th>
<th>Rec. (%)</th>
<th>F1(%)</th>
<th>AUC</th>
<th>Par/M</th>
</tr>
</thead>
<tbody>
<tr>
<td>MTAD-GAT</td>
<td>77.5</td>
<td>87.0</td>
<td>82.0</td>
<td>0.81</td>
<td>1.1</td>
</tr>
<tr>
<td>GAT-GRU</td>
<td>93.3</td>
<td>87.5</td>
<td>90.3</td>
<td>0.84</td>
<td>36.5</td>
</tr>
<tr>
<td>GLSL</td>
<td>94.5</td>
<td>87.0</td>
<td>90.6</td>
<td>0.93</td>
<td>0.6</td>
</tr>
<tr>
<td>Ours</td>
<td>91.04</td>
<td>94.09</td>
<td>92.52</td>
<td>0.95</td>
<td>0.80</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>A comparative analysis reveals specific limitations of the existing approaches that affect detection effectiveness. MTAD-GAT incorporates a graph attention mechanism to model temporal and feature dimensions, but its adjacency matrix is either static or pre-defined, lacking a data-driven graph structure learning mechanism, resulting in limited adaptability to true inter-node relationships. GAT-GRU attempts multidimensional graph modeling via a multi-level GAT fusion strategy; however, its topology remains rule-based without enhanced topological learning. Despite achieving an F1-score of 90.3% and a Precision of 93.3%, its Recall of 87.5% indicates potential omissions in capturing all anomalous points. GLSL constructs graph structures from a modal perspective and combines reconstruction with prediction mechanisms to enhance robustness, achieving an F1-score of 90.6%. Nonetheless, its modal partitioning is relatively fixed and cannot dynamically learn the coupling relationships or importance differences between time and frequency domains, leaving room for improvement in anomaly feature expressiveness.</p>
<p>In contrast, our method significantly outperforms the aforementioned approaches. With a Precision of 91.04% and a Recall of 94.09%, it ensures accurate anomaly identification while covering a higher proportion of anomalous samples. The improvement in AUC further demonstrates reliable performance across different thresholds, and the low parameter count confirms its efficiency under computationally constrained WSN environments.</p>
<p>In the frequency-domain feature processing stage, a Gaussian filter is introduced to suppress high-frequency noise. Considering that anomaly signals, especially point anomalies, may contain high-frequency components, a range of filter parameters is systematically evaluated, and &#x03C3; &#x003D; 200 is selected. Under this setting, the filter exhibits a weak low-pass characteristic, preserving the main signal components within 0&#x2013;30 Hz while smoothing high-frequency random noise. Given the properties of WSN data, background noise typically manifests as low-amplitude high-frequency fluctuations, whereas anomalies tend to show more significant amplitude variations; thus, the filtering strategy has limited impact on anomaly features. To further illustrate its effectiveness, <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents a comparison of time- and frequency-domain representations before and after filtering, where the positions and amplitudes of anomalies remain largely unchanged in the time domain, and the low-frequency components are well preserved in the frequency domain. In addition, the results in <xref ref-type="table" rid="table-5">Table 5</xref> show that the F1-score improves from 91.89% to 92.52% after applying the filtering strategy, indicating that it effectively reduces noise interference with minimal degradation of anomaly information, thereby enhancing detection performance.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparison of time and frequency domains before and after filtering: (<bold>a</bold>) Time domain before filtering, (<bold>b</bold>) time domain after filtering, (<bold>c</bold>) frequency domain before filtering, (<bold>d</bold>) frequency domain after filtering.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-6.tif"/>
</fig><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of experimental results with and without filtering.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Scheme</th>
<th>Precision (%)</th>
<th>Rec. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>With filtering</td>
<td>91.04</td>
<td>94.09</td>
<td>92.52</td>
</tr>
<tr>
<td>Without filtering</td>
<td>90.02</td>
<td>93.86</td>
<td>91.89</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To evaluate the computational efficiency of the proposed method, we conducted a comparative experiment with the Transformer model. Specifically, the RWKV module in our model was replaced with a standard Transformer, and the CFE module was substituted with a cross-attention mechanism. Inference experiments were then performed, and the single-sample inference time and peak memory usage were recorded. The results are summarized in <xref ref-type="table" rid="table-6">Table 6</xref>:</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparative experimental between RWKV-based and transformer-based models.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Single-Sample Inference Time</th>
<th>Peak Memory Usage</th>
</tr>
</thead>
<tbody>
<tr>
<td>RWKV-based</td>
<td>144 ms</td>
<td>1895.67 MB</td>
</tr>
<tr>
<td>Transformer-based</td>
<td>441 ms</td>
<td>2366.27 MB</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To evaluate the impact of the parameter K in the Top-K graph construction strategy, a systematic analysis was conducted under different K settings, as shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparative experimental under different K values.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>K Value</th>
<th>Precision (%)</th>
<th>Rec. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>A</td>
<td>1</td>
<td>87.41</td>
<td>77.81</td>
<td>82.33</td>
</tr>
<tr>
<td>B</td>
<td>2</td>
<td>91.67</td>
<td>82.26</td>
<td>86.71</td>
</tr>
<tr>
<td>C</td>
<td>3</td>
<td>91.04</td>
<td>94.09</td>
<td>92.52</td>
</tr>
<tr>
<td>D</td>
<td>4</td>
<td>81.86</td>
<td>84.48</td>
<td>83.15</td>
</tr>
<tr>
<td>E</td>
<td>5</td>
<td>75.91</td>
<td>82.81</td>
<td>79.21</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be observed that when K is relatively small, the constructed graph is sparse, and potential inter-node relationships are not fully captured, leading to suboptimal performance in terms of Precision, Recall, and F1-score. As K increases, the model performance improves and reaches its optimum at K &#x003D; 3, achieving the highest F1-score of 92.52%. However, when K becomes excessively large, the graph becomes increasingly dense, which may introduce redundant connections and noise, thereby interfering with feature propagation and degrading the overall performance. These results indicate that an appropriate choice of K can effectively balance the preservation of informative topological structures and the suppression of noise, and thus K &#x003D; 3 is selected as the default setting in this study.</p>
<p>To further validate the anomaly detection performance of the proposed TE-MSTAD network model across diverse application scenarios, this paper designed comparative experiments based on indoor and outdoor multi-scenario environments. The typical indoor sensor network dataset IBRL and the self-built outdoor LoRA-OSD real-world monitoring dataset were selected, covering complex data features from multi-modal, multi-node collection in wireless sensor networks under varying conditions. By training and testing the model on both datasets, we comprehensively evaluate the proposed method&#x2019;s generalization capability and robustness across multiple scenarios and data distributions. Specific experimental results are shown in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Comparative experimental results across different scenarios.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Precision (%)</th>
<th>Rec. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>IBRL</td>
<td>91.04</td>
<td>94.09</td>
<td>92.52</td>
</tr>
<tr>
<td>LoRA-OSD</td>
<td>92.57</td>
<td>94.05</td>
<td>93.28</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The table demonstrates that the proposed TE-MSTAD model achieves outstanding anomaly detection performance on both datasets, maintaining high levels across all metrics. For the indoor IBRL dataset, the model exhibits high recall in capturing complex multi-node, multi-modal relationships, indicating strong sensitivity to anomalous data. On the outdoor LoRA-OSD dataset with complex scenarios, the model demonstrates robust adaptability and stability, achieving an F1 score slightly higher than IBRL. This validates TE-MSTAD&#x2019;s generalization capability and robustness across diverse environments. This further demonstrates that TE-MSTAD can effectively enhance the accuracy and robustness of anomaly detection in wireless sensor networks, exhibiting strong practical potential in complex, multi-modal, multi-node industrial scenarios.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Visualization Analysis</title>
<p>In this section, several representative experimental samples are selected to conduct a case study on the anomaly detection performance of the proposed model. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, temperature and humidity data collected from node 21 over 4800 time steps are presented, where correlation anomalies are injected into the data. In the figure, the red and green curves represent the temperature and humidity measurements, respectively, the black line denotes the ground-truth anomaly labels, the blue line indicates the model predictions, and the orange markers highlight misclassified points. This visualization clearly demonstrates the effectiveness of the model in handling cross-modal correlation anomalies and further validates the collaborative functionality and rational design of the proposed modules.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Anomaly detection of multimodal correlations in WSN.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-7.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-8">Fig. 8a</xref> illustrates the temperature time series of node 21 in the IBRL dataset over 2240 time steps. The blue curve represents the original normal data, while the red, green, and purple curves correspond to injected point anomalies, contextual anomalies, and collective anomalies, respectively. In <xref ref-type="fig" rid="fig-8">Fig. 8b</xref>, a sliding window approach with a window size of 200 is employed to construct test samples, and anomaly detection is performed based on the model predictions. <xref ref-type="fig" rid="fig-8">Fig. 8b</xref> shows the distribution of anomaly labels for different anomaly types, while <xref ref-type="fig" rid="fig-8">Fig. 8c</xref> presents the detection results. The comparison indicates that the predicted labels are highly consistent with the ground-truth labels in most cases, with a small number of misclassifications mainly occurring when the sliding window first covers anomaly points or in segments with significant fluctuations in the original data. Overall, the proposed method is capable of accurately identifying multiple types of anomalies, demonstrating strong practicality and robustness.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>WSN single-node anomaly detection: (<bold>a</bold>) Test sample data; (<bold>b</bold>) sample labels (<bold>c</bold>) anomaly detection labels.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-8.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-9">Fig. 9a</xref> shows the temperature time series of three neighboring nodes (23, 25 and 27), where point anomalies are injected into node 23. <xref ref-type="fig" rid="fig-9">Fig. 9b</xref> presents the corresponding anomaly detection results. It can be observed that the model successfully detects all point anomalies at node 23, indicating strong capability in capturing local temporal perturbations. Meanwhile, no false positives are observed for nodes 25 and 27, suggesting that by explicitly modeling graph structural dependencies, the proposed method effectively constrains the propagation of anomaly responses, ensuring that anomalies remain localized to the affected nodes and avoiding spurious detections caused by spatial correlation propagation.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Correlation anomaly detection among multiple nodes in WSN: (<bold>a</bold>) Multi-node sample data; (<bold>b</bold>) multi-node anomaly detection label.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78282-fig-9.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This paper proposes a multi-modal spatio-temporal anomaly detection method that enhances topological information by integrating temporal and frequency domain features. Building upon an improved RWKV model, the method introduces cross-extraction blocks and employs a dual-branch network architecture. This approach fully exploits spatio-temporal dependencies among multi-node multi-modal data, significantly enhancing anomaly detection performance. Furthermore, leveraging information enhancement principles, the method integrates graph neural network submodels and jointly learns graph structure adjacency matrices across time and frequency domains, thereby strengthening its ability to capture spatial correlations among different nodes. Experimental results demonstrate superior performance of the proposed method compared to traditional multi-time-series detection models on public datasets. Furthermore, experiments applying the model to both public and real-world datasets achieve excellent detection outcomes, validating its robust detection capability and generalization performance. Future research may further explore lightweight architectures, spatio-temporal adaptive mechanisms, and diverse training strategies to enhance the model&#x2019;s detection performance, deployability, and real-time capability in complex scenarios such as industrial IoT.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their sincere gratitude to all collaborators for their valuable comments, suggestions and support that have greatly improved the quality of this paper.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research is funded in part by The National Natural Science Foundation of China (No. 62161006), Guangxi Science and Technology Program under Grant No. FN2504240022 and Innovation Project of GUET Graduate Education (No. 2025YCXS078).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Miao Ye and Ziheng Wang conceived the study, performed experiments and drafted the manuscript. Qiuxiang Jiang contributed analytical tools and analyzed the data. Xingsi Xue and Wenxi Liu assisted with data collection and validation. Yu Ning supervised statistical analysis. Cheng Zhu designed and supervised the overall project and revised the manuscript. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data and materials used to support the findings of this study are available from the corresponding author upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Behera</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Mohapatra</surname> <given-names>SK</given-names></string-name></person-group>. <article-title>A novel scheme for mitigation of energy hole problem in wireless sensor network for military application</article-title>. <source>Int J Commun</source>. <year>2021</year>;<volume>34</volume>(<issue>11</issue>):<fpage>e4886</fpage>. doi:<pub-id pub-id-type="doi">10.1002/dac.4886</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dang</surname> <given-names>TB</given-names></string-name>, <string-name><surname>Le</surname> <given-names>DT</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>TD</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>M</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Monotone split and conquer for anomaly detection in IoT sensory data</article-title>. <source>IEEE Internet Things J</source>. <year>2021</year>;<volume>8</volume>(<issue>20</issue>):<fpage>15468</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2021.3073705</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Haque</surname> <given-names>ME</given-names></string-name>, <string-name><surname>Asikuzzaman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>IU</given-names></string-name>, <string-name><surname>Ra</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Shah</surname> <given-names>SBH</given-names></string-name></person-group>. <article-title>Comparative study of IoT-based topology maintenance protocol in a wireless sensor network for structural health monitoring</article-title>. <source>Remote Sens</source>. <year>2020</year>;<volume>12</volume>(<issue>15</issue>):<fpage>2358</fpage>. doi:<pub-id pub-id-type="doi">10.3390/rs12152358</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wilson</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Radhamani</surname> <given-names>AS</given-names></string-name></person-group>. <article-title>Real time flood disaster monitoring based on energy efficient ensemble clustering mechanism in wireless sensor network</article-title>. <source>Softw Pract Exp</source>. <year>2022</year>;<volume>52</volume>(<issue>1</issue>):<fpage>254</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1002/spe.3019</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname> <given-names>T</given-names></string-name>, <string-name><surname>Irfan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shaf</surname> <given-names>A</given-names></string-name>, <string-name><surname>Saeed Alwadie</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sajid</surname> <given-names>A</given-names></string-name>, <string-name><surname>Awais</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A secure communication in IoT enabled underwater and wireless sensor network for smart cities</article-title>. <source>Sensors</source>. <year>2020</year>;<volume>20</volume>(<issue>15</issue>):<fpage>4309</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s20154309</pub-id>; <pub-id pub-id-type="pmid">32748819</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boubiche</surname> <given-names>DE</given-names></string-name>, <string-name><surname>Athmani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Boubiche</surname> <given-names>S</given-names></string-name>, <string-name><surname>Toral-Cruz</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Cybersecurity issues in wireless sensor networks: current challenges and solutions</article-title>. <source>Wirel Pers Commun</source>. <year>2021</year>;<volume>117</volume>(<issue>1</issue>):<fpage>177</fpage>&#x2013;<lpage>213</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11277-020-07213-5</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Security and application of wireless sensor network</article-title>. <source>Procedia Comput Sci</source>. <year>2021</year>;<volume>183</volume>(<issue>3</issue>):<fpage>486</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2021.02.088</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>M</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A novel anomaly detection method for multimodal WSN data flow via a dynamic graph neural network</article-title>. <source>Connect Sci</source>. <year>2022</year>;<volume>34</volume>(<issue>1</issue>):<fpage>1609</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1080/09540091.2022.2078281</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Multimodal adversarial learning based unsupervised time series anomaly detection</article-title>. <source>J Comput Res Dev</source>. <year>2021</year>;<volume>58</volume>(<issue>8</issue>):<fpage>1655</fpage>&#x2013;<lpage>67</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A novel self-supervised learning-based anomalous node detection method based on an autoencoder for wireless sensor networks</article-title>. <source>IEEE Syst J</source>. <year>2024</year>;<volume>18</volume>(<issue>1</issue>):<fpage>256</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSYST.2023.3347435</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmad</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alkhammash</surname> <given-names>EH</given-names></string-name></person-group>. <article-title>Online adaptive Kalman filtering for real-time anomaly detection in wireless sensor networks</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>15</issue>):<fpage>5046</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24155046</pub-id>; <pub-id pub-id-type="pmid">39124095</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>N</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Automatic fuzzy architecture design for defect detection via classifier-assisted multiobjective optimization approach</article-title>. <source>IEEE Trans Evol Comput</source>. <year>2025</year>;<volume>30</volume>(<issue>2</issue>):<fpage>479</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TEVC.2025.3530416</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A novel spatiotemporal correlation anomaly detection method based on time-frequency-domain feature fusion and a dynamic graph neural network in wireless sensor network</article-title>. <source>IEEE Sens J</source>. <year>2025</year>;<volume>25</volume>(<issue>9</issue>):<fpage>15548</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2025.3549220</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dener</surname> <given-names>M</given-names></string-name>, <string-name><surname>Okur</surname> <given-names>C</given-names></string-name>, <string-name><surname>Al</surname> <given-names>S</given-names></string-name>, <string-name><surname>Orman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>WSN-BFSF: a new data set for attacks detection in wireless sensor networks</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>11</volume>(<issue>2</issue>):<fpage>2109</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2023.3292209</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Javed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Larijani</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ahmadinia</surname> <given-names>A</given-names></string-name>, <string-name><surname>Emmanuel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Mannion</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gibson</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Design and implementation of a cloud enabled random neural network-based decentralized smart controller with intelligent sensor nodes for HVAC</article-title>. <source>IEEE Internet Things J</source>. <year>2017</year>;<volume>4</volume>(<issue>2</issue>):<fpage>393</fpage>&#x2013;<lpage>403</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2016.2627403</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Felici-Castell</surname> <given-names>S</given-names></string-name>, <string-name><surname>Segura-Garcia</surname> <given-names>J</given-names></string-name>, <string-name><surname>Perez-Solano</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Fayos-Jordan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Soriano-Asensi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alcaraz-Calero</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>AI-IoT low-cost pollution-monitoring sensor network to assist citizens with respiratory problems</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>23</issue>):<fpage>9585</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23239585</pub-id>; <pub-id pub-id-type="pmid">38067957</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hsiung</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>HY</given-names></string-name></person-group>. <article-title>GRNN-based detection of eavesdropping attacks in SWIPT-enabled smart grid wireless sensor networks</article-title>. <source>IEEE Internet Things J</source>. <year>2024</year>;<volume>11</volume>(<issue>22</issue>):<fpage>37381</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3443277</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. In: <conf-name>Proceedings of 31st International Conference on Neural Information Processing Systems (NIPS 2017)</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates</publisher-name>; <year>2017</year>. p. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>AutoGMM-RWKV: a detecting scheme based on attention mechanisms against selective forwarding attacks in wireless sensor networks</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>4</issue>):<fpage>4403</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2024.3484999</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gr&#x00F6;chenig</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Uncertainty principles for time-frequency representations</article-title>. In: <conf-name>Advances in Gabor Analysis</conf-name>. <publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>Birkh&#x00E4;user Boston</publisher-name>; <year>2003</year>. p. <fpage>11</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-1-4612-0133-5_2</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sha</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Anomaly detection method for industrial control system operation data based on time-frequency fusion feature attention encoding</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>18</issue>):<fpage>6131</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24186131</pub-id>; <pub-id pub-id-type="pmid">39338876</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kipf</surname> <given-names>TN</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <source>Semi-supervised classification with graph convolutional networks</source>. <comment>arXiv:1609.02907. 2016</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Veli&#x010D;kovi&#x0107;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cucurull</surname> <given-names>G</given-names></string-name>, <string-name><surname>Casanova</surname> <given-names>A</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li&#x00F2;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph attention networks</article-title>. In: <conf-name>Proceedings of the International Conference on Learning Representations (ICLR); 2018 Apr 30&#x2013;May 3</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Klicpera</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bojchevski</surname> <given-names>A</given-names></string-name>, <string-name><surname>G&#x00FC;nnemann</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Predict then propagate: graph neural networks meet personalized PageRank</article-title>. In: <conf-name>Proceedings of the International Conference on Learning Representations (ICLR); 2019 May 6&#x2013;9</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Alcaide</surname> <given-names>E</given-names></string-name>, <string-name><surname>Anthony</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Albalak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Arcadinho</surname> <given-names>S</given-names></string-name>, <string-name><surname>Biderman</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RWKV: reinventing RNNs for the transformer era</article-title>. <comment>arXiv:2305.13048. 2023</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thomas</surname> <given-names>D</given-names></string-name>, <string-name><surname>Shankaran</surname> <given-names>R</given-names></string-name>, <string-name><surname>Orgun</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Mukhopadhyay</surname> <given-names>SC</given-names></string-name></person-group>. <article-title>SEC2: a secure and energy efficient barrier coverage scheduling for WSN-based IoT applications</article-title>. <source>IEEE Trans Green Commun Netw</source>. <year>2021</year>;<volume>5</volume>(<issue>2</issue>):<fpage>622</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TGCN.2021.3067606</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>TKK</given-names></string-name>, <string-name><surname>Karami</surname> <given-names>A</given-names></string-name>, <string-name><surname>Armanfard</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Graph-based time-series anomaly detection: a survey and outlook</article-title>. <comment>arXiv:2302.00058. 2024</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Data-driven anomaly detection approach for time-series streaming data</article-title>. <source>Sensors</source>. <year>2020</year>;<volume>20</volume>(<issue>19</issue>):<fpage>5646</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s20195646</pub-id>; <pub-id pub-id-type="pmid">33023175</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>ITran: a novel transformer-based approach for industrial anomaly detection and localization</article-title>. <source>Eng Appl Artif Intell</source>. <year>2023</year>;<volume>125</volume>(<issue>1</issue>):<fpage>106677</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2023.106677</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bilal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Masked Swin transformer Unet for industrial anomaly detection</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2023</year>;<volume>19</volume>(<issue>2</issue>):<fpage>2200</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2022.3199228</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Shah</surname> <given-names>N</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A synergistic approach for graph anomaly detection with pattern mining and feature learning</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2022</year>;<volume>33</volume>(<issue>6</issue>):<fpage>2393</fpage>&#x2013;<lpage>405</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2021.3102609</pub-id>; <pub-id pub-id-type="pmid">34460385</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Interpretable spatial-temporal graph convolutional network for system log anomaly detection</article-title>. <source>Adv Eng Inform</source>. <year>2024</year>;<volume>62</volume>(<issue>3</issue>):<fpage>102803</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2024.102803</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choi</surname> <given-names>SH</given-names></string-name>, <string-name><surname>An</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>I</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Anomaly detection based on graph convolutional network-variational autoencoder model using time-series vibration and current data</article-title>. <source>Mathematics</source>. <year>2024</year>;<volume>12</volume>(<issue>23</issue>):<fpage>3750</fpage>. doi:<pub-id pub-id-type="doi">10.3390/math12233750</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ya&#x011F;ci</surname> <given-names>MY</given-names></string-name>, <string-name><surname>Ali Aydin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>EA-GAT: event aware graph attention network on cyber-physical systems</article-title>. <source>Comput Ind</source>. <year>2024</year>;<volume>159</volume>(<issue>1</issue>):<fpage>104097</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compind.2024.104097</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph attention network and informer for multivariate time series anomaly detection</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>5</issue>):<fpage>1522</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24051522</pub-id>; <pub-id pub-id-type="pmid">38475058</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nambo</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Plant biopotential sensing based on generative adversarial networks for environmental anomaly detection</article-title>. <source>IEEE Sens J</source>. <year>2023</year>;<volume>23</volume>(<issue>23</issue>):<fpage>29793</fpage>&#x2013;<lpage>803</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2023.3323147</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Long</surname> <given-names>X</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Adversarial regularized graph autoencoder for intelligent anomaly detection with multisensor signal fusion</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2025</year>;<volume>74</volume>:<fpage>3534011</fpage>. doi:<pub-id pub-id-type="doi">10.1109/TIM.2025.3565242</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Plaza</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Hyperspectral anomaly detection using reconstruction fusion of quaternion frequency domain analysis</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2024</year>;<volume>35</volume>(<issue>6</issue>):<fpage>8358</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2022.3227167</pub-id>; <pub-id pub-id-type="pmid">37022253</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmad</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alhasan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wazirali</surname> <given-names>R</given-names></string-name>, <string-name><surname>Almajalid</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A reliable approach for lightweight anomaly detection in sensors using continuous wavelet transform and vector clustering</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>15</issue>):<fpage>24921</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2024.3407158</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Le</surname> <given-names>KT</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>TB</given-names></string-name>, <string-name><surname>Le</surname> <given-names>DT</given-names></string-name>, <string-name><surname>Raza</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>M</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>H</given-names></string-name></person-group>. <article-title>VEAD: variance profile exploitation for anomaly detection in real-time IoT data streaming</article-title>. <source>Internet Things</source>. <year>2024</year>;<volume>25</volume>(<issue>3</issue>):<fpage>100994</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iot.2023.100994</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Park</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Lightweight machine sound anomaly detector based on parallel discrete wavelet transform</article-title>. <source>IEEE Sens J</source>. <year>2025</year>;<volume>25</volume>(<issue>10</issue>):<fpage>18529</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JSEN.2025.3556204</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Doroudi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lavassani</surname> <given-names>SHH</given-names></string-name>, <string-name><surname>Shahrouzi</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Optimal tuning of three deep learning methods with signal processing and anomaly detection for multi-class damage detection of a large-scale bridge</article-title>. <source>Struct Health Monit</source>. <year>2024</year>;<volume>23</volume>(<issue>5</issue>):<fpage>3227</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1177/14759217231216694</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Anomaly detection collaborating adaptive CEEMDAN feature exploitation with intelligent optimizing classification for IIoT sparse data</article-title>. <source>Wirel Commun Mob Comput</source>. <year>2021</year>;<volume>2021</volume>(<issue>1</issue>):<fpage>4329219</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/4329219</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>