<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">60564</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.060564</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Robust Deep One-Class Classification Time Series Anomaly Detection</article-title>
<alt-title alt-title-type="left-running-head">Robust Deep One-Class Classification Time Series Anomaly Detection</alt-title>
<alt-title alt-title-type="right-running-head">Robust Deep One-Class Classification Time Series Anomaly Detection</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Zhengdao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Xuewei</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Chen</surname><given-names>Yuling</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>ylchen3@gzu.edu.cn</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Dou</surname><given-names>Hui</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Sang</surname><given-names>Haiwei</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University</institution>, <addr-line>Guiyang, 550000</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Computer Science and Technology, Weifang University of Science and Technology</institution>, <addr-line>Weifang, 261000</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>School of Mathematics and Big Data, Guizhou Education University</institution>, <addr-line>Guiyang, 550018</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yuling Chen. Email: <email>ylchen3@gzu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>19</day><month>05</month><year>2025</year>
</pub-date>
<volume>83</volume>
<issue>3</issue>
<fpage>5181</fpage>
<lpage>5197</lpage>
<history>
<date date-type="received">
<day>04</day>
<month>11</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>3</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_60564.pdf"></self-uri>
<abstract>
<p>Anomaly detection (AD) in time series data is widely applied across various industries for monitoring and security applications, emerging as a key research focus within the field of deep learning. While many methods based on different normality assumptions perform well in specific scenarios, they often neglected the overall normality issue. Some feature extraction methods incorporate pre-training processes but they may not be suitable for time series anomaly detection, leading to decreased performance. Additionally, real-world time series samples are rarely free from noise, making them susceptible to outliers, which further impacts detection accuracy. To address these challenges, we propose a novel anomaly detection method called Robust One-Class Classification Detection (ROC). This approach utilizes an autoencoder (AE) to learn features while constraining the context vectors from the AE within a sufficiently small hypersphere, akin to One-Class Classification (OC) methods. By simultaneously optimizing two hypothetical objective functions, ROC captures various aspects of normality. We categorize the input raw time series into clean and outlier sequences, reducing the impact of outliers on compressed feature representation. Experimental results on public datasets indicate that our approach outperforms existing baseline methods and substantially improves model robustness.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Time series anomaly detection</kwd>
<kwd>self-supervised learning</kwd>
<kwd>robustness</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation</funding-source>
<award-id>62202118</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Guizhou Province Major Project (Qiankehe Major Project</funding-source>
<award-id>[2024]014</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Science and Scientific and Technological Research Projects</funding-source>
<award-id>[2023]003</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Science and Technology Department</funding-source>
<award-id>[2023]018</award-id>
</award-group>
<award-group id="awg5">
<funding-source>Guizhou Province Major Project</funding-source>
<award-id>[2024]003</award-id>
</award-group>
<award-group id="awg6">
<funding-source>Key Laboratory of Public Big Data Security Technology</funding-source>
<award-id>QJ202300001</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Analyzing time series allows us to understand the underlying processes that generate these sequences, thereby enhancing our comprehension of these processes. Anomaly detection in time series is a critical challenge in data mining, with applications across diverse fields such as transportation and manufacturing, where it serves to monitor system behavior. For instance, in aircraft engine fault detection, monitoring the time series of engine RPM, fuel pressure, temperature, and vibration signals enables the identification of abnormal fluctuations or patterns, allowing for early detection of mechanical failures and preventing accidents during flight. In railway transportation systems, monitoring the status of trains involves real-time analysis of data such as wheel temperature, axle vibration, and electrical signals. By detecting anomalies, potential hazards that could lead to derailments or failures can be identified and mitigated promptly. In manufacturing production lines, equipment health management is facilitated through the monitoring of vibration, temperature, and current of machinery via time series data, enabling the identification of potential failures and timely maintenance, thus minimizing downtime and ensuring production efficiency.</p>
<p>Time Series Anomaly Detection (TSAD) is also essential in applications like health monitoring and fraud detection, where it involves identifying unique time series instances that deviate from typical patterns [<xref ref-type="bibr" rid="ref-1">1</xref>]. For example, in health monitoring, detecting abnormal heart rates and brain waves in electrocardiogram (ECG) and electroencephalogram (EEG) data can reveal health issues such as arrhythmias and epilepsy. In credit card fraud detection, analyzing user transaction patterns and identifying unusual behaviors, such as large purchases in a short time or transactions at atypical locations, helps in detecting and preventing fraudulent activities. Additionally, in network traffic anomaly detection, analyzing time series data of network traffic can uncover abnormal data transmission behaviors, such as DDoS attacks, data breaches, or malicious software activities.</p>
<p>In recent years, deep learning-based approaches have achieved impressive results in time series anomaly detection, particularly with complex datasets. These methods excel at modeling long-term and nonlinear temporal patterns within the data, surpassing traditional methods such as similarity search and density-based clustering [<xref ref-type="bibr" rid="ref-2">2</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>]. Neural network-based approaches commonly use an encoder to compress the input time series into a compact latent representation, which is then decoded to reconstruct the original series. This encoder-decoder framework, commonly referred to as an autoencoder [<xref ref-type="bibr" rid="ref-6">6</xref>], compresses the initial input through a bottleneck layer, promoting the formation of compact latent representations. This structure enables the model to capture essential patterns within the time series while filtering out irrelevant or atypical patterns, such as anomalies [<xref ref-type="bibr" rid="ref-7">7</xref>]. This approach aids in the detection of anomalies by evaluating the reconstruction error between the original time series and its reconstructed counterpart. A larger reconstruction error suggests an increased probability that the related observations are anomalies. One-class (OC) techniques consolidate normal instances into a single category by minimizing the volume of the hypersphere that encompasses the feature representations. However, despite the strong performance demonstrated by deep learning techniques [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>] such as autoencoders, they encounter two major challenges.</p>
<p><bold>Single Hypothesis:</bold> A single hypothesis often captures only a limited aspect of sample normality, while anomalies can manifest in diverse forms, such as point anomalies, subsequence anomalies, and anomalies across entire time series. Inspired by the principles of ensemble anomaly detection methods, we suggest that detectors relying on a singular hypothesis may be insufficient to capture this diversity, potentially limiting their effectiveness in detecting varied anomaly types [<xref ref-type="bibr" rid="ref-11">11</xref>]. Consequently, while these methods may excel in detecting specific types of anomalies, their effectiveness often diminishes when encountering others.</p>
<p>An alternative approach divides the process into two phases: pre-training on the overall time series, followed by fine-tuning for anomaly detection (AD). For instance, deep SVDD [<xref ref-type="bibr" rid="ref-12">12</xref>] initially utilizes an autoencoder for feature extraction and subsequently refines these features for anomaly detection through a one-class loss function. Similarly, Reference [<xref ref-type="bibr" rid="ref-13">13</xref>] applies contrastive learning in the first phase and utilize one-class methods for detection in the second phase. Formally, these approaches combine feature extraction methods with the assumptions of one-class learning; however, the objectives of the two phases are often misaligned. While the pre-training phase may yield representations that align with typical patterns, these representations can also be influenced by extraneous features unrelated to anomaly detection. As a result, the performance of such methods may be constrained by these pre-trained features.</p>
<p><bold>Robustness:</bold> In unsupervised learning, training data often includes anomalies. Because the encoder compresses all observations in the input time series, even those that are anomalous, the hidden representation may become susceptible to these outliers. Particularly when anomalies are high in amplitude, even a few can compromise the latent information, leading to a risk that the latent representation itself reflects anomalous patterns from the training data.</p>
<p>Deep learning-based soft sensors for time series anomaly detection exhibit significant vulnerabilities to adversarial attacks. Knowledge-guided adversarial perturbations can be designed to subtly manipulate the input data distribution without causing noticeable changes, resulting in a substantial degradation of model performance. This exposes the insufficient robustness of current deep learning methods in time series anomaly detection [<xref ref-type="bibr" rid="ref-14">14</xref>]. Reference [<xref ref-type="bibr" rid="ref-15">15</xref>] proposes an adversarial training strategy using historical gradients and domain adaptation. By leveraging historical information to capture temporal dynamics and mapping input samples to a shared feature space, this approach effectively enhances model robustness against adversarial examples. This is particularly critical for time series anomaly detection, where data distribution changes introduce additional uncertainty. As a result, the model may produce lower reconstruction errors for specific anomalies, making it difficult to distinguish them from normal data, which in turn negatively affects accuracy. As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, training set data contaminated by anomalies can lead to the model generating smaller reconstruction errors for similar anomalies during testing, complicating their detection. To mitigate this issue, a robust solution is required to ensure that the latent representation is not influenced by anomalies present in the training data.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>A portion of the time series from the training set for anomaly detection, where the orange shading represents the actual anomalous segments and the red solid line represents the anomaly score. In unsupervised anomaly detection, the training set typically relies on benign samples, which results in the model learning characteristics associated with anomalies. Consequently, the model becomes less sensitive to certain anomalies present in the test set</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-1.tif"/>
</fig>
<p>We introduce a novel autoencoder framework to address the issues of robustness and single hypothesis in anomaly detection. In this study, we introduce a single-stage anomaly detection approach termed the Robust One-Class Classification (ROC) autoencoder. We hypothesize that normal samples will be more accurately reconstructed, with their projection vectors in the latent space forming a compact hypersphere.</p>
<p>Rather than reconstructing the input time series <italic>T</italic> directly, we decompose it into two components: the clean time series <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> and the anomalous time series <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>. Using an integrated autoencoder, we then reconstruct only the clean component <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula>, thereby ensuring that the latent representation remains unaffected by anomalous data. This approach enhances robustness and, consequently, improves the model&#x2019;s accuracy.</p>
<p>In summary, the main contributions of this paper are as follows:
<list list-type="bullet">
<list-item>
<p>The proposed method differs from traditional single-hypothesis autoencoder approaches by integrating a multi-hypothesis anomaly detection framework that combines autoencoders with OC methods, capturing richer feature representations in the time series.</p></list-item>
<list-item>
<p>Unlike standard denoising autoencoders, our method does not require additional noise-free training data. By optimizing the sparsity of <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> in the objective function, we can effectively separate the anomalous components in the training data, achieving the effect of noise-free training data.</p></list-item>
<list-item>
<p>The effectiveness of the proposed approach has been validated on publicly available time series datasets, where it demonstrates superior performance compared to existing methods.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Problem Definition</title>
<p>A time series <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo></mml:math></inline-formula> consists of <italic>C</italic> observations, with each observation <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> in <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>D</mml:mi></mml:msup></mml:math></inline-formula>. If <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, then <italic>T</italic> is a univariate time series; if <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>D</mml:mi><mml:mo>&gt;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, it represents a multivariate or multidimensional time series.</p>
<p>For a specified time series <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo></mml:math></inline-formula>, we aim to calculate the anomaly score <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>O</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for every observation <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. Observations with higher anomaly scores are more likely to be classified as anomalies. Anomalies are not specifically categorized as point anomalies or sequence anomalies; rather, a set of consecutive observations is considered a sequence anomaly if they share high anomaly scores.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deep Learning-Based Anomaly Detection</title>
<p>Traditional anomaly detection methods rely on statistical features, but these approaches are often unsuitable for time series data, leads to the feature information not being correctly represented. With the increase in data dimensionality and volume, deep learning methods have emerged. OC methods can capture complex features that represent &#x201C;normality.&#x201D; For instance, methods based on Generative Adversarial Networks (GANs) [<xref ref-type="bibr" rid="ref-16">16</xref>] and autoencoders [<xref ref-type="bibr" rid="ref-17">17</xref>] assume that normal samples can be reconstructed better by the model. Clustering methods, on the other hand, posit that normal samples cluster together as a large group, while outlier data points are classified as anomalies [<xref ref-type="bibr" rid="ref-18">18</xref>]. In contrast, contrastive learning approaches enhance the data by making positive samples closer to each other while causing the negative samples to be far apart from each other [<xref ref-type="bibr" rid="ref-19">19</xref>]. However, these assumptions can often be overly simplistic or effective only for specific types of anomalies. Additionally, another class of deep learning methods, such as Deep SVDD, employs a two-stage OC classifier. This involves feature extraction using a pre-trained encoder model, which is then classified using OC-SVM [<xref ref-type="bibr" rid="ref-12">12</xref>]. Nevertheless, this approach tends to separate the training objective from the downstream task, hindering the effective learning of diverse time series features.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Robust Principal Component Analysis</title>
<p>Given a matrix <italic>M</italic>, Principal Component Analysis (PCA) is capable of discovering a low-rank matrix that serves as an approximation of <italic>M</italic>. However, as PCA generally employs Singular Value Decomposition (SVD) to determine the low-rank matrix, it exhibits similar sensitivity to outliers as that observed in SVD. To enhance the effectiveness of Principal Component Analysis (PCA) in the presence of outliers, Robust Principal Component Analysis (RPCA) [<xref ref-type="bibr" rid="ref-20">20</xref>] has been introduced. The aim of RPCA is to decompose the matrix <italic>M</italic> into two components: a low-rank matrix <italic>L</italic> that represents the underlying clean data and a sparse matrix <italic>S</italic> that captures the outliers. Specifically, RPCA expresses the original matrix <italic>M</italic> as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mi>S</mml:mi></mml:math></disp-formula></p>
<p>In this decomposition, <italic>L</italic> serves as a low-rank matrix that approximates the clean data within the original matrix <italic>M</italic>, while <italic>S</italic> represents a sparse matrix composed of elements identified as outliers, which are not encapsulated by the low-rank matrix <italic>L</italic>. RPCA accomplishes this decomposition by resolving the optimization problem outlined in the following formula.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:munder><mml:mrow><mml:mtext>argmin</mml:mtext></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:mspace width="1em" /><mml:mrow><mml:mtext>rank</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>S</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>0</mml:mn></mml:msub><mml:mspace width="1em" /><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mi>S</mml:mi></mml:math></disp-formula></p>
<p>In this context, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mtext>rank</mml:mtext><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the rank of matrix <italic>L</italic>; <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>S</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> is the <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> norm of matrix <italic>S</italic>, which indicates the count of non-zero elements within <italic>S</italic>; and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula> serves as a parameter that adjusts the relative significance of <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>S</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula>. Additionally, given that <italic>M</italic> is expressed as the sum of <italic>L</italic> and <italic>S</italic>, the optimization is subject to the constraint <italic>M</italic> &#x003D; <italic>L</italic> &#x002B; <italic>S</italic>. To discover a low-rank matrix <italic>L</italic> that closely approximates the original matrix <italic>M</italic> and a sparse matrix <italic>S</italic> capturing the outliers, the loss function is minimized. Although RPCA is efficient in detecting and excluding outliers, it is limited by its lack of support for time series data and its restriction to linear transformations. Time series data, however, frequently involves complex, nonlinear variations that this approach cannot fully address.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<sec id="s3_1">
<label>3.1</label>
<title>Overall Approach</title>
<p>In unsupervised learning, training data may already contain anomalies. As the encoder compresses the time series, the hidden representation becomes highly sensitive to these anomalies. This means that the anomalous information in the training data can contaminate the latent representation, causing the model to learn abnormal features. As a result, the model may exhibit low reconstruction errors for anomalous samples, making it difficult to distinguish them from clean samples.</p>
<p>To address this issue, we propose a robust method that ensures the latent representation is not affected by anomalies in the training data. Drawing on the principles of RPCA, our proposed neural network is designed to partition the input data during training into two segments: the clean time series <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> and the anomalous component <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>, ensuring that <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>. We then input the clean series <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> into a deep one-class (OC) method to capture the overall features of the positive samples under multiple hypotheses. The specific approach is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The architecture of the proposed ROC model is depicted, where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>c</mml:mi></mml:math></inline-formula> represents the center of the hypersphere. Projection vectors of normal instances (e.g., <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>) are contained within the hypersphere, while those of anomalous instances are positioned beyond its boundary. Simultaneously, we separate <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> from <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula>, and only clean time series are used as input to the autoencoder, resulting in the reconstruction of the time series <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msubsup><mml:mi>T</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Deep OC</title>
<p>In our proposed deep one-class (OC) method, we combine two assumptions: that normal samples can be better reconstructed and that when mapped to a high-dimensional space, they will form a hypersphere with a smaller radius. This approach allows for a more comprehensive learning of the features of clean samples. The motivation behind this is that under a single assumption, the normality features learned by the model may be one-sided, leading to a situation where certain features of abnormal samples differ from those of normal samples, but those features are not recognized by the model. As a result, specific types of anomalies may go undetected by a model based on a single assumption. The specific objective function is as follows:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>O</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>For a given set of <italic>N</italic> time series training samples, the objective function consists of two parts: one for learning the features of the time series from the perspective of the autoencoder and the other from the one-class (OC) method. The parameter <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> controls the weights of these two components. <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the reconstruction error of the seq2seq model, defined as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mtext>ae</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>AE</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>AE</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the reconstructed time series, defined as <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>AE</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mtext>AE</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This is primarily calculated by measuring the mean squared error between the original sequence and the reconstructed sequence, ensuring the model&#x2019;s reconstruction is as close to the original input as possible.</p>
<p>The <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>O</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the OC error defined as:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>c</mml:mi><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:math></disp-formula></p>
<p>The data point <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>q</mml:mi></mml:math></inline-formula> is obtained by projecting the hidden representations of the training samples from the encoder into a high-dimensional feature space. The distance between <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>q</mml:mi></mml:math></inline-formula> and the center point <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>c</mml:mi></mml:math></inline-formula> is then calculated, with the objective of minimizing the size of the hypersphere centered at <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>c</mml:mi></mml:math></inline-formula> with <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>q</mml:mi></mml:math></inline-formula> as the radius. The center point is determined using a Gaussian mixture model, which can effectively model complex data distributions and handle noise and outliers more effectively.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the weight function based on the Gaussian distribution, specifically defined as:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>c</mml:mi><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the standard deviation of the Gaussian distribution, controlling the degree of fuzziness, and in the testing phase, the classification of the time series <italic>T</italic> as anomalous is based on the calculated anomaly score <italic>S</italic>.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>anomaly</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>S</mml:mi><mml:mo>&gt;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>normal</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>S</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> is the predefined classification threshold.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Robust Anomaly Detection</title>
<p>In Principal Component Analysis (PCA), a given matrix <italic>M</italic> can be approximated by identifying a low-rank matrix. To obtain a low-rank representation, PCA applies Singular Value Decomposition (SVD), which makes it inherently sensitive to outliers. To enhance robustness in the presence of outliers, Robust Principal Component Analysis (RPCA) has been proposed. RPCA aims to break down the original matrix <italic>M</italic> into two parts: a low-rank matrix <italic>L</italic> that represents the underlying clean structure of <italic>M</italic> and a sparse matrix <italic>S</italic> that contains the elements identified as anomalies.</p>
<p>Inspired by the approach of RPCA, we can separate the anomalous parts from the input time series and focus solely on learning the benign features from the samples. In this context, the clean time series <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> encapsulates the trends and periodic patterns present in the time series data, whereas <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> identifies the anomalous characteristics, which largely include random fluctuations that do not conform to established patterns. By eliminating this component&#x2019;s influence on the hidden representations, we can more accurately learn the information from clean samples and better differentiate them from anomalous samples. The objective function is as follows:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>0</mml:mn></mml:msub><mml:mspace width="1em" /><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> represents the encoder part, while <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> represents the decoder part. <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> are parameters used to control the balance between the sparsity of <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>. From our analysis, we observe that <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> plays a crucial role in separating the anomalous values in the time series. Specifically, when <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> is small, the objective function encourages more data to be classified as anomalous and separated from the original data. Conversely, when <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> is large, most of the data is retained, with only a small portion being isolated.</p>
<p>In <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-9">(9)</xref>, the loss functions include an <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>l</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> norm term to optimize the sparsity of anomalous values while ensuring their semantics. However, the <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>l</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> norm is non-convex, making optimization challenging. According to the [<xref ref-type="bibr" rid="ref-21">21</xref>], transforming the <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>l</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> norm to the <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>l</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> norm can provide a good approximation of the <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>l</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> norm. The formula is as follows:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>1</mml:mn></mml:msub><mml:mspace width="1em" /><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Algorithm</title>
<p>The optimization problems of ROC have constraints and thus cannot be solved by gradient descent based back-propagation (BACKPROP). The optimization task may instead be reformulated into two segments and approached using the Alternating Direction Method of Multipliers (ADMM). ADMM fundamentally works by breaking down the main objective into several sub-objectives, enabling the iterative optimization of each sub-objective while holding the remaining ones constant. Upon optimizing a given sub-objective, the method applies constraints to ensure consistency with the overall objective [<xref ref-type="bibr" rid="ref-22">22</xref>]. Furthermore, the Proximal Algorithm [<xref ref-type="bibr" rid="ref-23">23</xref>] is employed to address elements involving the <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>l</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> norm.</p>
<p>As shown in Algorithm 1, when optimizing the ROC, the process is as follows: first, optimize the integrated autoencoder part by minimizing <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>O</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>; then minimize <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula>; lastly, update <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> to maintain the constraint and provide the result as input for the subsequent iteration. The optimization process concludes based on two criteria: first, when <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> holds, and second, when both <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> remain constant, indicating that neither <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> nor <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>T</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula> is further changing, signifying that the anomalous values in <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> have stabilized at an optimal state.</p>
<fig id="fig-7">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-7.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<p>In terms of datasets, the AIOps dataset is a collection of 29 sub-datasets designed for detecting anomalies in web services based on business cloud KPIs. It includes 29 KPI time series collected from several large technology companies (such as Alibaba, Sogou, Tencent, Baidu, and eBay). These time series are sampled at 1 or 5-min intervals and divided into training and testing portions.</p>
<p>Another dataset we use the UCR Time Series Anomaly Archive, a recently launched repository containing 250 different time series datasets specifically for time series anomaly detection research. Each dataset contains anomalous events of varying lengths, ranging from 1 to 1700. Furthermore, these datasets cover various fields such as health, industry, and biology, exhibiting different types of anomalies with specific characteristics [<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
<p>We also use the multivariate time series dataset SMAP, which comes from a real-world expert-labeled dataset provided by NASA. Each dataset includes a training set and a testing set, with anomalies labeled in the testing set. It consists of data from 27 entities, each monitored by 55 metrics (variables). In all datasets, both point anomalies and collective anomalies are present, and true anomaly labels are available. Moreover, all methods are trained using time series data that contains anomalies, as the datasets do not provide clean time series without anomalies for training purposes. This configuration enables an investigation into the robustness of various algorithms when confronted with anomalies.</p>
<p><xref ref-type="table" rid="table-1">Table 1</xref> provides a systematic comparison of the key characteristics of the three datasets (AIOps, UCR, and SMAP) used in this study. It outlines the configurations of the sliding window parameters (window size and time step), the total number of samples, the data splits across training, validation, and testing sets, as well as the anomaly proportions in the training and testing datasets. Notably, the anomaly proportions vary significantly across datasets, with AIOps containing a small proportion of anomalies in both training and testing sets, UCR having no anomalies in the training set and a low proportion in the testing set, and SMAP featuring no anomalies in the training set but a relatively higher anomaly proportion in the testing set. This highlights the diverse nature of the datasets and their suitability for different anomaly detection tasks.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Details of dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th>AIOps</th>
<th>UCR</th>
<th>SMAP</th>
</tr>
</thead>
<tbody>
<tr>
<td>Window size</td>
<td>16</td>
<td>64</td>
<td>64</td>
</tr>
<tr>
<td>Time step</td>
<td>2</td>
<td>4</td>
<td>2</td>
</tr>
<tr>
<td>Total sample</td>
<td>2,961,039</td>
<td>4,830,858</td>
<td>281,400</td>
</tr>
<tr>
<td>Training/validation/testing</td>
<td>40%/10%/50%</td>
<td>24%/6%/70%</td>
<td>32%/8%/60%</td>
</tr>
<tr>
<td>Training/testing anomaly</td>
<td>2.98%/1.92%</td>
<td>0%/0.71%</td>
<td>0%/12.79%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Regarding baseline methods, we select two shallow machine learning methods, including OC-SVM [<xref ref-type="bibr" rid="ref-25">25</xref>] and Random Cut Forest (RCF) [<xref ref-type="bibr" rid="ref-26">26</xref>]. In deep learning algorithms, we compare our approach with five algorithms, including the deep one-class method SVDD [<xref ref-type="bibr" rid="ref-12">12</xref>], context-based anomaly detection for time series (TS-TCC) [<xref ref-type="bibr" rid="ref-27">27</xref>], and Ensemble Detection Method AOC [<xref ref-type="bibr" rid="ref-28">28</xref>]. Lastly, we chose two variants of the ROC method for ablation experiments: NoOC, which sets <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> to 0, representing a single hypothesis anomaly detection method; and NoRPCA, which directly uses the raw training data without any processing.</p>
<p>In our experiments, we primarily employed <bold>PA</bold>, <bold>PW</bold>, and <bold>Affiliation (precision recall and F1-score)</bold> [<xref ref-type="bibr" rid="ref-29">29</xref>] as evaluation metrics, as they align well with the unique requirements of time series anomaly detection tasks.</p>
<p><bold>PA</bold> measures the ratio of correctly classified points (both normal and anomalous) to the total number of points in the time series, offering a global perspective on the model&#x2019;s overall performance. Its formula is as follows:<disp-formula id="ueqn-12"><mml:math id="mml-ueqn-12" display="block"><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>Number of Correctly Classified Points</mml:mtext></mml:mrow><mml:mrow><mml:mtext>Total Number of Points</mml:mtext></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>PA is particularly suitable for scenarios where the primary goal is to assess the model&#x2019;s general classification accuracy across both normal and anomalous data. However, it may have limitations in datasets where normal points significantly outnumber anomalies, as the metric can dilute the model&#x2019;s anomaly detection performance by emphasizing overall accuracy.</p>
<p><bold>PW</bold> on the other hand, is designed to focus specifically on anomaly detection by emphasizing the precision and recall of the model when identifying anomalies. It provides a more refined measure of the model&#x2019;s effectiveness in distinguishing anomalous data from normal data [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. The formula for PW Precision is:<disp-formula id="ueqn-13"><mml:math id="mml-ueqn-13" display="block"><mml:mi>P</mml:mi><mml:mi>W</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>True Positives</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>True Positives</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>False Negatives</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>PW is particularly well-suited for time series anomaly detection tasks where the primary focus is on ensuring the accurate identification of anomalous samples. This metric is valuable in applications such as fault detection in industrial systems, where missing an anomaly (false negative) or misclassifying a normal event (false positive) can lead to significant consequences.</p>
<p>The choice of PA and PW as evaluation metrics reflects their ability to complement each other in time series anomaly detection scenarios. PA offers a holistic view of the model&#x2019;s classification accuracy, while PW ensures the model&#x2019;s stability and effectiveness in specifically detecting anomalies. This dual perspective allows for a balanced evaluation of the model&#x2019;s performance in time series tasks, particularly when addressing real-world applications with imbalanced data distributions or critical anomaly detection requirements.</p>
<p>By integrating these metrics, we can better assess the trade-offs between general accuracy and the precision-recall balance in anomaly detection, ensuring the model&#x2019;s applicability across diverse time series datasets and real-world tasks.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Implementation Details</title>
<p>In the ROC framework, we utilize two identical three-layer LSTMs (with a dropout rate of 0.45) as Seq2Seq autoencoders. The Adam optimizer was employed with a learning rate of <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, weight decay of <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.9</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.99</mml:mn></mml:math></inline-formula>, and <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>&#x03F5;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. All methods were implemented using Python 3.9, with PyTorch 1.7 for all neural network-based approaches. Additionally, Sklearn 0.24 was used for OCSVM, while Numpy 1.19 was used for EMA, SSA, and MP. Finally, Statsmodels 0.13 was utilized for STL. All experiments were conducted on a Linux workstation equipped with an Intel 32-core CPU, 256 GB RAM, and a single NVIDIA GeForce RTX 3090 GPU.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Results</title>
<p>In our study, we presented the prediction accuracy of affiliation, along with the corresponding point accuracy (PA) and range prediction (PW) scores. The results indicate that our method performs well across multiple datasets. Although it did not achieve the best results on the UCR dataset, it maintained a strong competitive edge relative to other methods. It is noteworthy that RCF and LSTM exhibited excellent performance on the UCR dataset, but their accuracy significantly declined on the AIOps dataset. This phenomenon can be attributed to the fact that the UCR dataset typically contains only a single anomaly segment and lacks anomalous samples in the training set. In contrast, the AIOps dataset features multiple anomaly segments, and the training set includes some anomalous characteristics, leading to higher false negative rates for methods that lack robustness.</p>
<p>To enhance the robustness of the model, the AOC method employs a soft-boundary strategy, which effectively improves the model&#x2019;s adaptability to anomalies. In contrast, our approach integrates the RPCA method to successfully filter out a significant portion of anomalous features in the training set. This approach also yielded favorable results on the multivariate time series dataset SMAP, further validating the broad applicability of our method.</p>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, in comparisons with various baseline models, we draw the following conclusions: RCF, as a shallow model, demonstrated outstanding performance on the UCR dataset, even surpassing some deep learning models. Meanwhile, two-stage anomaly detection methods, including SVDD and TS-TCC, did not achieve ideal results in time series anomaly detection, revealing the limitations of staged approaches for time series data and thereby constraining the performance of deep models. Additionally, our proposed ROC method performed well across all three datasets, confirming the effectiveness of the ensemble approach and its robustness against contaminated training sets.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Results summary</title>
</caption>
<table>
<colgroup>
<col/>
<col width="15mm"/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Metric</th>
<th>SVM</th>
<th>RCF</th>
<th>LSTM</th>
<th>DAGMM</th>
<th>SVDD</th>
<th>TS-TCC</th>
<th>AOC</th>
<th>ROC</th>
<th>NoRPCA</th>
<th>NoOC</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td>Precision</td>
<td>45.8</td>
<td>52.6</td>
<td>52.2</td>
<td>44.7</td>
<td>47.3</td>
<td>50.5</td>
<td>90.3</td>
<td><bold>96.2</bold></td>
<td>94.3</td>
<td>92.2</td>
</tr>
<tr>
<td></td>
<td>Recall</td>
<td>17.5</td>
<td>25.7</td>
<td>25.3</td>
<td>30.4</td>
<td>32.1</td>
<td>23.4</td>
<td><bold>38.6</bold></td>
<td>36.7</td>
<td>35.1</td>
<td>34.8</td>
</tr>
<tr>
<td>AIOps</td>
<td>F1-score</td>
<td>25.4</td>
<td>34.5</td>
<td>34.1</td>
<td>36.2</td>
<td>38.2</td>
<td>31.9</td>
<td><bold>54.0</bold></td>
<td>53.1</td>
<td>51.2</td>
<td>50.5</td>
</tr>
<tr>
<td></td>
<td>PA</td>
<td>53.4</td>
<td>53.2</td>
<td>76.1</td>
<td>14.2</td>
<td>14.3</td>
<td>17.5</td>
<td>80.1</td>
<td><bold>86.3</bold></td>
<td>81.6</td>
<td>62.3</td>
</tr>
<tr>
<td></td>
<td>PW</td>
<td>8.7</td>
<td>16.6</td>
<td>6.4</td>
<td>5.8</td>
<td>6.4</td>
<td>13.4</td>
<td>45.5</td>
<td><bold>47.2</bold></td>
<td>45.3</td>
<td>45.1</td>
</tr>
<tr>
<td></td>
<td>Precision</td>
<td>47.6</td>
<td>59.1</td>
<td><bold>67.8</bold></td>
<td>51.2</td>
<td>37.2</td>
<td>44.3</td>
<td>61.7</td>
<td>65.3</td>
<td>50.3</td>
<td>51.2</td>
</tr>
<tr>
<td></td>
<td>Recall</td>
<td><bold>84.0</bold></td>
<td>57.7</td>
<td>66.0</td>
<td>96.7</td>
<td>37.1</td>
<td>44.3</td>
<td>61.5</td>
<td>63.3</td>
<td>60.3</td>
<td>61.8</td>
</tr>
<tr>
<td>UCR</td>
<td>F1-score</td>
<td>60.3</td>
<td>58.4</td>
<td><bold>66.9</bold></td>
<td>66.9</td>
<td>37.2</td>
<td>44.3</td>
<td>61.6</td>
<td>64.1</td>
<td>54.8</td>
<td>56.0</td>
</tr>
<tr>
<td></td>
<td>PA</td>
<td>10.1</td>
<td><bold>98.2</bold></td>
<td>97.8</td>
<td>11.3</td>
<td>78.2</td>
<td>62.3</td>
<td>62.6</td>
<td>94.9</td>
<td>63.2</td>
<td>94.8</td>
</tr>
<tr>
<td></td>
<td>PW</td>
<td>2.02</td>
<td>35.3</td>
<td><bold>43.2</bold></td>
<td>6.6</td>
<td>8.3</td>
<td>14.2</td>
<td>15.3</td>
<td>22.6</td>
<td>17.6</td>
<td>16.8</td>
</tr>
<tr>
<td></td>
<td>Precision</td>
<td>43.3</td>
<td>42.2</td>
<td>84.3</td>
<td>40.6</td>
<td>51.9</td>
<td>45.4</td>
<td>91.3</td>
<td><bold>95.6</bold></td>
<td>93.7</td>
<td>90.6</td>
</tr>
<tr>
<td></td>
<td>Recall</td>
<td>34.2</td>
<td><bold>52.3</bold></td>
<td>24.3</td>
<td>15.1</td>
<td>46.9</td>
<td>17.2</td>
<td>36.3</td>
<td>41.7</td>
<td>38.6</td>
<td>39.1</td>
</tr>
<tr>
<td>SMAP</td>
<td>F1-score</td>
<td>38.2</td>
<td>46.7</td>
<td>37.7</td>
<td>22.0</td>
<td>49.3</td>
<td>24.9</td>
<td>51.9</td>
<td><bold>58.1</bold></td>
<td>54.7</td>
<td>54.6</td>
</tr>
<tr>
<td></td>
<td>PA</td>
<td>97.1</td>
<td>90.2</td>
<td><bold>98.5</bold></td>
<td>86.0</td>
<td>86.6</td>
<td>94.4</td>
<td>86.1</td>
<td>90.2</td>
<td>88.0</td>
<td>87.6</td>
</tr>
<tr>
<td></td>
<td>PW</td>
<td>14.1</td>
<td>7.4</td>
<td><bold>49.0</bold></td>
<td>7.0</td>
<td>13.1</td>
<td>12.2</td>
<td>37.2</td>
<td>40.9</td>
<td>37.6</td>
<td>35.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>On some datasets, although our methods are not the best, they do not fall behind much. The reason why our method does not achieve the best performance on the UCR dataset may lie in the fact that many time series in the UCR dataset exhibit strong contextual dependencies [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. For instance, in motion sensor data, transitions between different actions, or in weather data, long-term trends and seasonal patterns play a significant role. Models like LSTM, which excel at handling strongly time-dependent sequential data, can effectively capture critical patterns through learning temporal transitions between states. This capability allows LSTM to achieve higher accuracy in tasks such as behavior classification and anomaly detection. As a result, methods like LSTM are more suitable for datasets with strong temporal dependencies, such as UCR.</p>
<p>However, for datasets like AIOps, which encompass rich operational data and diverse task scenarios, our method demonstrates superior performance. This is due to its ability to handle large-scale data with highly diverse anomaly samples and to tackle more complex tasks. In such cases, our method significantly outperforms LSTM and other approaches that rely solely on temporal dependencies.</p>
<p>Finally, the results from NoOC and NoRPCA indicate that combining multiple normality assumptions with anomaly filtering models significantly enhances anomaly detection (AD) performance, this further validates the efficacy and importance of the different elements within our model.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Hyper-Parameter Analysis</title>
<p>In this section, we perform a hyperparameter analysis on the AIOps dataset, with a specific focus on examining two key parameters: <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> in the equations. <xref ref-type="fig" rid="fig-3">Fig. 3a</xref> illustrates the results of varying <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> for RPCA, showing that the model performs best when <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>. We hypothesize that the underlying reason for the observed performance may be that when the score from the one-class method constitutes a larger proportion of the anomaly score and exceeds the threshold value of 0.01, the model&#x2019;s performance tends to approximate that of shallow methods such as SVM. <xref ref-type="fig" rid="fig-3">Fig. 3b</xref> demonstrates the impact of <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> on overall performance, with the model achieving optimal results when <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>, identifying the best threshold for filtering anomaly features from the training set. The <italic>y</italic>-axis in the figure represents PA and PW precision metric.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>We conducted a hyperparameter analysis on AIOps for <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> (<bold>a</bold>) and <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> (<bold>b</bold>), focusing on the impact of anomaly filtering on the model&#x2019;s precision</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-3.tif"/>
</fig>
<p>We also conducted detailed experiments to investigate the reasons behind the performance decline associated with varying <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. As a parameter that adjusts the sparsity in <italic>S</italic>, <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> plays a crucial role in our analysis. Specifically, a smaller <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> encourages a large amount of data to be isolated as noise or anomalies in <italic>S</italic>, which minimizes the reconstruction error of the autoencoder; however, this can severely distort the original time series, resulting in inadequate anomaly detection due to low anomaly scores. Conversely, a larger <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> prevents data from being classified as noise or anomalies, leading to increased reconstruction errors. As shown in the <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, when <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> is too large, the reconstruction error rises, and since only a few anomalies are isolated, the results resemble those prior to RPCA, thus losing the filtering effect and causing a decline in performance. The <italic>y</italic>-axis in the figure represents the model&#x2019;s anomaly scores across different batches.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The figure illustrates how the anomaly scores change with variations in <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, where the green curve represents the scenario in which <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> is at its optimal value. This configuration more accurately reflects the anomaly scores of the time series, providing a clearer indication of the actual anomalies present</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Optimization Algorithm Analysis</title>
<p>We employ the <bold>gradient descent method</bold> as the core optimization algorithm to minimize the objective function. Gradient descent iteratively updates the parameters by computing the gradient of the loss function with respect to the model&#x2019;s parameters, ensuring a systematic approach to minimizing loss.</p>
<p>In addition, we incorporate a <bold>dynamically adjusted learning rate</bold> during the optimization process. The dynamic adjustment of the learning rate allows the algorithm to take larger steps when far from the optimal solution to accelerate convergence, while automatically reducing the step size as it approaches the optimal solution. This mechanism helps to avoid overshooting the minimum and improves stability near the global optimum. More importantly, the dynamically adjusted learning rate mitigates the risk of the algorithm getting trapped in local minima, a common issue in non-convex optimization problems.</p>
<p>By comparing the effects of dynamic and fixed learning rates, <xref ref-type="fig" rid="fig-5">Fig. 5</xref> provides a visual representation of how the convergence behavior differs under different learning rate strategies. The dynamic learning rate strategy demonstrates faster convergence and better adaptability to the optimization landscape, particularly in complex scenarios.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Figure demonstrates the convergence process of the optimization algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-5.tif"/>
</fig>
<p>In addition, we visualized the actual convergence process of the model, demonstrating the step-by-step reduction in the loss function values during optimization. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> not only provides an intuitive comparison between the performance of dynamic learning rate adjustment and fixed learning rate strategies but also strongly supports the effectiveness of the dynamic learning rate. Specifically, the dynamic learning rate facilitates a faster reduction in loss values and exhibits greater stability as it approaches the global optimum. This indicates that the dynamic learning rate has significant advantages in optimizing non-convex problems and handling complex objective functions. Furthermore, it validates the applicability of this approach in addressing challenging tasks.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The optimization process with dynamic learning rate and fixed learning rate as a function of epochs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60564-fig-6.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper introduces a robust time series anomaly detection method, <bold>ROC</bold>, which is grounded in multiple hypotheses and eliminates the need for pre-training. The proposed method projects the hidden representation layer of the autoencoder and integrates the objectives of both the autoencoder and one-class (OC) methods. By filtering out anomalous segments of the input time series, <bold>ROC</bold> avoids the contamination of the compression layer by anomalous features during training and resolves potential inconsistencies between the two hypotheses. This approach effectively captures normal patterns from multiple perspectives, allowing the model to learn a more comprehensive representation of typical time series data. As a result, the method demonstrates an enhanced ability to detect diverse types of anomalies. Experimental evaluations on three real-world datasets validate the superior performance of the proposed approach.</p>
<p>In future work, we plan to further enhance the method&#x2019;s robustness against adversarial attacks in time series anomaly detection. Drawing inspiration from state-of-the-art techniques, we aim to explore feature learning from various forms of time series representations, such as residuals and frequency domains. Additionally, we intend to combine these advanced feature extraction techniques with our robust approach to filter anomalous features from multiple perspectives, ultimately improving the model&#x2019;s effectiveness and adaptability in complex scenarios.</p>
</sec>
</body>
<back>
<ack>
<p>The authors appreciate the reviewers and editors for their valuable feedback on this work. We also acknowledge the providers of datasets, including AIOps and UCR.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the National Natural Science Foundation (62202118), Guizhou Province Major Project (Qiankehe Major Project [2024]014), Science and Scientific and Technological Research Projects from Guizhou Education Department (Qianiao ji [2023]003), Hundred-level Innovative Talent Project of Guizhou Provincial Science and Technology Department (Qiankehe Platform Talent-GCC[2023]018) and Guizhou Province Major Project (Qiankehe Major Project [2024]003), Foundation of Chongqing Key Laboratory of Public Big Data Security Technology (CQKL-QJ202300001).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: method proposal and implementation, experimental proof, and manuscript writing: Zhengdao Yang; experimental setting and grant support: Yuling Chen, Xuewei Wang; manuscript revision: Hui Dou, Haiwei Sang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data used to support the findings of this study are available from the corresponding author upon request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Nomenclature</title>
<def-list>
<def-item>
<term>AE</term>
<def>
<p>Autoencoder</p>
</def>
</def-item>
<def-item>
<term>AD</term>
<def>
<p>Anomaly Detection</p>
</def>
</def-item>
<def-item>
<term>ADMM</term>
<def>
<p>Alternating Direction Method of Multipliers</p>
</def>
</def-item>
<def-item>
<term>AIOps</term>
<def>
<p>Artificial Intelligence for IT (Information Technology) Operations Performance Score</p>
</def>
</def-item>
<def-item>
<term>AOC</term>
<def>
<p>deep Autoencoding One Class</p>
</def>
</def-item>
<def-item>
<term>BACKPROP</term>
<def>
<p>Backpropagation</p>
</def>
</def-item>
<def-item>
<term>DDoS</term>
<def>
<p>Distributed Denial of Service</p>
</def>
</def-item>
<def-item>
<term>DAGMM</term>
<def>
<p>Deep Autoencoding Gaussian Mixture Model</p>
</def>
</def-item>
<def-item>
<term>EMA</term>
<def>
<p>Exponential Moving Average</p>
</def>
</def-item>
<def-item>
<term>ECG</term>
<def>
<p>Electrocardiogram</p>
</def>
</def-item>
<def-item>
<term>EEG</term>
<def>
<p>Electroencephalogram</p>
</def>
</def-item>
<def-item>
<term>GAN</term>
<def>
<p>Generative Adversarial Network</p>
</def>
</def-item>
<def-item>
<term>KPI</term>
<def>
<p>Key Performance Indicator</p>
</def>
</def-item>
<def-item>
<term>LSTM</term>
<def>
<p>Long Short-Term Memory</p>
</def>
</def-item>
<def-item>
<term>MP</term>
<def>
<p>Matrix Profile</p>
</def>
</def-item>
<def-item>
<term>OC</term>
<def>
<p>One-Class Classification</p>
</def>
</def-item>
<def-item>
<term>OS</term>
<def>
<p>Outlier Score</p>
</def>
</def-item>
<def-item>
<term>PCA</term>
<def>
<p>Principal Component Analysis</p>
</def>
</def-item>
<def-item>
<term>PA</term>
<def>
<p>Point-adjusted metrics</p>
</def>
</def-item>
<def-item>
<term>PW</term>
<def>
<p>Point-wise metrics</p>
</def>
</def-item>
<def-item>
<term>RCF</term>
<def>
<p>Random Cut Forest</p>
</def>
</def-item>
<def-item>
<term>RPCA</term>
<def>
<p>Robust Principal Component Analysis</p>
</def>
</def-item>
<def-item>
<term>ROC</term>
<def>
<p>Robust One-Class Classification Detection</p>
</def>
</def-item>
<def-item>
<term>SVD</term>
<def>
<p>Singular Value Decomposition</p>
</def>
</def-item>
<def-item>
<term>SVM</term>
<def>
<p>Support Vector Machine</p>
</def>
</def-item>
<def-item>
<term>SSA</term>
<def>
<p>Singular Spectrum Analysis</p>
</def>
</def-item>
<def-item>
<term>STL</term>
<def>
<p>Seasonal and Trend decomposition using Loess</p>
</def>
</def-item>
<def-item>
<term>SVDD</term>
<def>
<p>Support Vector Data Description</p>
</def>
</def-item>
<def-item>
<term>TSAD</term>
<def>
<p>Time Series Anomaly Detection</p>
</def>
</def-item>
<def-item>
<term>TS-TCC</term>
<def>
<p>Time-Series representation learning via Temporal and Contextual Contrasting</p>
</def>
</def-item>
<def-item>
<term>UCR</term>
<def>
<p>University of California, Riverside Time Series Anomaly Archive</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grubbs</surname> <given-names>FE</given-names></string-name></person-group>. <article-title>Procedures for detecting outlying observations in samples</article-title>. <source>Technometrics</source>. <year>1969</year>;<volume>11</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1080/00401706.1969.10490657</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hengel</surname> <given-names>AVD</given-names></string-name></person-group>. <article-title>Deep learning for anomaly detection: a review</article-title>. <source>ACM Comput Surv</source>. <year>2021</year>;<volume>54</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3439950</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>DRL: dynamic rebalance learning for adversarial robustness of UAV with long-tailed distribution</article-title>. <source>Comput Commun</source>. <year>2023</year>;<volume>205</volume>(<issue>6</issue>):<fpage>14</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.comcom.2023.04.002</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>RPU-PVB: robust object detection based on a unified metric perspective with bilinear interpolation</article-title>. <source>J Cloud Comput</source>. <year>2023</year>;<volume>12</volume>(<issue>1</issue>):<fpage>169</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s13677-023-00534-3</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>X</given-names></string-name></person-group>. <article-title>BCEAD: a blockchain-empowered ensemble anomaly detection for wireless sensor network via isolation forest</article-title>. <source>Secur Commun Netw</source>. <year>2021</year>;<volume>2021</volume>(<issue>1</issue>):<fpage>9430132</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/9430132</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>RR</given-names></string-name></person-group>. <article-title>Reducing the dimensionality of data with neural networks</article-title>. <source>Science</source>. <year>2006</year>;<volume>313</volume>(<issue>5786</issue>):<fpage>504</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1126/science.1127647</pub-id>; <pub-id pub-id-type="pmid">16873662</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tishby</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pereira</surname> <given-names>FC</given-names></string-name>, <string-name><surname>Bialek</surname> <given-names>W</given-names></string-name></person-group>. <article-title>The information bottleneck method</article-title>. <comment>arXiv:physics/0004057. 2000</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.physics/0004057</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Maqsood</surname> <given-names>M</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rho</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Advancing autoencoder architectures for enhanced anomaly detection in multivariate industrial time series</article-title>. <source>Comput Mater Contin</source>. <year>2024</year>;<volume>81</volume>(<issue>1</issue>):<fpage>1275</fpage>&#x2013;<lpage>300</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2024.054826</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Prasad</surname> <given-names>P</given-names></string-name>, <string-name><surname>Sayeed</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Enhancing internet of things intrusion detection using artificial intelligence</article-title>. <source>Comput, Mater Contin</source>. <year>2024</year>;<volume>81</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2024.053861</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>PUNet: a semi-supervised anomaly detection model for network anomaly detection based on positive unlabeled data</article-title>. <source>Comput Mater Contin</source>. <year>2024</year>;<volume>81</volume>(<issue>1</issue>):<fpage>327</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2024.054558</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bl&#x00E1;zquez-Garc&#x00ED;a</surname> <given-names>A</given-names></string-name>, <string-name><surname>Conde</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mori</surname> <given-names>U</given-names></string-name>, <string-name><surname>Lozano</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>A review on outlier/anomaly detection in time series data</article-title>. <source>ACM Comput Surv</source>. <year>2021</year>;<volume>54</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3444690</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ruff</surname> <given-names>L</given-names></string-name>, <string-name><surname>Vandermeulen</surname> <given-names>R</given-names></string-name>, <string-name><surname>Goernitz</surname> <given-names>N</given-names></string-name>, <string-name><surname>Deecke</surname> <given-names>L</given-names></string-name>, <string-name><surname>Siddiqui</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Binder</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep one-class classification</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2018</year>; <publisher-name>PMLR</publisher-name>. p. <fpage>4393</fpage>&#x2013;<lpage>402</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sohn</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>CL</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Pfister</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Learning and evaluating representations for deep one-class classification</article-title>. <year>arXiv:2011.02578. 2020</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2011.02578</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>When deep learning-based soft sensors encounter reliability challenges: a practical knowledge-guided adversarial attack and its defense</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2023</year>;<volume>20</volume>(<issue>2</issue>):<fpage>2702</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2023.3297663</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Adversarial robustness enhancement for deep learning-based soft sensors: an adversarial training strategy using historical gradients and domain adaptation</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>12</issue>):<fpage>3909</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24123909</pub-id>; <pub-id pub-id-type="pmid">38931693</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Schlegl</surname> <given-names>T</given-names></string-name>, <string-name><surname>Seeb&#x00F6;ck</surname> <given-names>P</given-names></string-name>, <string-name><surname>Waldstein</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Schmidt-Erfurth</surname> <given-names>U</given-names></string-name>, <string-name><surname>Langs</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Unsupervised anomaly detection with generative adversarial networks to guide marker discovery</article-title>. In: <conf-name>International Conference on Information Processing in Medical Imaging</conf-name>; <year>2017</year>; <publisher-name>Springer</publisher-name>. p. <fpage>146</fpage>&#x2013;<lpage>57</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Malhotra</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramakrishnan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Anand</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vig</surname> <given-names>L</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Shroff</surname> <given-names>G</given-names></string-name></person-group>. <article-title>LSTM-based encoder-decoder for multi-sensor anomaly detection</article-title>. <comment>arXiv:1607.00148. 2016</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1607.00148</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Min</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lumezanu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep autoencoding gaussian mixture model for unsupervised anomaly detection</article-title>. In: <conf-name>International Conference on Learning Representations</conf-name>; <year>2018</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Qiu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Pfrommer</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kloft</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mandt</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rudolph</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Neural transformation learning for deep anomaly detection beyond images</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2021</year>; <publisher-name>PMLR</publisher-name>. p. <fpage>8703</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cand&#x00E8;s</surname> <given-names>EJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wright</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Robust principal component analysis?</article-title> <source>J ACM</source>. <year>2011</year>;<volume>58</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/1970392.197039</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Donoho</surname> <given-names>DL</given-names></string-name></person-group>. <article-title>For most large underdetermined systems of linear equations the minimal l1-norm solution is also the sparsest solution</article-title>. <source>Commun Pure Appl Math: A J Issued Courant Institute Math Sci</source>. <year>2006</year>;<volume>59</volume>(<issue>6</issue>):<fpage>797</fpage>&#x2013;<lpage>829</lpage>. doi:<pub-id pub-id-type="doi">10.1002/cpa.20132</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boyd</surname> <given-names>S</given-names></string-name>, <string-name><surname>Parikh</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>E</given-names></string-name>, <string-name><surname>Peleato</surname> <given-names>B</given-names></string-name>, <string-name><surname>Eckstein</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Distributed optimization and statistical learning via the alternating direction method of multipliers</article-title>. <source>Found Trends<sub>&#x00AE;</sub> Mach Learn</source>. <year>2011</year>;<volume>3</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>122</lpage>. doi:<pub-id pub-id-type="doi">10.1561/2200000016</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Parikh</surname> <given-names>N</given-names></string-name>, <string-name><surname>Boyd</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Proximal algorithms</article-title>. <source>Found Trends<sub>&#x00AE;</sub> Optim</source>. <year>2014</year>;<volume>1</volume>(<issue>3</issue>):<fpage>127</fpage>&#x2013;<lpage>239</lpage>. doi:<pub-id pub-id-type="doi">10.1561/2400000003</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dau</surname> <given-names>HA</given-names></string-name>, <string-name><surname>Bagnall</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kamgar</surname> <given-names>K</given-names></string-name>, <string-name><surname>Yeh</surname> <given-names>CCM</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gharghabi</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The UCR time series archive</article-title>. <source>IEEE/CAA J Autom Sin</source>. <year>2019</year>;<volume>6</volume>(<issue>6</issue>):<fpage>1293</fpage>&#x2013;<lpage>305</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JAS.2019.1911747</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sch&#x00F6;lkopf</surname> <given-names>B</given-names></string-name>, <string-name><surname>Williamson</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Smola</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shawe-Taylor</surname> <given-names>J</given-names></string-name>, <string-name><surname>Platt</surname> <given-names>J</given-names></string-name></person-group>. <chapter-title>Support vector method for novelty detection</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Solla</surname> <given-names>S</given-names></string-name>, <string-name><surname>Leen</surname> <given-names>T</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>K</given-names></string-name></person-group>, editors. <source>Advances in neural information processing systems</source>. <publisher-name>MIT Press</publisher-name>; <year>1999</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>N</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>G</given-names></string-name>, <string-name><surname>Schrijvers</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Robust random cut forest based anomaly detection on streams</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2016</year>; <publisher-name>PMLR</publisher-name>. p. <fpage>2712</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Eldele</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ragab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kwoh</surname> <given-names>CK</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Time-series representation learning via temporal and contextual contrasting</article-title>. <comment>arXiv:2106.14112. 2021</comment>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2021</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wo</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep autoencoding one-class time series anomaly detection</article-title>. In: <conf-name>ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>; <year>2023</year>, <publisher-name>IEEE</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huet</surname> <given-names>A</given-names></string-name>, <string-name><surname>Navarro</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Rossi</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Local evaluation of time series anomaly detection algorithms</article-title>. In: <conf-name>Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining</conf-name>; <year>2022</year>. p. <fpage>635</fpage>&#x2013;<lpage>45</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Schmidl</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wenig</surname> <given-names>P</given-names></string-name>, <string-name><surname>Papenbrock</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Anomaly detection in time series: a comprehensive evaluation</article-title>. <source>Proc VLDB Endow</source>. <year>2022</year>;<volume>15</volume>(<issue>9</issue>):<fpage>1779</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.14778/3538598.3538602</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Pei</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Robust anomaly detection for multivariate time series through stochastic recurrent neural network</article-title>. In: <conf-name>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x0026; Data Mining</conf-name>; <year>2019</year>. p. <fpage>2828</fpage>&#x2013;<lpage>37</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Keogh</surname> <given-names>EJ</given-names></string-name></person-group>. <article-title>Current time series anomaly detection benchmarks are flawed and are creating the illusion of progress</article-title>. <source>IEEE Trans Knowl Data Eng</source>. <year>2021</year>;<volume>35</volume>(<issue>3</issue>):<fpage>2421</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TKDE.2021.3112126</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lai</surname> <given-names>KH</given-names></string-name>, <string-name><surname>Zha</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Revisiting time series outlier detection: definitions and benchmarks</article-title>. In: <conf-name>Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)</conf-name>; <year>2021</year>.</mixed-citation></ref>
</ref-list>
</back></article>