<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">69097</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.069097</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Data Augmentation: A Multi-Perspective Survey on Data, Methods, and Applications</article-title>
<alt-title alt-title-type="left-running-head">Data Augmentation: A Multi-Perspective Survey on Data, Methods, and Applications</alt-title>
<alt-title alt-title-type="right-running-head">Data Augmentation: A Multi-Perspective Survey on Data, Methods, and Applications</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Cui</surname><given-names>Canlin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yao</surname><given-names>Junyu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>yaojunyu0205@163.com</email></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Xia</surname><given-names>Heng</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>hengxia@mail.tsinghua.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information Science and Technology, Beijing University of Technology</institution>, <addr-line>Beijing, 100124</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Automation, Tsinghua University</institution>, <addr-line>Beijing, 100084</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Junyu Yao. Email: <email>yaojunyu0205@163.com</email>; Heng Xia. Email: <email>hengxia@mail.tsinghua.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>10</month><year>2025</year>
</pub-date>
<volume>85</volume>
<issue>3</issue>
<fpage>4275</fpage>
<lpage>4306</lpage>
<history>
<date date-type="received">
<day>14</day>
<month>6</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_69097.pdf"></self-uri>
<abstract>
<p>High-quality data is essential for the success of data-driven learning tasks. The characteristics, precision, and completeness of the datasets critically determine the reliability, interpretability, and effectiveness of subsequent analyzes and applications, such as fault detection, predictive maintenance, and process optimization. However, for many industrial processes, obtaining sufficient high-quality data remains a significant challenge due to high costs, safety concerns, and practical constraints. To overcome these challenges, data augmentation has emerged as a rapidly growing research area, attracting considerable attention across both academia and industry. By expanding datasets, data augmentation techniques improve greater generalization and more robust performance in actual applications. This paper provides a comprehensive, multi-perspective review of data augmentation methods for industrial processes. For clarity and organization, existing studies are systematically grouped into four categories: small sample with low dimension, small sample with high dimension, large sample with low dimension, and large sample with high dimension. Within this framework, the review examines current research from both methodological and application-oriented perspectives, highlighting main methods, advantages, and limitations. By synthesizing these findings, this review offers a structured overview for scholars and practitioners, serving as a valuable reference for newcomers and experienced researchers seeking to explore and advance data augmentation techniques in industrial processes.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Data-driven</kwd>
<kwd>data augmentation</kwd>
<kwd>big data</kwd>
<kwd>industrial application</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Postdoctoral Fellowship Program (Grade B) of China</funding-source>
<award-id>GZB20250435</award-id>
</award-group>
<award-group id="awg2">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62403270</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Data-driven modeling aims to uncover latent patterns within the data, enabling reliable decision support for industrial operations [<xref ref-type="bibr" rid="ref-1">1</xref>]. The development of accurate and robust models depends on high-quality data. In most cases, industrial systems operate safely and stably under normal conditions, whereas abnormal or fault conditions occur infrequently [<xref ref-type="bibr" rid="ref-2">2</xref>]. This results in a long-tailed data distribution, which can bias the models toward majority classes and increase the risk of overfitting [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. Furthermore, due to technological and cost constraints, certain critical parameters are difficult to measure online, leading to a limited number of labeled samples for model training [<xref ref-type="bibr" rid="ref-5">5</xref>]. As a result, data-driven modeling for industrial processes often contends with the challenge of &#x2018;big data, small sample&#x2019;.</p>
<p>Data augmentation offers a promising strategy for expanding datasets and improving modeling performance. A keyword search for &#x2018;industrial&#x2019; and &#x2018;data augmentation&#x2019; in the Web of Science (WoS) database reveals the number of related publications and citations, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The upward trend in both metrics reflects a growing scholarly interest in this field. In particular, the application of data augmentation in industrial processes has increased significantly between 2019 and 2024.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Development trends of data augmentation for industrial processes (2006&#x2013;2025)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-1.tif"/>
</fig>
<p>Data augmentation aims to approximate the true data distribution by generating virtual samples. Over the past decade, numerous review articles have examined various data augmentation techniques. For example, reference [<xref ref-type="bibr" rid="ref-6">6</xref>] surveyed several data augmentation methods for text classification across data and feature spaces. Reference [<xref ref-type="bibr" rid="ref-7">7</xref>] provided a comprehensive analysis of data augmentation techniques in medical imaging. Reference [<xref ref-type="bibr" rid="ref-8">8</xref>] systematically reviewed mix-based data augmentation methods, focusing on multi-modal data such as images, text, and video. In the context of mechanical equipment, references [<xref ref-type="bibr" rid="ref-9">9</xref>] and [<xref ref-type="bibr" rid="ref-10">10</xref>] synthesized state-of-the-art data augmentation strategies for fault diagnosis and predictive maintenance. However, there remains a lack of comprehensive reviews that specifically address the unique challenges and applications of data augmentation in industrial processes. To fill this gap, this survey offers an in-depth review of data augmentation techniques, with a particular emphasis on their applications in industrial processes. The main contributions are as follows: 1) It highlights the impact of data characteristics on modeling performance and introduces a framework based on sample size and feature dimensionality. 2) It categorizes and analyzes recent studies, summarizing targeted solutions to key challenges. 3) It reviews the application of data augmentation methods in representative industrial domains, providing insight and guidance to researchers and engineers.</p>
<p>For this survey review, a thematic approach was employed, which is defined by four data topics (details are provided in <xref ref-type="sec" rid="s2">Section 2</xref>). Our objective was to cover both classical foundations and emerging trends in data augmentation strategies by utilizing the four data characteristics outlined in the survey. The review process involved a structured literature search in major academic databases, including WoS, Scopus, IEEE Xplore, and ScienceDirect. We identified relevant literature using keyword combinations such as &#x2018;small sample learning&#x2019;, &#x2018;industrial process&#x2019;, and &#x2018;data augmentation&#x2019;. Subsequently, each paper was tagged according to a four-quadrant framework based on its sample size and feature dimensionality, and then classified by methodological category. The final step was to synthesize key themes and practical industrial applications, providing readers with both breadth and depth of understanding.</p>
<p>This survey is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> defines the data characteristics and analyzes existing challenges from the perspective of sample size and feature dimensionality. <xref ref-type="sec" rid="s3">Section 3</xref> reviews recent studies on data augmentation and evaluates the strengths and limitations of various methods. <xref ref-type="sec" rid="s4">Section 4</xref> explores industrial applications of these techniques and compares their characteristics across different domains. <xref ref-type="sec" rid="s5">Section 5</xref> discusses current challenges and proposes potential directions for future research. Finally, <xref ref-type="sec" rid="s6">Section 6</xref> concludes the survey.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Data Characteristics and Problem Formulation</title>
<p>A typical dataset is denoted as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="bold-italic">X</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, with <italic>N</italic> samples and <italic>M</italic> feature dimension, as illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. A sample comprises a set of data collected by the system at a specific time, whereas a feature represents the recorded values of the system over a given period.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Diagram of the process data</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-2.tif"/>
</fig>
<p><bold>Remark:</bold> Determining the appropriate sample size and feature dimensionality is critical before starting modeling tasks [<xref ref-type="bibr" rid="ref-11">11</xref>]. Reference [<xref ref-type="bibr" rid="ref-12">12</xref>] introduced the probably approximately correct theory, which estimates the sample size required to ensure optimal model performance. This theory emphasizes that sufficient samples are an essential prerequisite for data-driven models. However, collecting sufficient samples in actual industrial processes is often challenging, particularly in a setting with small sample sizes and high-dimensional data [<xref ref-type="bibr" rid="ref-13">13</xref>]. For the small sample, reference [<xref ref-type="bibr" rid="ref-14">14</xref>] defined a training set with fewer than 100 samples as a small sample case. In [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>], the small sample problem was characterized by sample sizes smaller than 50 in engineering applications and smaller than 30 in academic research. In this survey, a sample size of <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> is defined as the demarcation point, where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>N</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> denotes a small sample case and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>N</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> denotes a large sample case. Based on this basis, reference [<xref ref-type="bibr" rid="ref-12">12</xref>] considered the ratio of sample size to feature dimension, defining the index as <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>M</mml:mi></mml:math></inline-formula>. Reference [<xref ref-type="bibr" rid="ref-18">18</xref>] indicated that <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> is a typical value. Therefore, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>M</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> represents a low-dimensional case, and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>M</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> represents a high-dimensional case in this survey.</p>
<p>Without loss of generality, data can be classified into four categories based on sample size and feature dimensionality: small sample with low dimensionality (SSLD), small sample with high dimensionality (SSHD), large sample with low dimensionality (LSLD), and large sample with high dimensionality (LSHD), as illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Coordinate diagram of data characteristics</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-3.tif"/>
</fig>
<p><bold>Case 1: Large sample with low dimension (LSLD)</bold></p>
<p>In Case 1, the large sample size and low-dimensional datasets are defined below:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mn>10</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> denotes the dataset that follows the LSLD characteristics, <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>i</mml:mi></mml:math></inline-formula>-th sample, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the sample size of <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the feature dimension of <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> denotes the evaluate index.</p>
<p>LSLD data are commonly encountered in the monitoring of industrial equipment, particularly rotating machinery. Abnormal or fault conditions typically occur infrequently and have a short duration, resulting in a low proportion of fault samples. As a result, the numerous redundant normal data and the scarcity of fault data hinder the development of accurate detection models. This imbalance often leads to class distribution problems [<xref ref-type="bibr" rid="ref-19">19</xref>]. A representative example is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, where minority fault classes introduce challenges such as decision boundary overlap and outliers.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Diagram of the class imbalance for LSLD</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-4.tif"/>
</fig>
<p><bold>Case 2: Small sample with high dimension (SSHD)</bold></p>
<p>In Case 2, the small sample size and high-dimensional datasets are defined below:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>10</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p>This scenario introduces the curse of dimensionality (CoD) [<xref ref-type="bibr" rid="ref-20">20</xref>], resulting in a sparse sample space, as illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. As dimensionality increases, sample sparsity grows exponentially. Moreover, high dimensionality often causes multicollinearity, further complicating model training. Therefore, SSHD modeling is highly susceptible to overfitting, making it difficult to achieve strong generalization performance.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The sample space is sparse due to the increase in dimension</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-5.tif"/>
</fig>
<p><bold>Case 3: Small sample with low dimension (SSLD)</bold></p>
<p>In Case 3, the small sample size and low-dimensional datasets are defined below:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mn>0.1</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>100</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p>SSLD modeling is highly susceptible to noise interference, especially due to sensor drift. The statistical confidence of the process data is low, and the available information is insufficient to support reliable real-time decision-making. SSLD data commonly occurs during early-stage device monitoring, such as in the initial phase of a process or during the debugging stage of new equipment, when the sample size and feature dimensionality are typically limited. This scenario often causes overfitting during the modeling process, as illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Diagram of the overfitting problem</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-6.tif"/>
</fig>
<p><bold>Case 4: Large sample with high dimension (LSHD)</bold></p>
<p>In Case 4, the large sample size and high-dimensional datasets are defined below:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">D</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p>LSHD modeling not only inherits the challenges present in SSLD, SSHD, and LSLD cases, but also imposes higher demands on data analysis and mining. LSHD data are commonly encountered in plant-level monitoring and multi-plant federated learning tasks. Industrial systems often integrate multi-source data, such as sensor readings, production logs, image inspections, and maintenance records, resulting in a high-dimensional, heterogeneous feature space. This complexity gives rise to challenges such as long-tailed data distributions in intelligent monitoring and maintenance, as illustrated in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Diagram of the class imbalance distribution problem</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-7.tif"/>
</fig>
</sec>
<sec id="s3">
<label>3</label>
<title>Methods</title>
	<p>Existing data augmentation methods can be classified into five categories based on their generation mechanisms, as illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>: transform-based, statistical-based, deep generative-based, transfer learning-based, and physical model-based methods. Statistical methods were first introduced in 2008. With the rise of deep learning, deep generative models became a key focus in data augmentation from 2018 onward. In 2019, transform-based methods emerged to address the small sample problem in time series datasets. Physical model-based approaches began to appear in 2021. Since 2022, transfer learning has been employed to extract domain knowledge and augment minority classes. Based on four data characteristics, the existing data augmentation techniques are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Classification of data augmentation methods</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-8.tif"/>
</fig><table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of data augmentation techniques based on four data categories</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Data characteristic</th>
<th>Data augmentation technique</th>
<th>Implementation method</th>
<th>Reference</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="11">LSLD</td>
<td rowspan="5">Transform technique</td>
<td>Noise injection method</td>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
</tr>
<tr>
<td>Geometric scaling method</td>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
</tr>
<tr>
<td>Zero-masking method</td>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
</tr>
<tr>
<td>Time-shift method</td>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
</tr>
<tr>
<td>Flip method</td>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
</tr>
<tr>
<td rowspan="3">Deep generative model</td>
<td>Variational autoencoder</td>
<td>[<xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
</tr>
<tr>
<td>Generative adversarial network</td>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>&#x2013;<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>Diffusion model</td>
<td>[<xref ref-type="bibr" rid="ref-41">41</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
</tr>
<tr>
<td>Transfer learning</td>
<td>&#x2014;</td>
<td>[<xref ref-type="bibr" rid="ref-45">45</xref>&#x2013;<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
</tr>
<tr>
<td rowspan="2">Physical model</td>
<td>Numerical simulations</td>
<td>[<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
</tr>
<tr>
<td>Digital twins</td>
<td>[<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-55">55</xref>]</td>
</tr>
<tr>
<td rowspan="6">SSHD</td>
<td rowspan="3">Statistical technique</td>
<td>Feature selection method</td>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>&#x2013;<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
</tr>
<tr>
<td>Feature extraction method</td>
<td>[<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-59">59</xref>&#x2013;<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
</tr>
<tr>
<td>Multi-model method</td>
<td>[<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
</tr>
<tr>
<td rowspan="2">Deep generative model</td>
<td>Variational autoencoder</td>
<td>[<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-66">66</xref>&#x2013;<xref ref-type="bibr" rid="ref-69">69</xref>]</td>
</tr>
<tr>
<td>Generative adversarial network</td>
<td>[<xref ref-type="bibr" rid="ref-70">70</xref>&#x2013;<xref ref-type="bibr" rid="ref-74">74</xref>]</td>
</tr>
<tr>
<td>Transfer learning</td>
<td>&#x2014;</td>
<td>[<xref ref-type="bibr" rid="ref-75">75</xref>&#x2013;<xref ref-type="bibr" rid="ref-77">77</xref>]</td>
</tr>
<tr>
<td rowspan="4">SSLD</td>
<td rowspan="2">Statistical technique</td>
<td>Distributional assumption method</td>
<td>[<xref ref-type="bibr" rid="ref-78">78</xref>&#x2013;<xref ref-type="bibr" rid="ref-82">82</xref>]</td>
</tr>
<tr>
<td>Interpolation method</td>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-83">83</xref>&#x2013;<xref ref-type="bibr" rid="ref-87">87</xref>]</td>
</tr>
<tr>
<td rowspan="2">Deep generative model</td>
<td>Classification model</td>
<td>[<xref ref-type="bibr" rid="ref-88">88</xref>&#x2013;<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
</tr>
<tr>
<td>Regression model</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>&#x2013;<xref ref-type="bibr" rid="ref-95">95</xref>]</td>
</tr>
<tr>
<td rowspan="3">LSHD</td>
<td rowspan="2">Deep generative model</td>
<td>Variational autoencoder</td>
<td>[<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
</tr>
<tr>
<td>Generative adversarial network</td>
<td>[<xref ref-type="bibr" rid="ref-97">97</xref>&#x2013;<xref ref-type="bibr" rid="ref-101">101</xref>]</td>
</tr>
<tr>
<td>Physical model</td>
<td>&#x2014;</td>
<td>[<xref ref-type="bibr" rid="ref-102">102</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3_1">
<label>3.1</label>
<title>Data Augmentation for LSLD</title>
<p>LSLD datasets often exhibit significant class imbalance, with minority class samples representing only a small portion of the overall dataset. This imbalance leads to models biased toward the majority class, whereas overlooking the minority class. To address this problem, various data augmentation methods have been proposed, which can be categorized into transform-based, deep generative model-based, transfer learning-based, and physical model-based approaches.</p>
<p><bold>(a) LSLD based on transform techniques</bold></p>
<p><bold>1)</bold> The noise injection method enhances data generalization by superimposing random noise on the original data. The core idea is to simulate the uncertainty inherent in actual scenarios. The noise injection process is described as follows:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">&#x03B7;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the augmented data after noise injection, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the original data, and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">&#x03B7;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the random noise. Gaussian noise is commonly used in existing research [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-104">104</xref>]. The noise intensity often requires dynamic adjustment based on physical constraints. To this end, references [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>] employed signal-to-noise ratios to regulate the noise level, as shown below:
<disp-formula id="eqn-6a"><label>(6a)</label><mml:math id="mml-eqn-6a" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mrow><mml:mrow><mml:mtext>SNR</mml:mtext></mml:mrow></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:mrow><mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x03C6;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-6b"><label>(6b)</label><mml:math id="mml-eqn-6b" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:msqrt><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>SNR</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">&#x03B7;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>noise</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mrow><mml:mrow><mml:mtext>{SNR}</mml:mtext></mml:mrow></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the signal-to-noise ratio, and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> denotes the noise level.</p>
<p><bold>2)</bold> The geometric scaling method generates new samples by adjusting the amplitude or time scale of the original data. This method aims to simulate variations in equipment operating conditions through techniques such as amplitude scaling and time scaling.</p>
<p>&#x2022;&#x2002;The amplitude scaling is defined as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the augmented data after amplitude scaling, and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>amp</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the amplitude scaling factor. This method simulates variations in sensor range or fluctuations in equipment load [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]. The scaling factor must be restricted within the operational limits of the equipment or sensor to ensure realistic data generation.</p>
<p>&#x2022;&#x2002;The time scaling is defined as follows:
<disp-formula id="eqn-8a"><label>(8a)</label><mml:math id="mml-eqn-8a" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8b"><label>(8b)</label><mml:math id="mml-eqn-8b" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x230B;</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x230B;</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> represents the augmented data after time scaling, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the time scaling factor, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x230B;</mml:mo></mml:mrow></mml:math></inline-formula> represents the floor function, and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>time</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>&#x230B;</mml:mo></mml:mrow></mml:math></inline-formula>. Time scaling simulates variations in the equipment&#x2019;s operating rate while preserving essential physical characteristics [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. This technique improves the generalization of the model under variable conditions and equipment speeds.</p>
<p><bold>3)</bold> The zero-masking method generates new samples by randomly setting continuous segments of the original data to zero, thereby simulating sensor failures or communication interruptions [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-29">29</xref>]. The process is defined as follows:
<disp-formula id="eqn-9a"><label>(9a)</label><mml:math id="mml-eqn-9a" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>trunc</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">l</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>trunc</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>&#x03BB;</mml:mi><mml:mo>+</mml:mo><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9b"><label>(9b)</label><mml:math id="mml-eqn-9b" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>trunc</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x03BB;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>+</mml:mo><mml:mi>L</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mrow><mml:mtext>others</mml:mtext></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>L</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula> is the truncation start position, <italic>L</italic> is the truncation length, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">l</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>trunc</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the binary masking vector, and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> denotes the Hadamard product.</p>
<p><bold>4)</bold> The time-shift method generates new samples by shifting the original time series data to the left or right, thus simulating different time delays [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>]. Given an original sample <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:math></inline-formula>, the shifted samples are defined as:
<disp-formula id="eqn-10a"><label>(10a)</label><mml:math id="mml-eqn-10a" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mrow><mml:mtext>right</mml:mtext></mml:mrow></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-10b"><label>(10b)</label><mml:math id="mml-eqn-10b" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mrow><mml:mtext>left</mml:mtext></mml:mrow></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>right</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>left</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denote the right-shifted and left-shifted samples, respectively, and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>k</mml:mi></mml:math></inline-formula> is the shift length.</p>
<p><bold>5)</bold> The flip method reverses the order of the time series data to improve the model&#x2019;s generalization capability [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]. Given an original sample <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:math></inline-formula>, the flipped sample is defined as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>flip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mtext>org</mml:mtext></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:math></disp-formula></p>
<p>Transform-based data augmentation methods apply various transformation strategies within the bounds of physical constraints to generate new samples and expand the dataset. Although these techniques improve the diversity of the time series data, their applicability is typically limited to specific scenarios. Therefore, it remains challenging to establish a unified framework for their general application.</p>
<p><bold>(b) LSLD based on deep generative models</bold></p>
<p>Deep generative models based on neural networks have emerged as a prominent direction in data augmentation research [<xref ref-type="bibr" rid="ref-105">105</xref>]. These methods aim to learn the latent data distribution using neural architectures, generating new samples through sampling from them. Among the various models, the variational autoencoder (VAE) and the generative adversarial network (GAN) are two representative and widely used methods, as illustrated in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. Moreover, several studies [<xref ref-type="bibr" rid="ref-106">106</xref>] have transformed one-dimensional (1-D) signals into two-dimensional (2-D) images using techniques such as the fast Fourier transform [<xref ref-type="bibr" rid="ref-107">107</xref>], the continuous wavelet transform [<xref ref-type="bibr" rid="ref-108">108</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>], and the Gramian angular field [<xref ref-type="bibr" rid="ref-110">110</xref>]. These transformations enable diffusion models, originally developed for image data, to be applied to LSLD scenarios. The details of three deep generative models are presented in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Diagram of the typical deep generative models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-9.tif"/>
</fig><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Characteristics of various deep generative models</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th colspan="3">Generation process</th>
</tr>
</thead>
<tbody>
<tr>
<td>VAE</td>
<td colspan="3">The VAE consists of an encoder and a decoder. The encoder learns the distributional characteristics of the original data and employs the reparameterization trick to obtain latent variables. The decoder then reconstructs the data based on latent variables [<xref ref-type="bibr" rid="ref-111">111</xref>].</td>
</tr>
<tr>
<td></td>
<td><bold>Generation quality</bold></td>
<td><bold>Training stability</bold></td>
<td><bold>Generation speed</bold></td>
</tr>
<tr>
<td></td>
<td>Middle</td>
<td>High</td>
<td>Fast</td>
</tr>
<tr>
<td></td>
<td colspan="3"><bold>Generation process</bold></td>
</tr>
<tr>
<td>GAN</td>
<td colspan="3">The GAN consists of a generator and a discriminator. The generator receives random noise as input and generates virtual samples that approximate real samples. The discriminator receives real and virtual samples and tries to distinguish between them [<xref ref-type="bibr" rid="ref-112">112</xref>].</td>
</tr>
<tr>
<td></td>
<td><bold>Generation quality</bold></td>
<td><bold>Training stability</bold></td>
<td><bold>Generation speed</bold></td>
</tr>
<tr>
<td></td>
<td>High</td>
<td>Low</td>
<td>Fast</td>
</tr>
<tr>
<td></td>
<td colspan="3"><bold>Generation process</bold></td>
</tr>
<tr>
<td>Diffusion model</td>
<td colspan="3">The diffusion model generates data by iteratively denoising random noise through a reverse diffusion process. This approach offers better stability and sample quality by decomposing the generation process into a series of gradual and deterministic steps [<xref ref-type="bibr" rid="ref-113">113</xref>].</td>
</tr>
<tr>
<td></td>
<td><bold>Generation quality</bold></td>
<td><bold>Training stability</bold></td>
<td><bold>Generation speed</bold></td>
</tr>
<tr>
<td></td>
<td>High</td>
<td>High</td>
<td>Slow</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>1)</bold> VAEs can learn the original data distribution and generate new samples by sampling the latent space [<xref ref-type="bibr" rid="ref-31">31</xref>]. To improve the performance of VAEs for industrial data, existing studies have primarily focused on optimizing the objective function and network architectures.</p>
<p>&#x2022;&#x2002;In terms of objective functions, Dixit et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] proposed a fault diagnosis framework based on a modified conditional VAE with centroid loss, which effectively enhances the generation of virtual samples. Zhao et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] developed a modified Wasserstein autoencoder that integrates a squeeze-and-excitation attention mechanism and utilizes the sliced Wasserstein distance along with a gradient penalty to improve sample similarity. In another study, Zhao et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] further improved the representational capacity of VAEs by introducing a Gaussian mixture prior and optimizing it via the expectation-maximization algorithm.</p>
<p>&#x2022;&#x2002;In terms of network architectures, Karamti et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] introduced a stacked VAE framework for multi-fault machinery identification, in which a VAE is used for data augmentation and two sparse autoencoders are used for feature extraction. Han et al. [<xref ref-type="bibr" rid="ref-114">114</xref>] proposed a VAE incorporating long short-term memory units to capture temporal dependencies and generate realistic virtual samples. To address the limitations of non-Gaussian signals and sparse fault data, Luo et al. [<xref ref-type="bibr" rid="ref-115">115</xref>] developed a mixture network VAE by integrating Gaussian mixture models and Weibull distributions. Wang et al. [<xref ref-type="bibr" rid="ref-116">116</xref>] combined the noise robustness of VAE with the expressive capacity of convolutional neural networks (CNNs) to enhance diagnostic precision and robustness. Zhang et al. [<xref ref-type="bibr" rid="ref-117">117</xref>] proposed a multi-scale dilated variational convolutional autoencoder, which integrates a multi-scale dilated convolutional attention mechanism and a graph convolution module to capture both hierarchical features and inter-sensor dependencies. For CNN structure, Zeng et al. [<xref ref-type="bibr" rid="ref-118">118</xref>] introduced an attention-guided hierarchical wavelet CNN that integrates multi-layer wavelet decomposition, a time-frequency attention module, and gradient-weight fusion to simultaneously suppress noise and highlight fault-relevant frequency bands.</p>
<p><bold>2)</bold> The training process of GANs is formulated as a minimax game between two neural networks: the generator and the discriminator. The overall performance of GANs is influenced by the structures of both components. Moreover, the stability and convergence of the training process depend on the design of loss functions. Consequently, existing studies have focused on three aspects: objective functions, generator structures, and discriminator structures. The development trend of these studies is summarized in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Development trend of GANs for LSLD scenarios [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-64">64</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-10.tif"/>
</fig>
<p>&#x2022;&#x2002;Regarding objective functions, the Wasserstein distance has been used as a replacement for the Kullback-Leibler (KL) divergence, effectively alleviating problems such as gradient vanishing and mode collapse [<xref ref-type="bibr" rid="ref-36">36</xref>]. Based on this advancement, [<xref ref-type="bibr" rid="ref-37">37</xref>] integrated the Wasserstein distance with a hierarchical feature matching loss to constrain local category similarity, thus improving the quality and validity of generated samples. Subsequently, reference [<xref ref-type="bibr" rid="ref-38">38</xref>] introduced a gradient penalty term to further stabilize the training process and reduce the risk of mode collapse. Moreover, studies such as [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>] incorporated the Wasserstein distance and the gradient penalty into the time series GAN and the auxiliary classifier GAN, respectively, resulting in improved model convergence.</p>
<p>&#x2022;&#x2002;Regarding generator structures, Pan et al. [<xref ref-type="bibr" rid="ref-119">119</xref>] enhanced the generator by incorporating an additional sequence into the condition, allowing the generation of more abundant samples. Zareapoor et al. [<xref ref-type="bibr" rid="ref-120">120</xref>] proposed a minority oversampling GAN, where the generator learns a mixture data distribution to generate samples from the minority class. Zhang et al. [<xref ref-type="bibr" rid="ref-121">121</xref>] introduced an adaptive decoupling strategy in the generator, which adjusts each intermediate output using labels to prevent mode collapse and improve sample diversity. Xu et al. [<xref ref-type="bibr" rid="ref-122">122</xref>] employed multiple generators to generate virtual samples from the real data distribution, thus increasing sample variety and reducing the risk of mode collapse. Ren et al. [<xref ref-type="bibr" rid="ref-123">123</xref>] improved the generator through pre-training based on the majority classes, followed by fine-tuning with anchor samples. This approach preserves the learned distribution from pre-training while ensuring that the generated samples remain close to real ones. Huo et al. [<xref ref-type="bibr" rid="ref-124">124</xref>] incorporated a residual mixed self-attention module into the generator to effectively extract time- and frequency-domain features. Chen et al. [<xref ref-type="bibr" rid="ref-125">125</xref>] utilized a serial CNN-Transformer architecture as the foundation of the generator to capture long-range dependencies and improve understanding of local features.</p>
<p>&#x2022;&#x2002;Regarding discriminator structures, Wang et al. [<xref ref-type="bibr" rid="ref-126">126</xref>] proposed a GAN framework where the generator generates new samples to expend the dataset, and a stacked denoising autoencoder serves as the discriminator to extract robust fault-related features, thereby enhancing the model&#x2019;s fault classification capability. Ding et al. [<xref ref-type="bibr" rid="ref-127">127</xref>] introduced a multi-discriminator structure designed to learn data distributions associated with different health states, improving performance through a semi-supervised learning strategy.</p>
<p><bold>3)</bold> The denoising diffusion probabilistic model (DDPM) is a representative diffusion-based generative model widely applied in data augmentation, as illustrated in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. Reference [<xref ref-type="bibr" rid="ref-41">41</xref>] integrated a classifier-free denoising diffusion implicit model with multiclass contrastive learning to tackle challenges posed by limited fault samples and variable operating conditions. Reference [<xref ref-type="bibr" rid="ref-42">42</xref>] combined DDPM with a physical simulation model to generate diverse fault samples. In addition, reference [<xref ref-type="bibr" rid="ref-43">43</xref>] proposed a DDPM-based method that generates high-fidelity images, effectively enhancing the diversity of feature representations. Reference [<xref ref-type="bibr" rid="ref-44">44</xref>] introduced a lightweight DDPM variant that incorporates multi-dconv head transposed attention, reducing computational complexity while improving the model&#x2019;s ability to capture fine local details.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Diagram of the DDPM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-11.tif"/>
</fig>
<p><bold>(c) LSLD based on transfer learning</bold></p>
<p>Transfer learning-based methods generate target domain data through cross-domain knowledge transfer. These methods utilize the distribution and feature relationships of the source domain to generate new samples in the target domain that conform to actual operating conditions. Recently, Li et al. [<xref ref-type="bibr" rid="ref-128">128</xref>] proposed a cross-domain augmentation method, which employs convex combinations of data and feature-label pairs to create an augmented domain. He et al. [<xref ref-type="bibr" rid="ref-45">45</xref>] introduced an attention-based cross-domain adaptive GAN that applies attention mechanisms for adaptive feature selection and incorporates correlation alignment regularization to enable transferable data augmentation between varying machining parameters. Ge et al. [<xref ref-type="bibr" rid="ref-46">46</xref>] developed a transfer learning method based on multiple mixed augmentations, which combines auto-augment-driven dynamic strategies with multi-stage data mixing. This method enhances cross-condition fault diagnosis by adaptively extracting transferable features from limited vibration data. Jian et al. [<xref ref-type="bibr" rid="ref-47">47</xref>] proposed an auxiliary classifier GAN (ACGAN) to generate open-set samples. This method used the multi-domain mixup to expand fault representation diversity across operating conditions, while ACGAN generates targeted open-set data, as illustrated in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>. Mu et al. [<xref ref-type="bibr" rid="ref-48">48</xref>] proposed a task-oriented meta-learning network based on the Theil index, incorporating a gradient calibration strategy to address the challenge of cross-domain fault diagnosis in rotating machinery under extremely limited sample conditions. This method integrates a fine-grained feature learner, a task-specific Theil index, and a gradient calibration mechanism to enhance model generalization and stability across diverse industrial scenarios. Wang et al. [<xref ref-type="bibr" rid="ref-49">49</xref>] proposed a dynamic collaborative adversarial domain adaptation network that adaptively adjusts the generator and adversarial components to enable unsupervised fault diagnosis of rotating machinery across multiple source domains, without requiring labeled data from the target domain.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Diagram of generating open-set data using the ACGAN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-12.tif"/>
</fig>
<p><bold>(d) LSLD Data augmentation methods based on physical models</bold></p>
<p>Physical model-based methods simulate the behavior of actual systems by establishing mathematical models. These methods are generally divided into two categories: numerical simulations and digital twins. Numerical simulations reproduce the dynamic behavior of physical systems through mathematical equations and computational techniques, while digital twins create real-time, dynamic digital representations of physical entities via virtual-real mapping, as illustrated in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>. Current research primarily focuses on developing physical models for rotating machinery, such as gears, bearings, and rotors, due to their well-defined physical characteristics and mechanical principles, which facilitate the construction of numerical simulations and digital twin models [<xref ref-type="bibr" rid="ref-129">129</xref>,<xref ref-type="bibr" rid="ref-130">130</xref>].</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Diagram of the digital twin model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-13.tif"/>
</fig>
<p><bold>1)</bold> Regarding numerical simulations, references [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>] developed simulation models of rotors and propellers to generate virtual fault samples, thus alleviating data imbalance problems in neural network-based fault diagnosis.</p>
<p><bold>2)</bold> Regarding digital twins, Cai et al. [<xref ref-type="bibr" rid="ref-52">52</xref>] developed a digital twin model of a triplex pump to generate diverse fault data. Qin et al. [<xref ref-type="bibr" rid="ref-53">53</xref>] proposed a data augmentation method based on digital twin for rolling bearings to address class imbalance in fault diagnosis. Zhang et al. [<xref ref-type="bibr" rid="ref-54">54</xref>] constructed a digital twin-based rolling bearing model that integrates virtual modeling with transformer-based discrepancy learning and multi-loss alignment. Ming et al. [<xref ref-type="bibr" rid="ref-55">55</xref>] introduced a virtual-physical component fusion method that incorporates frequency-adaptive filtering and a feature fusion self-attention network to align subdomain features and effectively alleviate distribution discrepancies in imbalanced bearing fault diagnosis.</p>
<p>In this context, various data augmentation methods have been developed to address the challenges of LSLD data. Transform technique-based methods quickly generate new samples through simple signal transformations, offering high computational efficiency and strong interpretability. However, they often lack data diversity and struggle to provide a unified, generalizable framework. Deep generative model-based methods learn the data distribution to generate complex virtual samples, improving model generalization. Despite these benefits, such methods have high training complexity and unstable generation quality. Transfer learning-based methods use cross-domain knowledge transfer to alleviate data scarcity, enhancing adaptability under different operational conditions. However, their effectiveness is sensitive to domain discrepancies and feature alignment. Physical model-based methods expand datasets with strong physical interpretability and high fidelity but involve high modeling costs and offer limited adaptability.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Augmentation for SSHD</title>
<p>The sparsity introduced by the CoD often causes overfitting in data-driven models. Therefore, data augmentation methods for SSHD data need to address the challenges associated with high dimensionality during sample generation processes. In this regard, statistical technique-based, deep generative model-based, and transfer learning-based approaches have been developed to effectively augment SSHD data and improve model generalization.</p>
<p><bold>(a) SSHD based on statistical techniques</bold></p>
<p>Statistical technique-based methods aim to learn the latent distribution of the original dataset and generate new samples, using statistical models in sparse areas, as illustrated in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>. However, directly applying traditional statistical methods to high-dimensional data poses significant challenges due to CoD. Feature engineering serves as an effective solution to alleviate these limitations by reducing dimensionality and highlighting relevant features.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Diagram of the data augmentation based on statistical techniques</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-14.tif"/>
</fig>
<p>Feature selection methods aim to obtain a subset of features that are most relevant to the target variables [<xref ref-type="bibr" rid="ref-56">56</xref>&#x2013;<xref ref-type="bibr" rid="ref-58">58</xref>]. In contrast, feature extraction methods transform the original high-dimensional data into a low-dimensional space to capture essential patterns and reduce redundancy [<xref ref-type="bibr" rid="ref-59">59</xref>]. These techniques are effective in capturing nonlinear characteristics and enhancing low-dimensional representations. For example, reference [<xref ref-type="bibr" rid="ref-60">60</xref>] employed locally linear embedding to map high-dimensional data into a lower-dimensional space, which facilitated the identification of sparse areas for virtual sample generation (VSG). Similarly, reference [<xref ref-type="bibr" rid="ref-17">17</xref>] applied Isomap to embed data in a two-dimensional space, enabling a sparse area detection and interpolation-based augmentation. In [<xref ref-type="bibr" rid="ref-61">61</xref>], <italic>t</italic>-SNE was used for feature extraction and then virtual samples were generated by interpolation in the reduced space. In addition, reference [<xref ref-type="bibr" rid="ref-62">62</xref>] used discriminant locality preserving projection to handle high-dimensional data and introduced Monte Carlo sampling for VSG. Reference [<xref ref-type="bibr" rid="ref-63">63</xref>] proposed detecting sparse regions based on projection point spacing and then generated new samples based on midpoint and radial basis function interpolation. Finally, reference [<xref ref-type="bibr" rid="ref-131">131</xref>] utilized singular value decomposition to extract principal features, thus effectively addressing the challenge of high dimensionality.</p>
<p>In addition, multi-model methods have been explored to address the complexities associated with SSHD data. For example, reference [<xref ref-type="bibr" rid="ref-64">64</xref>] proposed a co-training strategy that detects sparse areas and fills them through interpolation. This method dynamically selects virtual samples based on model performance, improving the quality and relevance of the generated data. In another study, reference [<xref ref-type="bibr" rid="ref-65">65</xref>] introduced an adaptive data augmentation approach using generalized correntropy, which effectively utilizes <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>-order statistics to learn relationships between samples. This method is suited for scenarios such as high dimensionality, non-Gaussian noise, and data uncertainty.</p>
<p><bold>(b) SSHD based on deep generative models</bold></p>
<p>High-dimensional process data exhibit more complex and diverse characteristics compared to low-dimensional data, which presents significant challenges for training deep generative models. Therefore, the development and training of such models for SSHD data should be designed for specific tasks.</p>
<p><bold>1)</bold> VAEs learn the low-dimensional latent distribution of the original data and generate new samples accordingly. To address information loss during layer-wise pretraining, Yuan et al. [<xref ref-type="bibr" rid="ref-66">66</xref>] generated virtual samples at each layer through linear interpolation between adjacent samples, thereby improving the training of stacked autoencoders. Jiang et al. [<xref ref-type="bibr" rid="ref-67">67</xref>] introduced a causality-informed VAE to alleviate data sparsity and high dimensionality in just-in-time learning for soft sensor applications. Tian et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] developed an adaptive loss function and utilized kernel mean matching to assign weights to virtual samples, thus improving data quality and learning performance. Moreover, Xu et al. [<xref ref-type="bibr" rid="ref-68">68</xref>,<xref ref-type="bibr" rid="ref-69">69</xref>] integrated VAE with SMOTE and Glow models to stabilize the training process and enhance sample generation.</p>
<p><bold>2)</bold> Training GANs with SSHD data often faces significant challenges, such as gradient vanishing and mode collapse. To address them, references [<xref ref-type="bibr" rid="ref-70">70</xref>,<xref ref-type="bibr" rid="ref-71">71</xref>] introduced a gradient penalty into the Wasserstein GAN (WGAN-GP), which stabilizes the training process and enhances the quality of generated samples. Based on this foundation, reference [<xref ref-type="bibr" rid="ref-72">72</xref>] incorporated a deep neural network regressor into a conditional WGAN-GP framework to address the small sample problem in soft sensor. Subsequently, Jiang et al. [<xref ref-type="bibr" rid="ref-73">73</xref>] extended this architecture to a multi-generator setting, enabling more effective handling of complex data imbalances. To address the any-shot learning problem in industrial fault diagnosis, Zhuo et al. [<xref ref-type="bibr" rid="ref-74">74</xref>] proposed a data augmentation method that integrates auxiliary fault attribute information with GANs. Moreover, Cui et al. [<xref ref-type="bibr" rid="ref-132">132</xref>] introduced fuzzy set theory into the GAN framework, which effectively alleviates uncertainty and high-dimensional challenges in data augmentation. Liu et al. [<xref ref-type="bibr" rid="ref-133">133</xref>] developed a robust VAE to constrain the generator&#x2019;s sampling space, while density-based spatial clustering of applications with noise clustering was introduced to guide the generation process. Finally, Zhang et al. [<xref ref-type="bibr" rid="ref-134">134</xref>] incorporated regression loss assistance into a conditional StyleGAN to generate high-quality virtual samples.</p>
<p><bold>(c) SSHD based on transfer learning</bold></p>
<p>Recently, transfer learning has gained attention as a promising solution for SSHD data. Ren et al. [<xref ref-type="bibr" rid="ref-75">75</xref>] proposed a low-rank joint domain adaptation network, which enhances the effectiveness of industrial small-sample datasets by extracting discriminative features and aligning cross-domain sample distributions. Furthermore, Zhu et al. [<xref ref-type="bibr" rid="ref-76">76</xref>] introduced a data self-generating transfer learning framework designed for chemical process fault diagnosis under limited sample conditions. Li et al. [<xref ref-type="bibr" rid="ref-77">77</xref>] developed the dual adversarial and contrastive network, a data augmentation method that combines adversarial and contrastive learning. It is designed to generate virtual samples while simultaneously extracting domain-invariant feature representations from single-modality data.</p>
<p>In this context, various data augmentation methods have been developed to address the challenges associated with SSHD data. Statistical technique-based methods typically perform dimensionality reduction before VSG and employ interpolation in reduced spaces. These methods are valued for their simplicity and interpretability but struggle to capture the characteristics inherent in high-dimensional data. Deep generative model-based methods, on the other hand, utilize neural networks to extract features and learn latent distributions to generate virtual samples. Transfer learning-based methods address domain discrepancies by transferring knowledge from source to target domains, incorporating adversarial or contrastive learning mechanisms to generate domain-invariant features and samples. In summary, these methods alleviate overfitting risks and improve model robustness in SSHD scenarios.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Data Augmentation for SSLD</title>
<p>SSLD data present challenges due to their limited latent information, which causes underfitting and hinders model performance. Data augmentation methods based on statistical techniques and deep generative models have been introduced to address the above problems and expand the size of the dataset.</p>
<p><bold>(a) SSLD based on statistical techniques</bold></p>
<p>In statistical techniques, data augmentation methods are generally classified into distributional assumption-based and interpolation-based approaches, depending on their generative principles. These studies are summarized in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Summarization of SSLD based on statistical techniques</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Categorization</th>
<th>Methods</th>
<th>Years</th>
<th>References</th>
</tr>
</thead>
<tbody>
<tr>
<td>Distributional assumption-based methods</td>
<td>Generalized internalized kernel density estimator (GIKDE)</td>
<td>2008</td>
<td>[<xref ref-type="bibr" rid="ref-78">78</xref>]</td>
</tr>
<tr>
<td></td>
<td>Maximal <italic>p</italic>-value with Weibull distribution</td>
<td>2013</td>
<td>[<xref ref-type="bibr" rid="ref-79">79</xref>]</td>
</tr>
<tr>
<td></td>
<td>Small Johnson data transformation (SJDT)</td>
<td>2016</td>
<td>[<xref ref-type="bibr" rid="ref-80">80</xref>]</td>
</tr>
<tr>
<td></td>
<td>Mega-trend diffusion (MTD)</td>
<td>2007</td>
<td>[<xref ref-type="bibr" rid="ref-135">135</xref>]</td>
</tr>
<tr>
<td></td>
<td>MTD with particle swarm optimization (PSO)</td>
<td>2020</td>
<td>[<xref ref-type="bibr" rid="ref-81">81</xref>]</td>
</tr>
<tr>
<td></td>
<td><italic>k</italic>-Nearest Neighbor MTD (<italic>k</italic>NNMTD)</td>
<td>2022</td>
<td>[<xref ref-type="bibr" rid="ref-82">82</xref>]</td>
</tr>
<tr>
<td>Interpolation-based methods</td>
<td>Weighted kernel-based SMOTE (WK-SMOTE)</td>
<td>2018</td>
<td>[<xref ref-type="bibr" rid="ref-83">83</xref>]</td>
</tr>
<tr>
<td></td>
<td>Range-controlled SMOTE (RCSMOTE)</td>
<td>2021</td>
<td>[<xref ref-type="bibr" rid="ref-84">84</xref>]</td>
</tr>
<tr>
<td></td>
<td>Multi-vector stochastic exploration oversampling (MSEO)</td>
<td>2024</td>
<td>[<xref ref-type="bibr" rid="ref-85">85</xref>]</td>
</tr>
<tr>
<td></td>
<td>Nonlinear interpolation virtual sample generation (NIVSG)</td>
<td>2018</td>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
</tr>
<tr>
<td></td>
<td>Kriging-VSG</td>
<td>2020</td>
<td>[<xref ref-type="bibr" rid="ref-86">86</xref>]</td>
</tr>
<tr>
<td></td>
<td>Data augmentation and weighted interpolation (DAWI-VSG)</td>
<td>2023</td>
<td>[<xref ref-type="bibr" rid="ref-87">87</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>1)</bold> Distributional assumption-based methods fit a probability distribution to existing data and then generate new samples by sampling from the fitted distribution. The low dimensionality of SSLD data facilitates the modeling of their distributional characteristics. For example, in [<xref ref-type="bibr" rid="ref-78">78</xref>], the generalized internalized kernel density estimation was used to incorporate time-dependent data properties and address the challenges of knowledge acquisition in early-stage manufacturing systems. Reference [<xref ref-type="bibr" rid="ref-79">79</xref>] introduced a maximal <italic>p</italic>-value method based on the two-parameter Weibull distribution to generate virtual samples and evaluate product lifetime performance in scenarios with limited datasets. In [<xref ref-type="bibr" rid="ref-80">80</xref>], the small Johnson data transformation was proposed to normalize small datasets for effective VSG.</p>
<p>Information diffusion theory [<xref ref-type="bibr" rid="ref-136">136</xref>], which integrates distributional assumptions with fuzzy set theory, has been applied to estimate feature expansion ranges. Based on this, the mega-trend diffusion (MTD) method introduces an asymmetric diffusion mechanism to generate virtual samples [<xref ref-type="bibr" rid="ref-135">135</xref>], as illustrated in <xref ref-type="fig" rid="fig-15">Fig. 15a</xref>. Liu et al. [<xref ref-type="bibr" rid="ref-81">81</xref>] further improved the MTD method by integrating PSO, allowing an accurate prediction of forming forces in single-point incremental forming processes under the conditions of limited data. To address challenges in supervised and unsupervised learning with small samples, Sivakumar et al. [<xref ref-type="bibr" rid="ref-82">82</xref>] proposed a modified version of MTD, termed k-Nearest Neighbor MTD.</p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Diagram of typical statistical technique-based methods</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-15.tif"/>
</fig>
<p><bold>2)</bold> Interpolation-based methods represent another prominent direction in data augmentation research. These techniques generate new samples based on relationships between existing data points, using linear or nonlinear interpolation.</p>
<p>&#x2022;&#x2002;A representative linear approach is the synthetic minority over-sampling technique (SMOTE) [<xref ref-type="bibr" rid="ref-137">137</xref>], which interpolates between samples from minority class, as illustrated in <xref ref-type="fig" rid="fig-15">Fig. 15b</xref>. To address limitations of standard SMOTE, Mathew et al. [<xref ref-type="bibr" rid="ref-83">83</xref>] proposed weighted kernel-based SMOTE, which performs oversampling in the feature space of support vector machines to better preserve minority class characteristics. Soltanzadeh et al. [<xref ref-type="bibr" rid="ref-84">84</xref>] introduced range-controlled SMOTE to alleviate class overlap near decision boundaries. Li et al. [<xref ref-type="bibr" rid="ref-85">85</xref>] proposed multi-vector stochastic exploration oversampling, which generates diverse virtual samples through randomized direction and scaling vectors. Khan [<xref ref-type="bibr" rid="ref-138">138</xref>] proposed a balanced weighted extreme learning machine (WELM) that integrates k-fold sampling strategies, such as SMOTE and TomekLinks, with a cost-sensitive WELM framework to reduce data complexity and improve minority class classification in imbalanced datasets.</p>

<p>&#x2022;&#x2002;For nonlinear interpolation, He et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] developed the nonlinear interpolation VSG method, integrated with neural networks to improve the accuracy of energy prediction. In another study, Zhu et al. [<xref ref-type="bibr" rid="ref-86">86</xref>] utilized a distance-based criterion to detect sparse data regions and employed Kriging interpolation to generate virtual samples. Similarly, Song et al. [<xref ref-type="bibr" rid="ref-87">87</xref>] proposed data augmentation with weighted interpolation for the VSG method, which improves sample quality for soft sensing models.</p>
<p><bold>(b) SSLD based on deep generative models</bold></p>
<p>For classification tasks, deep generative models play a crucial role in expanding minority class datasets to alleviate class imbalance and improve model performance. In [<xref ref-type="bibr" rid="ref-88">88</xref>], an augmented generator was introduced to produce high-quality virtual samples representing abnormal states. This approach incorporated augmented filter layers and batching techniques to enhance the diversity and reliability of the generated data. Then, reference [<xref ref-type="bibr" rid="ref-89">89</xref>] integrated a multi-head attention mechanism into a GAN framework, effectively improving the sample quality and modeling performance. Furthermore, reference [<xref ref-type="bibr" rid="ref-90">90</xref>] proposed two augmentation strategies, individual-based and concatenation-based, which enhanced the representation of minority classes and effectively balanced the dataset.</p>
<p>For regression tasks, deep generative models must generate virtual input-output pairs that maintain their mapping relationship. In [<xref ref-type="bibr" rid="ref-91">91</xref>], Gaussian noise was injected into the features extracted by a target-relevant autoencoder to generate informative input-output sample pairs. In [<xref ref-type="bibr" rid="ref-92">92</xref>], labels were incorporated as conditional information in a GAN framework, allowing the generation of labeled samples specifically designed to populate the sparse regions of the feature space. Then, reference [<xref ref-type="bibr" rid="ref-93">93</xref>] integrated GAN with vine copula regression to address the challenge of limited labeled data in complex chemical processes. In [<xref ref-type="bibr" rid="ref-94">94</xref>], a supervised variational autoencoder and a WGAN-GP were combined to generate high-quality labeled samples, thus improving the accuracy and robustness of soft sensor models. In [<xref ref-type="bibr" rid="ref-95">95</xref>], the TimeCVAE model incorporated time-dependent features and conditional information into the virtual sample generation process, improving the performance of time series prediction.</p>
<p>In this context, various data augmentation methods have been proposed for SSLD data. Statistical technique-based methods emphasize simplicity and computational efficiency. These methods are well-suited to low-dimensional data, where they can effectively infer latent information. However, they often face limitations when modeling complex, high-dimensional relationships. In contrast, deep generative model-based methods aim to learn latent data distributions through neural networks. Although these models are highly expressive and capable of capturing complex patterns, they require careful hyperparameter tuning to ensure training stability. Consequently, statistical methods provide greater interpretability and reliability in low-dimensional scenarios, whereas deep generative models demonstrate superior adaptability and performance.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Data Augmentation for LSHD</title>
<p>High dimensionality in industrial datasets often results from multi-source data fusion across different processes, causing massive and complex datasets. However, these high-dimensional datasets contain substantial noise and redundant features. As a result, the research focus should shift from achieving &#x2018;sufficient data size&#x2019; to ensuring &#x2018;sufficient data quality and scenario coverage&#x2019;. In this context, deep generative model-based and physical model-based data augmentation methods have been developed to address these challenges.</p>
<p><bold>(a) LSHD based on deep generative models</bold></p>
<p>For LSHD data, developing effective deep generative models remains a challenging task due to data sparsity and variability. Recent studies have addressed these limitations through task-specific model designs. For example, reference [<xref ref-type="bibr" rid="ref-96">96</xref>] proposed three data-level augmentation strategies:
<list list-type="bullet">
<list-item>
<p>a VAE-based augmentation approach that reconstructs virtual samples by learning latent feature distributions;</p></list-item>
<list-item>
<p>a CVAE-based balancing approach that employs label-guided generation to expand minority class;</p></list-item>
<list-item>
<p>a hybrid CVAE approach combined with random undersampling, which reduces the dominance of majority classes while retaining essential information.</p></list-item>
</list></p>
<p>In [<xref ref-type="bibr" rid="ref-97">97</xref>], a complementary classifier was incorporated into a basic GAN framework to augment minority class and improve classification accuracy through collaborative adversarial training. Then, reference [<xref ref-type="bibr" rid="ref-98">98</xref>] proposed a data augmentation framework that integrates cooperative and competitive strategies. The cooperative strategy utilizes cross-training and parallel training among multiple GANs to improve diversity and efficiency. In contrast, the competitive strategy introduces a filtering mechanism to retain only high-quality generated samples. Zaman et al. [<xref ref-type="bibr" rid="ref-99">99</xref>] developed a comprehensive framework combining preprocessing, GAN-based augmentation, automated feature selection, and predictive modeling to alleviate class imbalance and enhance generalization to unseen data. To support privacy-preserving learning, Feng et al. [<xref ref-type="bibr" rid="ref-100">100</xref>] integrated CVAE and GAN to generate balanced virtual samples in federated learning environments and used geometric median aggregation to ensure privacy preservation. Furthermore, Zhang et al. [<xref ref-type="bibr" rid="ref-101">101</xref>] employed Isomap for dimensionality reduction while preserving critical data structures and used CGAN to generate additional fault samples.</p>
<p>Furthermore, for high-precision quality prediction tasks, the availability of sufficient training samples is critical. To address this, reference [<xref ref-type="bibr" rid="ref-139">139</xref>] proposed a two-phase framework that first introduced TimeGAN to generate temporally consistent data to impute missing values and then used minimal gated unit to enable efficient quality prediction with reduced computational complexity.</p>
<p><bold>(b) LSHD based on physical models</bold></p>
<p>Huang et al. [<xref ref-type="bibr" rid="ref-102">102</xref>] proposed an edge-intelligent digital twin framework that utilizes edge-cloud collaboration and real-time data processing to identify early-stage faults in automation systems. In another study, Krespach et al. [<xref ref-type="bibr" rid="ref-103">103</xref>] introduced a hybrid data augmentation method that combines historical data with virtual data generated from a digital twin. This method addresses the limitations of traditional data-driven predictive control by generating virtual samples that cover previously unexplored operational conditions.</p>
<p>In this context, data augmentation methods for LSHD data have been preliminarily explored. Among them, deep generative model-based methods effectively address the challenges of high-dimensional complexity and class imbalance by learning latent data distributions. However, these methods often involve complex training procedures, high computational costs, and risks of pattern collapse. In contrast, physical model-based methods generate high-quality samples grounded in physical mechanisms, enabling the capture of complex dynamic behaviors and improving prediction reliability. However, their applicability to unseen operating conditions remains limited due to strict model structures and high modeling overhead.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Application of Data Augmentation</title>
<p>This section summarizes current advances in data augmentation methods across three critical industrial domains: chemical processes, rotating machinery, and municipal solid waste incineration (MSWI). In addition, it highlights emerging research trends within each domain, aiming to offer valuable insights and practical references for researchers and engineers working in related fields.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Chemical Process</title>
<p>The chemical process involves the transformation of raw materials into final products through chemical and physical operations. It is a critical industry where accurate process monitoring and key performance index prediction are essential. Numerous studies have focused on these tasks using simulated and actual datasets. A typical benchmark is the Tennessee Eastman (TE) platform, a simulation system modeled on an actual chemical reaction process [<xref ref-type="bibr" rid="ref-140">140</xref>,<xref ref-type="bibr" rid="ref-141">141</xref>], as shown in <xref ref-type="fig" rid="fig-16">Fig. 16</xref>. The TE dataset includes 11 manipulated and 41 measured variables across various operating conditions, as it is suitable for evaluating data augmentation methods. Beyond simulations, researchers have also investigated actual industrial scenarios, such as purified terephthalic acid and ethylene production processes, to assess the practical effectiveness of proposed data augmentation methods.</p>
<fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>Diagram of the TE process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-16.tif"/>
</fig>
<p><bold>1)</bold> To alleviate the class imbalance in TE datasets, various data augmentation methods have been proposed. For example, Jiang et al. [<xref ref-type="bibr" rid="ref-73">73</xref>] developed a multi-generator framework to overcome the limitations of single-generator models in handling complex imbalance scenarios. Zhuo et al. [<xref ref-type="bibr" rid="ref-74">74</xref>] introduced auxiliary fault attributes into a GAN framework, enabling VSG for rare and unseen faults, thus supporting zero-shot and few-shot fault diagnosis. Furthermore, Xu et al. [<xref ref-type="bibr" rid="ref-68">68</xref>] proposed a deep encoder-decoder network integrated with SMOTE to balance data distribution, complemented by an ensemble classifier to improve diagnostic accuracy. Transfer learning has also gained advances in augmented TE datasets. Ren et al. [<xref ref-type="bibr" rid="ref-75">75</xref>] extracted simplified and accurate fault features from multiple conditions to augment data under the current condition. Similarly, Zhu et al. [<xref ref-type="bibr" rid="ref-76">76</xref>] utilized adversarial learning to transfer features of real samples for VSG, paired with a model-based transfer learning approach to improve robustness against low-quality generated samples. In addition, Li et al. [<xref ref-type="bibr" rid="ref-77">77</xref>] addressed single-source domain generalization by generating virtual samples and learning domain-invariant feature representations, thereby enabling cross-mode fault diagnosis.</p>
<p><bold>2)</bold> In purified terephthalic acid (PTA) processes, the conductivity of the solvent dehydration tower is a key quality index, so the development of accurate soft sensors is important. However, due to the SSHD characteristics of the PTA data, multiple learning techniques such as locally linear embedding [<xref ref-type="bibr" rid="ref-60">60</xref>], Isomap [<xref ref-type="bibr" rid="ref-17">17</xref>], and <italic>t</italic>-SNE [<xref ref-type="bibr" rid="ref-61">61</xref>] have been widely employed to reduce dimensionality and identify sparse areas to generate samples based on interpolation. For example, reference [<xref ref-type="bibr" rid="ref-63">63</xref>] utilized projection point spacing in the low-dimensional feature space to detect sparsity and then generated virtual samples using midpoint and radial basis function interpolation. Similarly, reference [<xref ref-type="bibr" rid="ref-131">131</xref>] applied the singular value decomposition to extract the principal features and expand datasets and generated the corresponding output values using the gradient boosting decision tree. In another study, reference [<xref ref-type="bibr" rid="ref-64">64</xref>] generated virtual samples through interpolation and used co-trained KNN regressors to generate the associated outputs. Furthermore, Xu et al. [<xref ref-type="bibr" rid="ref-69">69</xref>] proposed a Gaussian VAE-based VSG method, which employs improved least squares regression to generate high-quality virtual outputs.</p>
<p><bold>3)</bold> In ethylene production processes, accurate energy analysis is vital to optimize industrial operations. However, the limited data on energy consumption pose significant challenges. To address this, data augmentation techniques have been widely utilized. For example, reference [<xref ref-type="bibr" rid="ref-16">16</xref>] applied interpolation within the hidden layers of neural networks to generate virtual input samples, while employing the Moore-Penrose generalized inverse to estimate the corresponding output values. Similarly, reference [<xref ref-type="bibr" rid="ref-86">86</xref>] utilized dimension-wise interpolation via Kriging to create feasible virtual samples in sparse regions, thus improving the prediction accuracy of soft sensors. In addition, several studies have focused on deep generative models to further improve the quality of generated samples. Reference [<xref ref-type="bibr" rid="ref-91">91</xref>] introduced Gaussian noise into latent features extracted from a target-relevant VAE, generating informative input-output pairs. Meanwhile, in [<xref ref-type="bibr" rid="ref-92">92</xref>], a combination of local outlier factor and K-means clustering was used to identify sparse regions for input generation, and a CGAN was used to generate the corresponding output values. These approaches significantly improved the performance of data-driven soft sensor models under limited data.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Rotating Machinery</title>
<p>Rotating machinery includes mechanical systems that perform energy conversion or transmission through rotational motion. As essential components in modern industrial systems, their health monitoring and fault diagnosis are crucial to ensuring operational safety and reliability [<xref ref-type="bibr" rid="ref-142">142</xref>,<xref ref-type="bibr" rid="ref-143">143</xref>]. To facilitate theoretical research and validation of diagnostic methods, several benchmark datasets have been developed. Among them, the Case Western Reserve University (CWRU) bearing dataset is one of the most widely utilized in the diagnosis of rotating machinery faults [<xref ref-type="bibr" rid="ref-144">144</xref>], as illustrated in <xref ref-type="fig" rid="fig-17">Fig. 17</xref>. This dataset includes signals collected under normal conditions as well as various fault scenarios, such as 12k and 48k drive-end bearing faults and fan-end bearing faults. Similarly, the University of Connecticut (UoC) gearbox dataset provides experimental data from a two-stage gearbox system equipped with interchangeable gears [<xref ref-type="bibr" rid="ref-145">145</xref>]. It involves a series of conditions, including healthy operation, missing teeth, root cracks, spalling, and chipped teeth. Due to the limited availability of fault data, numerous data augmentation methods have been proposed to expand datasets and enhance robustness and generalization.</p>
<fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>Diagram of the CWRU motor experimental platform</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-17.tif"/>
</fig>
<p>The data augmentation methods based on transform technique generate new samples by applying various transformations to original vibration signals. Common techniques include noise injection, geometric scaling, zero-masking, time-shifting, and signal flipping [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-104">104</xref>]. In recent years, deep generative models have become a major focus in machine fault diagnosis. Among these, VAEs employ an encoder to extract distributional representations and a decoder to reconstruct low-dimensional latent variables [<xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-114">114</xref>&#x2013;<xref ref-type="bibr" rid="ref-117">117</xref>]. GANs use adversarial learning between a generator and a discriminator to generate high-quality virtual samples [<xref ref-type="bibr" rid="ref-36">36</xref>&#x2013;<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-119">119</xref>&#x2013;<xref ref-type="bibr" rid="ref-127">127</xref>]. Based on these, the conversion of 1-D vibration signals into two-dimensional 2-D images has enabled the use of more advanced deep learning algorithms. Recently, diffusion models based on Markov chain processes have emerged as an effective generative framework that offers stable training and high-quality image generation [<xref ref-type="bibr" rid="ref-41">41</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>]. Meanwhile, transfer learning techniques use data and knowledge from a source domain to generate more diverse and informative samples for a target domain [<xref ref-type="bibr" rid="ref-45">45</xref>&#x2013;<xref ref-type="bibr" rid="ref-49">49</xref>,<xref ref-type="bibr" rid="ref-128">128</xref>]. Finally, several studies have introduced numerical simulations [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>] and digital twin models [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-55">55</xref>], which are based on physical laws and mechanistic knowledge, to expand datasets and improve model generalization capabilities.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Municipal Solid Waste Incineration Process (MSWI)</title>
<p>The MSWI process has emerged as the primary approach to managing MSW, due to its advantages in innocuity, reduction, and reuse. The MSWI process comprises several stages: feeding, combustion, heat exchange, gas cleaning, and emissions [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-146">146</xref>], as illustrated in <xref ref-type="fig" rid="fig-18">Fig. 18</xref>. Among emission by-products, dioxins (DXNs) represent a critical environmental pollutant because of their high toxicity and persistence. However, DXNs are difficult to detect online, resulting in a limited number of effective samples with true values [<xref ref-type="bibr" rid="ref-5">5</xref>]. Moreover, since DXN emissions are influenced by the whole MSWI process, the DXN data are high-dimensional. Therefore, data augmentation methods have become essential for enabling data-driven modeling of DXN emissions.</p>
<fig id="fig-18">
<label>Figure 18</label>
<caption>
<title>Diagram of the MSWI process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_69097-fig-18.tif"/>
</fig>
<p>For high-dimensional DXN data, dimensionality reduction is a critical preprocessing step. In [<xref ref-type="bibr" rid="ref-56">56</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>,<xref ref-type="bibr" rid="ref-147">147</xref>,<xref ref-type="bibr" rid="ref-148">148</xref>], features were selected based on domain knowledge and expert experience to reduce data complexity. In [<xref ref-type="bibr" rid="ref-132">132</xref>], the random forest algorithm was introduced to achieve a subset of features relevant to DXN emissions by modifying the model performance. In [<xref ref-type="bibr" rid="ref-59">59</xref>], the principal component analysis was applied to extract low-dimensional features that capture the dominant variance in the data.</p>
<p>Based on this basic, reference [<xref ref-type="bibr" rid="ref-56">56</xref>] improved the MTD method by expanding the range of input features and employing equal-interval interpolation to generate virtual input-output pairs. Similarly, reference [<xref ref-type="bibr" rid="ref-58">58</xref>] used MTD and interpolation techniques to generate virtual samples and further enhanced the PSO algorithm to select high-quality samples. Then, reference [<xref ref-type="bibr" rid="ref-147">147</xref>] used a multi-objective PSO approach to identify optimal virtual samples for soft sensor modeling, balancing model accuracy with minimal sample selection. Reference [<xref ref-type="bibr" rid="ref-59">59</xref>] generated virtual samples along independent principal components using kernel density estimation and obtained virtual outputs through mapping models. Furthermore, reference [<xref ref-type="bibr" rid="ref-148">148</xref>] integrated the active learning mechanism with GAN to generate virtual samples, building an accurate and robust DXN warning model. In reference [<xref ref-type="bibr" rid="ref-132">132</xref>], fuzzy set theory was incorporated into the GAN framework to address uncertainty and facilitate the generation of high-quality samples.</p>
<p>Building on this foundation, the applications of data augmentation techniques in representative industrial processes are summarized in <xref ref-type="table" rid="table-4">Table 4</xref>. From a managerial perspective, practitioners can make informed decisions regarding model deployment by applying data augmentation techniques to specific data regimes and industrial conditions, as well as guiding data acquisition planning and digital strategies. This survey is intended to serve as a comprehensive reference for mapping data-related challenges to suitable data augmentation strategies within a range of industrial constraints.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Applications of data augmentation techniques in industrial processes</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Industrial process</th>
<th>Data characteristic</th>
<th>Methodology</th>
<th>Reference</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Chemical process</td>
<td rowspan="3">SSHD</td>
<td>Statistical technique</td>
<td>[<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-131">131</xref>]</td>
</tr>
<tr>
<td>Deep generative model</td>
<td>[<xref ref-type="bibr" rid="ref-68">68</xref>,<xref ref-type="bibr" rid="ref-69">69</xref>,<xref ref-type="bibr" rid="ref-73">73</xref>,<xref ref-type="bibr" rid="ref-74">74</xref>]</td>
</tr>
<tr>
<td>Transfer learning</td>
<td>[<xref ref-type="bibr" rid="ref-75">75</xref>&#x2013;<xref ref-type="bibr" rid="ref-77">77</xref>]</td>
</tr>
<tr>
<td rowspan="2">SSLD</td>
<td>Statistical technique</td>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-86">86</xref>]</td>
</tr>
<tr>
<td>Deep generative model</td>
<td>[<xref ref-type="bibr" rid="ref-91">91</xref>,<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
</tr>
<tr>
<td rowspan="4">Rotating machinery</td>
<td rowspan="4">LSLD</td>
<td>Transform technique</td>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
</tr>
<tr>
<td>Deep generative model</td>
<td>[<xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>&#x2013;<xref ref-type="bibr" rid="ref-43">43</xref>,<xref ref-type="bibr" rid="ref-123">123</xref>&#x2013;<xref ref-type="bibr" rid="ref-127">127</xref>]</td>
</tr>
<tr>
<td>Transfer learning</td>
<td>[<xref ref-type="bibr" rid="ref-45">45</xref>&#x2013;<xref ref-type="bibr" rid="ref-49">49</xref>,<xref ref-type="bibr" rid="ref-128">128</xref>]</td>
</tr>
<tr>
<td>Physical model</td>
<td>[<xref ref-type="bibr" rid="ref-50">50</xref>&#x2013;<xref ref-type="bibr" rid="ref-54">54</xref>]</td>
</tr>
<tr>
<td rowspan="2">MSWI process</td>
<td rowspan="2">SSHD</td>
<td>Statistical technique</td>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>,<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-147">147</xref>,<xref ref-type="bibr" rid="ref-148">148</xref>]</td>
</tr>
<tr>
<td>Deep generative model</td>
<td>[<xref ref-type="bibr" rid="ref-132">132</xref>,<xref ref-type="bibr" rid="ref-148">148</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion and Analysis</title>
<p>This review has examined four categories of data characteristics&#x2014;SSLD, SSHD, LSLD, and LSHD&#x2014;along with their corresponding data augmentation strategies. Despite notable advancements, several critical limitations remain in existing approaches. Drawing on these insights, we propose four key directions for future research:</p>
<p><bold>(1) <italic>Data Augmentation for Few-Shot Scenarios:</italic></bold> In actual industrial scenarios, safety constraints often restrict data collection under extreme or failure conditions, resulting in long-tailed distributions and limited model generalization. Traditional generative models struggle to capture these rare dynamics, particularly in systems characterized by nonlinear and multi-physics behaviors. Future research should focus on developing knowledge-guided or physics-informed augmentation strategies capable of simulating rare events with greater fidelity, while ensuring physical plausibility and adherence to safety constraints.</p>
<p><bold>(2) <italic>Cooperative Multi-Modal Data Augmentation:</italic></bold> The heterogeneity and inconsistency inherent in multi-modal data (e.g., sensor signals, images, and logs) pose significant challenges. Existing augmentation techniques often overlook cross-modal dependencies, increasing the risk of semantic misalignment. Advancing cooperative augmentation methods that align representations across modalities while preserving domain-specific semantics remains a critical need, particularly for robust modeling in data-scarce scenarios.</p>
<p><bold>(3) <italic>Evaluation and Closed-Loop Optimization of Generated Samples:</italic></bold> The absence of rigorous evaluation criteria undermines confidence in generated virtual samples. Furthermore, few existing methods incorporate feedback loops for iterative refinement of data augmentation. There is a critical need for closed-loop frameworks that integrate expert constraints, dynamic evaluation metrics, and system feedback to continuously validate and enhance generated data, particularly in safety-critical applications.</p>
<p><bold>(4) <italic>ChatGPT to Data Conversion Augmentation:</italic></bold> Recent advances in large language models (LLMs), such as ChatGPT, have enabled the transformation of unstructured expert knowledge into structured data, opening new avenues for data augmentation. However, significant challenges persist in semantic grounding, domain adaptation, and reliability of LLM-generated data. Future research should investigate hybrid frameworks that integrate LLM-based generation with domain-specific rules to enhance interpretability and facilitate knowledge integration in data-driven modeling.</p>
<p>Current data augmentation techniques still face challenges in generalizing to operating conditions, handling data heterogeneity, and maintaining interpretability. To fill these gaps, future research should pursue hybrid, domain-informed, and feedback-integrated frameworks that leverage both data-driven and knowledge-based approaches.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>Data augmentation plays a vital role in improving the performance and robustness of artificial intelligence (AI) models in industrial applications, particularly under conditions of data scarcity and imbalance. This survey reviews data augmentation techniques on four representative data characteristics and five methodological categories, emphasizing their applications and limitations. Despite recent advancements, significant challenges remain in cross-domain generalization, multi-modal data fusion, evaluation of generated data, and the integration of domain knowledge. Future research should prioritize the development of hybrid augmentation frameworks, the enforcement of cross-modal consistency, feedback-driven validation mechanisms, and the application of LLM to integrate expert knowledge with data-driven modeling. Addressing these challenges is essential for the advancement of reliable and adaptable industrial AI systems.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the Postdoctoral Fellowship Program (Grade B) of China (GZB20250435) and the National Natural Science Foundation of China (62403270).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Study conception and design: Canlin Cui, Heng Xia; Literature collection and analysis: Canlin Cui; Draft manuscript preparation: Canlin Cui, Junyu Yao, Heng Xia; Funding acquisition and supervision: Heng Xia, Junyu Yao. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Augmented industrial data-driven modeling under the curse of dimensionality</article-title>. <source>IEEE/CAA J Automat Sinica</source>. <year>2023</year>;<volume>10</volume>(<issue>6</issue>):<fpage>1445</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jas.2023.123396</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>CI</given-names></string-name></person-group>. <article-title>Nonparametric control limits incorporating exceedance probability criterion for statistical process monitoring with commonly employed small to moderate sample sizes</article-title>. <source>Quality Eng</source>. <year>2025</year>;<volume>37</volume>(<issue>1</issue>):<fpage>145</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1080/08982112.2024.2363828</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Progressively balanced supervised contrastive representation learning for long-tailed fault diagnosis</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2022</year>;<volume>71</volume>(<issue>1</issue>):<fpage>3506112</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2022.3151946</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A novel data augmentation approach to fault diagnosis with class-imbalance problem</article-title>. <source>Reliab Eng Syst Saf</source>. <year>2024</year>;<volume>243</volume>:<fpage>109832</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ress.2023.109832</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xia</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aljerf</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ukaogo</surname> <given-names>PO</given-names></string-name></person-group>. <article-title>Dioxin emission modeling using feature selection and simplified DFR with residual error fitting for the grate-based MSWI process</article-title>. <source>Waste Manage</source>. <year>2023</year>;<volume>168</volume>(<issue>4</issue>):<fpage>256</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.wasman.2023.05.056</pub-id>; <pub-id pub-id-type="pmid">37327519</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bayer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kaufhold</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Reuter</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A survey on data augmentation for text classification</article-title>. <source>ACM Comput Surv</source>. <year>2022</year>;<volume>55</volume>(<issue>7</issue>):<fpage>1</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3544558</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garcea</surname> <given-names>F</given-names></string-name>, <string-name><surname>Serra</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lamberti</surname> <given-names>F</given-names></string-name>, <string-name><surname>Morra</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Data augmentation for medical imaging: a systematic literature review</article-title>. <source>Comput Biol Med</source>. <year>2023</year>;<volume>152</volume>(<issue>1</issue>):<fpage>106391</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.106391</pub-id>; <pub-id pub-id-type="pmid">36549032</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>F</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A survey of mix-based data augmentation: taxonomy, methods, applications, and explainability</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>57</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3696206</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ju</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ju</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A systematic review of data augmentation methods for intelligent fault diagnosis of rotating machinery under limited data conditions</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>35</volume>(<issue>12</issue>):<fpage>122004</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad7a97</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Schwarz</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rahal</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Sahelices</surname> <given-names>B</given-names></string-name>, <string-name><surname>Barroso-Garc&#x00ED;a</surname> <given-names>V</given-names></string-name>, <string-name><surname>Weis</surname> <given-names>R</given-names></string-name>, <string-name><surname>Duque Ant&#x00F3;n</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Data augmentation in predictive maintenance applicable to hydrogen combustion engines: a review</article-title>. <source>Artif Intell Rev</source>. <year>2025</year>;<volume>58</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-024-11021-9</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Raudys</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>AK</given-names></string-name></person-group>. <article-title>Small sample size effects in statistical pattern recognition: recommendations for practitioners</article-title>. <source>IEEE Transact Pattern Anal Mach Intell</source>. <year>1991</year>;<volume>13</volume>(<issue>3</issue>):<fpage>252</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1109/34.75512</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shawe-Taylor</surname> <given-names>J</given-names></string-name>, <string-name><surname>Anthony</surname> <given-names>M</given-names></string-name>, <string-name><surname>Biggs</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Bounding sample size with the Vapnik-Chervonenkis dimension</article-title>. <source>Discrete Appl Mathem</source>. <year>1993</year>;<volume>42</volume>(<issue>1</issue>):<fpage>65</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0166-218x(93)90179-r</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Application of small sample virtual expansion and spherical mapping model in wind turbine fault diagnosis</article-title>. <source>Expert Syst Appl</source>. <year>2021</year>;<volume>183</volume>(<issue>14</issue>):<fpage>115397</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2021.115397</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Fine</surname> <given-names>TL</given-names></string-name></person-group>. <article-title>Neural-network design for small training sets of high dimension</article-title>. <source>IEEE Transact Neu Netw</source>. <year>1998</year>;<volume>9</volume>(<issue>2</issue>):<fpage>266</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/72.661122</pub-id>; <pub-id pub-id-type="pmid">18252451</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>B</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A PSO based virtual sample generation method for small sample sets: applications to regression datasets</article-title>. <source>Eng Appl Artif Intell</source>. <year>2017</year>;<volume>59</volume>:<fpage>236</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2016.12.024</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A novel and effective nonlinear interpolation virtual sample generation method for enhancing energy prediction and analysis on small data problem: a case study of Ethylene industry</article-title>. <source>Energy</source>. <year>2018</year>;<volume>147</volume>(<issue>1</issue>):<fpage>418</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.energy.2018.01.059</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Novel manifold learning based virtual sample generation for optimizing soft sensor with small data</article-title>. <source>ISA Trans</source>. <year>2021</year>;<volume>109</volume>(<issue>5</issue>):<fpage>229</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.isatra.2020.10.006</pub-id>; <pub-id pub-id-type="pmid">33070985</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Modeling high dimensional frequency spectral data based on virtual sample generation technique</article-title>. In: <conf-name>2015 IEEE International Conference on Information and Automation; 2015 Aug 8&#x2013;10</conf-name>; <publisher-loc>Lijiang, China</publisher-loc>. p. <fpage>1090</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>WA</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>DJ</given-names></string-name></person-group>. <article-title>ALADA: a lite automatic data augmentation framework for industrial defect detection</article-title>. <source>Adv Eng Inform</source>. <year>2023</year>;<volume>58</volume>(<issue>2</issue>):<fpage>102205</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2023.102205</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Altman</surname> <given-names>N</given-names></string-name>, <string-name><surname>Krzywinski</surname> <given-names>M</given-names></string-name></person-group>. <article-title>The curse(s) of dimensionality</article-title>. <source>Nat Methods</source>. <year>2018</year>;<volume>15</volume>(<issue>6</issue>):<fpage>399</fpage>&#x2013;<lpage>400</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41592-018-0019-x</pub-id>; <pub-id pub-id-type="pmid">29855577</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Intelligent rotating machinery fault diagnosis based on deep learning using data augmentation</article-title>. <source>J Intell Manufact</source>. <year>2020</year>;<volume>31</volume>(<issue>2</issue>):<fpage>433</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10845-018-1456-1</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Fault feature extractor based on bootstrap your own latent and data augmentation algorithm for unlabeled vibration signals</article-title>. <source>IEEE Transact Indust Elect</source>. <year>2022</year>;<volume>69</volume>(<issue>9</issue>):<fpage>9547</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tie.2021.3111567</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abeysinghe</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tohmuang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Davy</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Fard</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Data augmentation on convolutional neural networks to classify mechanical noise</article-title>. <source>Appl Acoust</source>. <year>2023</year>;<volume>203</volume>(<issue>16</issue>):<fpage>109209</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apacoust.2023.109209</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nematirad</surname> <given-names>R</given-names></string-name>, <string-name><surname>Behrang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Pahwa</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Acoustic-based online monitoring of cooling fan malfunction in air-forced transformers using learning techniques</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>26384</fpage>&#x2013;<lpage>400</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3366807</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alharbi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wheeler</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Belt conveyor idlers fault detection using acoustic analysis and deep learning algorithm with the YAMNet pretrained network</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>19</issue>):<fpage>31379</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2024.3439509</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Self-supervised simple siamese framework for fault diagnosis of rotating machinery with unlabeled samples</article-title>. <source>IEEE Transact Neural Netw Learn Syst</source>. <year>2024</year>;<volume>35</volume>(<issue>5</issue>):<fpage>6380</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2022.3209332</pub-id>; <pub-id pub-id-type="pmid">36197866</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Russell</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jawahir</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Mixed-up experience replay for adaptive online condition monitoring</article-title>. <source>IEEE Transact Indust Elect</source>. <year>2024</year>;<volume>71</volume>(<issue>2</issue>):<fpage>1979</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tie.2023.3260351</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Domain knowledge-guided contrastive learning framework based on complementary views for fault diagnosis with limited labeled data</article-title>. <source>IEEE Transact Indust Inform</source>. <year>2024</year>;<volume>20</volume>(<issue>5</issue>):<fpage>8055</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2024.3369704</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cabrera</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cerrada</surname> <given-names>M</given-names></string-name>, <string-name><surname>S&#x00E1;nchez</surname> <given-names>RV</given-names></string-name>, <string-name><surname>Sancho</surname> <given-names>F</given-names></string-name>, <string-name><surname>Estupinan</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Fault diagnosis generalization improvement through contrastive learning for a multistage centrifugal pump</article-title>. <source>IEEE Transact Reliab</source>. <year>2025</year>;<volume>74</volume>(<issue>7</issue>):<fpage>2373</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2024.3381014</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A deep convolutional neural network based fusion method of two-direction vibration signal data for health state identification of planetary gearboxes</article-title>. <source>Measurement</source>. <year>2019</year>;<volume>146</volume>(<issue>7553</issue>):<fpage>268</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2019.04.093</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Enhanced data-driven fault diagnosis for machines with small and unbalanced data based on variational auto-encoder</article-title>. <source>Meas Sci Technol</source>. <year>2020</year>;<volume>31</volume>(<issue>3</issue>):<fpage>035004</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ab55f8</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dixit</surname> <given-names>S</given-names></string-name>, <string-name><surname>Verma</surname> <given-names>NK</given-names></string-name></person-group>. <article-title>Intelligent condition-based monitoring of rotary machines with few samples</article-title>. <source>IEEE Sens J</source>. <year>2020</year>;<volume>20</volume>(<issue>23</issue>):<fpage>14337</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2020.3008177</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>K</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A new data generation approach with modified Wasserstein auto-encoder for rotating machinery fault diagnosis with limited fault data</article-title>. <source>Knowl Based Syst</source>. <year>2022</year>;<volume>238</volume>(<issue>1</issue>):<fpage>107892</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2021.107892</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A semi-supervised Gaussian mixture variational autoencoder method for few-shot fine-grained fault diagnosis</article-title>. <source>Neural Netw</source>. <year>2024</year>;<volume>178</volume>(<issue>8</issue>):<fpage>106482</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2024.106482</pub-id>; <pub-id pub-id-type="pmid">38945116</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karamti</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lashin</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Alrowais</surname> <given-names>FM</given-names></string-name>, <string-name><surname>Mahmoud</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>A new deep stacked architecture for multi-fault machinery identification with imbalanced samples</article-title>. <source>IEEE Access</source>. <year>2021</year>;<volume>9</volume>:<fpage>58838</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2021.3071796</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Wasserstein generative adversarial network and convolutional neural network (WG-CNN) for bearing fault diagnosis</article-title>. <source>Math Probl Eng</source>. <year>2020</year>;<volume>2020</volume>(<issue>1</issue>):<fpage>2604191</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2020/2604191</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A novel bearing imbalance Fault-diagnosis method based on a Wasserstein conditional generative adversarial network</article-title>. <source>Measurement</source>. <year>2022</year>;<volume>192</volume>(<issue>6</issue>):<fpage>110924</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2022.110924</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jalayer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kaboli</surname> <given-names>A</given-names></string-name>, <string-name><surname>Orsenigo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vercellis</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Fault detection and diagnosis with imbalanced and noisy data: a hybrid framework for rotating machinery</article-title>. <source>Machines</source>. <year>2022</year>;<volume>10</volume>(<issue>4</issue>):<fpage>237</fpage>. doi:<pub-id pub-id-type="doi">10.3390/machines10040237</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>An imbalanced data augmentation and assessment method for industrial process fault classification with application in air compressors</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>3521510</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2023.3288257</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>ACWGAN-GP for milling tool breakage monitoring with imbalanced data</article-title>. <source>Robot Comput Integr Manuf</source>. <year>2024</year>;<volume>85</volume>:<fpage>102624</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.rcim.2023.102624</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Small-sample-oriented multi-condition fault diagnosis framework based on classifier-free denoising diffusion implicit model with multi-class contrastive learning</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>(<issue>24</issue>):<fpage>41635</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2024.3487209</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Mei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>F</given-names></string-name></person-group>. <article-title>A novel data augmentation method based on denoising diffusion probabilistic model for fault diagnosis under imbalanced data</article-title>. <source>IEEE Transact Indust Inform</source>. <year>2024</year>;<volume>20</volume>(<issue>5</issue>):<fpage>7820</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2024.3366991</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic model-enabled data augmentation method for intelligent machine fault diagnosis</article-title>. <source>Eng Appl Artif Intell</source>. <year>2025</year>;<volume>139</volume>(<issue>4</issue>):<fpage>109520</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.109520</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A novel lightweight DDPM-based data augmentation method for rotating machinery fault diagnosis with small sample</article-title>. <source>Mech Syst Signal Process</source>. <year>2025</year>;<volume>232</volume>:<fpage>112741</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2025.112741</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Adaptive weighted generative adversarial network with attention mechanism: a transfer data augmentation method for tool wear prediction</article-title>. <source>Mech Syst Signal Process</source>. <year>2024</year>;<volume>212</volume>(<issue>3&#x2013;4</issue>):<fpage>111288</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2024.111288</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ge</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A new multiple mixed augmentation-based transfer learning method for machinery fault diagnosis</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>35</volume>(<issue>8</issue>):<fpage>086141</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad4d15</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jian</surname> <given-names>C</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mo</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Open-set domain generalization for fault diagnosis through data augmentation and a dual-level weighted mechanism</article-title>. <source>Adv Eng Inform</source>. <year>2024</year>;<volume>62</volume>:<fpage>102703</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2024.102703</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A task-oriented theil index-based meta-learning network with gradient calibration strategy for rotating machinery fault diagnosis with limited samples</article-title>. <source>Adv Eng Inform</source>. <year>2024</year>;<volume>62</volume>(<issue>10</issue>):<fpage>102870</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2024.102870</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A dynamic collaborative adversarial domain adaptation network for unsupervised rotating machinery fault diagnosis</article-title>. <source>Reliab Eng Syst Saf</source>. <year>2025</year>;<volume>255</volume>:<fpage>110662</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ress.2024.110662</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lim</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jung</surname> <given-names>W</given-names></string-name>, <string-name><surname>Bae</surname> <given-names>J</given-names></string-name>, <string-name><surname>Park</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Utilization of high-fidelity simulation data for data augmentation of artificial neural net-based rotor faults diagnosis</chapter-title>. In: <source>Active and passive smart structures and integrated systems XVI</source>. Vol. <volume>12043</volume>. <publisher-loc>Bellingham, WA, USA</publisher-loc>: <publisher-name>SPIE</publisher-name>; <year>2022</year>. p. <fpage>441</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Fault diagnosis of controllable pitch propeller as few-shot classification with mechanism simulation data augmentation</article-title>. In: <conf-name>2023 IEEE 2nd Industrial Electronics Society Annual On-Line Conference (ONCON); 2023 Dec 8&#x2013;10; SC, USA</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A novel fault diagnosis method for denoising autoencoder assisted by digital twin</article-title>. <source>Comput Intell Neurosci</source>. <year>2022</year>;<volume>2022</volume>(<issue>1</issue>):<fpage>5077134</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2022/5077134</pub-id>; <pub-id pub-id-type="pmid">35909837</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Faulty rolling bearing digital twin model and its application in fault diagnosis with imbalanced samples</article-title>. <source>Adv Eng Inform</source>. <year>2024</year>;<volume>61</volume>(<issue>6</issue>):<fpage>102513</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2024.102513</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Contrastive learning-enabled digital twin framework for fault diagnosis of rolling bearing</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>36</volume>(<issue>1</issue>):<fpage>015026</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad8f52</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ming</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Digital twin-assisted fault diagnosis framework for rolling bearings under imbalanced data</article-title>. <source>Appl Soft Comput</source>. <year>2025</year>;<volume>168</volume>:<fpage>112528</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112528</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Virtual sample generation method based on improved megatrend diffusion and hidden layer interpolation with its application</article-title>. <source>CIESC J</source>. <year>2020</year>;<volume>71</volume>(<issue>12</issue>):<fpage>5681</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Damarla</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A Gaussian mixture model based virtual sample generation approach for small datasets in industrial processes</article-title>. <source>Inform Sci</source>. <year>2021</year>;<volume>581</volume>(<issue>4</issue>):<fpage>262</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ins.2021.09.014</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Prediction of dioxin emission concentration in the municipal solid waste incineration process based on optimal selection of virtual samples</article-title>. <source>J Beijing Univ Technol</source>. <year>2021</year>;<volume>47</volume>(<issue>5</issue>):<fpage>431</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ccdc52312.2021.9601628</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Virtual sample generation method using reduced feature probability density distribution</article-title>. <source>Control Theory Appl</source>. <year>2024</year>;<volume>41</volume>(<issue>11</issue>):<fpage>2165</fpage>&#x2013;<lpage>73</lpage>. <comment>(In Chinese)</comment>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel virtual sample generation based on locally linear embedding for optimizing the small sample problem: case of soft sensor applications</article-title>. <source>Indus Eng Chem Res</source>. <year>2020</year>;<volume>59</volume>(<issue>40</issue>):<fpage>17977</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.iecr.0c01942</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hua</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Enhanced virtual sample generation based on manifold features: applications to developing soft sensor using small data</article-title>. <source>ISA Transact</source>. <year>2022</year>;<volume>126</volume>(<issue>4</issue>):<fpage>398</fpage>&#x2013;<lpage>406</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.isatra.2021.07.033</pub-id>; <pub-id pub-id-type="pmid">34334185</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Novel discriminant locality preserving projection integrated with Monte Carlo sampling for fault diagnosis</article-title>. <source>IEEE Transact Reliab</source>. <year>2021</year>;<volume>72</volume>(<issue>1</issue>):<fpage>166</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2021.3115108</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel space projection interpolation based virtual sample generation for solving the small data problem in developing soft sensor</article-title>. <source>Chemometr Intell Lab Syst</source>. <year>2021</year>;<volume>217</volume>(<issue>35</issue>):<fpage>104425</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.chemolab.2021.104425</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Co-training based virtual sample generation for solving the small sample size problem in process industry</article-title>. <source>ISA Transact</source>. <year>2023</year>;<volume>134</volume>(<issue>C</issue>):<fpage>290</fpage>&#x2013;<lpage>301</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.isatra.2022.08.021</pub-id>; <pub-id pub-id-type="pmid">36064497</pub-id></mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Quan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>From regression to classification: fuzzy multikernel subspace learning for robust prediction and drug screening</article-title>. <source>IEEE Transact Indust Inform</source>. <year>2024</year>;<volume>20</volume>(<issue>3</issue>):<fpage>4137</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2023.3321332</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ou</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gui</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A layer-wise data augmentation strategy for deep learning networks and its soft sensor application in an industrial hydrocracking process</article-title>. <source>IEEE Transact Neural Netw Learn Syst</source>. <year>2021</year>;<volume>32</volume>(<issue>8</issue>):<fpage>3296</fpage>&#x2013;<lpage>305</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2019.2951708</pub-id>; <pub-id pub-id-type="pmid">31841424</pub-id></mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Improving the performance of just-in-time learning-based soft sensor through data augmentation</article-title>. <source>IEEE Transact Indust Elect</source>. <year>2022</year>;<volume>69</volume>(<issue>12</issue>):<fpage>13716</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tie.2021.3139194</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>DeepSMOTE with Laplacian matrix decomposition for imbalance instance fault diagnosis</article-title>. <source>Chemometr Intell Lab Syst</source>. <year>2025</year>;<volume>259</volume>(<issue>3</issue>):<fpage>105338</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.chemolab.2025.105338</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>W</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Virtual sample generation for soft-sensing in small sample scenarios using glow-embedded variational autoencoder</article-title>. <source>Comput Chem Eng</source>. <year>2025</year>;<volume>193</volume>(<issue>1</issue>):<fpage>108925</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compchemeng.2024.108925</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Data augmentation in fault diagnosis based on the Wasserstein generative adversarial network with gradient penalty</article-title>. <source>Neurocomputing</source>. <year>2020</year>;<volume>396</volume>(<issue>99</issue>):<fpage>487</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2018.10.109</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Augmented multidimensional convolutional neural network for industrial soft sensing</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2021</year>;<volume>70</volume>:<fpage>2508410</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2021.3075515</pub-id>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>A novel virtual sample generation method based on a modified conditional Wasserstein GAN to address the small sample size problem in soft sensing</article-title>. <source>J Process Cont</source>. <year>2022</year>;<volume>113</volume>(<issue>4</issue>):<fpage>18</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jprocont.2022.03.008</pub-id>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Data augmentation classifier for imbalanced fault classification</article-title>. <source>IEEE Transact Automat Sci Eng</source>. <year>2021</year>;<volume>18</volume>(<issue>3</issue>):<fpage>1206</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tase.2020.2998467</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhuo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Auxiliary information-guided industrial data augmentation for any-shot fault learning and diagnosis</article-title>. <source>IEEE Transact Indust Inform</source>. <year>2021</year>;<volume>17</volume>(<issue>11</issue>):<fpage>7535</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2021.3053106</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>LJDA-net: a low-rank joint domain adaptation network for industrial sample enhancement</article-title>. <source>IEEE Sens J</source>. <year>2022</year>;<volume>22</volume>(<issue>12</issue>):<fpage>11881</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2022.3170085</pub-id>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A model transfer learning based fault diagnosis method for chemical processes with small samples</article-title>. <source>Int J Control Autom Syst</source>. <year>2023</year>;<volume>21</volume>(<issue>12</issue>):<fpage>4080</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12555-022-0798-9</pub-id>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Atoui</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Dual adversarial and contrastive network for single-source domain generalization in fault diagnosis</article-title>. <source>Adv Eng Inform</source>. <year>2025</year>;<volume>65</volume>(<issue>3</issue>):<fpage>103140</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2025.103140</pub-id>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>DC</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>YS</given-names></string-name></person-group>. <article-title>Learning management knowledge for manufacturing systems in the early stages using time series data</article-title>. <source>European J Operat Res</source>. <year>2008</year>;<volume>184</volume>(<issue>1</issue>):<fpage>169</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ejor.2006.10.008</pub-id>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>DC</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>LS</given-names></string-name></person-group>. <article-title>A new approach to assess product lifetime performance for small data sets</article-title>. <source>European J Operat Res</source>. <year>2013</year>;<volume>230</volume>(<issue>2</issue>):<fpage>290</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ejor.2013.04.016</pub-id>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>DC</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>WC</given-names></string-name></person-group>. <article-title>A novel data transformation model for small data-set learning</article-title>. <source>Int J Product Res</source>. <year>2016</year>;<volume>54</volume>(<issue>24</issue>):<fpage>7453</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1080/00207543.2016.1192301</pub-id>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Small data-driven modeling of forming force in single point incremental forming using neural networks</article-title>. <source>Eng Comput</source>. <year>2020</year>;<volume>36</volume>(<issue>4</issue>):<fpage>1589</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00366-019-00781-6</pub-id>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sivakumar</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ramamurthy</surname> <given-names>K</given-names></string-name>, <string-name><surname>Radhakrishnan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Won</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Synthetic sampling from small datasets: a modified mega-trend diffusion approach using k-nearest neighbors</article-title>. <source>Knowl Based Syst</source>. <year>2022</year>;<volume>236</volume>(<issue>4</issue>):<fpage>107687</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2021.107687</pub-id>.</mixed-citation></ref>
<ref id="ref-83"><label>[83]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mathew</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>CK</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>M</given-names></string-name>, <string-name><surname>Leong</surname> <given-names>WH</given-names></string-name></person-group>. <article-title>Classification of imbalanced data by oversampling in kernel space of support vector machines</article-title>. <source>IEEE Transact Neural Netw Learn Syst</source>. <year>2018</year>;<volume>29</volume>(<issue>9</issue>):<fpage>4065</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2017.2751612</pub-id>; <pub-id pub-id-type="pmid">29028213</pub-id></mixed-citation></ref>
<ref id="ref-84"><label>[84]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Soltanzadeh</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hashemzadeh</surname> <given-names>M</given-names></string-name></person-group>. <article-title>RCSMOTE: range-Controlled synthetic minority over-sampling technique for handling the class imbalance problem</article-title>. <source>Inform Sci</source>. <year>2021</year>;<volume>542</volume>:<fpage>92</fpage>&#x2013;<lpage>111</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ins.2020.07.014</pub-id>.</mixed-citation></ref>
<ref id="ref-85"><label>[85]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Augmenting the diversity of imbalanced datasets via multi-vector stochastic exploration oversampling</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>583</volume>(<issue>4</issue>):<fpage>127600</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2024.127600</pub-id>.</mixed-citation></ref>
<ref id="ref-86"><label>[86]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Rajabifard</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Dealing with small sample size problems in process industry using virtual sample generation: a Kriging-based approach</article-title>. <source>Soft Comput</source>. <year>2020</year>;<volume>24</volume>(<issue>9</issue>):<fpage>6889</fpage>&#x2013;<lpage>902</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00500-019-04326-3</pub-id>.</mixed-citation></ref>
<ref id="ref-87"><label>[87]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel virtual sample generation method based on data augmentation and weighted interpolation for soft sensing with small data</article-title>. <source>Expert Syst Appl</source>. <year>2023</year>;<volume>225</volume>:<fpage>120085</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.120085</pub-id>.</mixed-citation></ref>
<ref id="ref-88"><label>[88]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>CB</given-names></string-name></person-group>. <article-title>Augmented time regularized generative adversarial network (atr-gan) for data augmentation in online process anomaly detection</article-title>. <source>IEEE Transact Autom Sci Eng</source>. <year>2022</year>;<volume>19</volume>(<issue>4</issue>):<fpage>3338</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tase.2021.3118635</pub-id>.</mixed-citation></ref>
<ref id="ref-89"><label>[89]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Attention-stacked generative adversarial network (AS-GAN)-empowered sensor data augmentation for online monitoring of manufacturing system</article-title>. <comment>arXiv:2306.06268. 2023</comment>.</mixed-citation></ref>
<ref id="ref-90"><label>[90]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>SKS</given-names></string-name>, <string-name><surname>Tsai</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Yeh</surname> <given-names>PC</given-names></string-name></person-group>. <article-title>Effective variational-autoencoder-based generative models for highly imbalanced fault detection data in semiconductor manufacturing</article-title>. <source>IEEE Transact Semicond Manufact</source>. <year>2023</year>;<volume>36</volume>(<issue>2</issue>):<fpage>205</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsm.2023.3238555</pub-id>.</mixed-citation></ref>
<ref id="ref-91"><label>[91]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel virtual sample generation using target-relevant autoencoder for small data-based soft sensor</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2021</year>;<volume>70</volume>:<fpage>2515910</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2021.3120135</pub-id>.</mixed-citation></ref>
<ref id="ref-92"><label>[92]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel virtual sample generation using conditional GAN for developing soft sensor with small data</article-title>. <source>Eng Appl Artif Intell</source>. <year>2021</year>;<volume>106</volume>(<issue>2</issue>):<fpage>104497</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2021.104497</pub-id>.</mixed-citation></ref>
<ref id="ref-93"><label>[93]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A soft sensor regression model for complex chemical process based on generative adversarial nets and vine copula</article-title>. <source>J Taiwan Inst Chem Eng</source>. <year>2022</year>;<volume>138</volume>:<fpage>104483</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jtice.2022.104483</pub-id>.</mixed-citation></ref>
<ref id="ref-94"><label>[94]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Soft sensor modeling for small data scenarios based on data enhancement and selective ensemble</article-title>. <source>Chem Eng Sci</source>. <year>2023</year>;<volume>279</volume>(<issue>12</issue>):<fpage>118958</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ces.2023.118958</pub-id>.</mixed-citation></ref>
<ref id="ref-95"><label>[95]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Data augmentation using time conditional variational autoencoder for soft sensor of industrial processes with limited data</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2024</year>;<volume>73</volume>(<issue>1</issue>):<fpage>2524714</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2024.3427765</pub-id>.</mixed-citation></ref>
<ref id="ref-96"><label>[96]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Antypenko</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sushko</surname> <given-names>I</given-names></string-name>, <string-name><surname>Zakharchenko</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Intrusion detection system after data augmentation schemes based on the VAE and CVAE</article-title>. <source>IEEE Transact Reliab</source>. <year>2022</year>;<volume>71</volume>(<issue>2</issue>):<fpage>1000</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2022.3164877</pub-id>.</mixed-citation></ref>
<ref id="ref-97"><label>[97]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Distribution bias aware collaborative generative adversarial network for imbalanced deep learning in industrial IoT</article-title>. <source>IEEE Transact Indust Inform</source>. <year>2023</year>;<volume>19</volume>(<issue>1</issue>):<fpage>570</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2022.3170149</pub-id>.</mixed-citation></ref>
<ref id="ref-98"><label>[98]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhuang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Ensemble data augmentation for imbalanced fault diagnosis</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>3528312</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2023.3307757</pub-id>.</mixed-citation></ref>
<ref id="ref-99"><label>[99]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zaman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Upadhyay</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lung</surname> <given-names>CH</given-names></string-name></person-group>. <article-title>Validation of a machine learning-based IDS design framework using ORNL datasets for power system with SCADA</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>118414</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2023.3326751</pub-id>.</mixed-citation></ref>
<ref id="ref-100"><label>[100]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>CGFL: a robust federated learning approach for intrusion detection systems based on data generation</article-title>. <source>Appl Sci</source>. <year>2025</year>;<volume>15</volume>(<issue>5</issue>):<fpage>2416</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app15052416</pub-id>.</mixed-citation></ref>
<ref id="ref-101"><label>[101]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multi-class data augmentation and fault diagnosis of wind turbine blades based on ISOMAP-CGAN under high-dimensional imbalanced samples</article-title>. <source>Renew Energy</source>. <year>2025</year>;<volume>243</volume>(<issue>1</issue>):<fpage>122609</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.renene.2025.122609</pub-id>.</mixed-citation></ref>
<ref id="ref-102"><label>[102]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Digital twin-driven online anomaly detection for an automation system based on edge intelligence</article-title>. <source>J Manufact Syst</source>. <year>2021</year>;<volume>59</volume>:<fpage>138</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jmsy.2021.02.010</pub-id>.</mixed-citation></ref>
<ref id="ref-103"><label>[103]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Krespach</surname> <given-names>V</given-names></string-name>, <string-name><surname>Blum</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pottmann</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rehfeldt</surname> <given-names>S</given-names></string-name>, <string-name><surname>Klein</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Improving extrapolation capabilities of a data-driven prediction model for control of an air separation unit</article-title>. <source>Comput Chem Eng</source>. <year>2025</year>;<volume>194</volume>:<fpage>108953</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compchemeng.2024.108953</pub-id>.</mixed-citation></ref>
<ref id="ref-104"><label>[104]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Efficient fine-tuned preventive monitoring models of bearing failures without prior on-site fault data</article-title>. <source>Measurement</source>. <year>2025</year>;<volume>242</volume>(<issue>5</issue>):<fpage>116067</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2024.116067</pub-id>.</mixed-citation></ref>
<ref id="ref-105"><label>[105]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bond-Taylor</surname> <given-names>S</given-names></string-name>, <string-name><surname>Leach</surname> <given-names>A</given-names></string-name>, <string-name><surname>Long</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Willcocks</surname> <given-names>CG</given-names></string-name></person-group>. <article-title>Deep generative modelling: a comparative review of VAEs, GANs, normalizing flows, energy-based and autoregressive models</article-title>. <source>IEEE Transact Pattern Analy Mach Intell</source>. <year>2021</year>;<volume>44</volume>(<issue>11</issue>):<fpage>7327</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2021.3116668</pub-id>; <pub-id pub-id-type="pmid">34591756</pub-id></mixed-citation></ref>
<ref id="ref-106"><label>[106]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Application of generative adversarial networks for intelligent fault diagnosis</article-title>. In: <conf-name>2018 IEEE 14th International Conference on Automation Science and Engineering (CASE); 2018 Aug 20&#x2013;24</conf-name>; <publisher-loc>Munich, Germany</publisher-loc>. p. <fpage>711</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-107"><label>[107]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>T</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CBAM-CRLSGAN: a novel fault diagnosis method for planetary transmission systems under small samples scenarios</article-title>. <source>Measurement</source>. <year>2024</year>;<volume>234</volume>:<fpage>114795</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2024.114795</pub-id>.</mixed-citation></ref>
<ref id="ref-108"><label>[108]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Data augmentation and intelligent fault diagnosis of planetary gearbox using ILoFGAN under extremely limited samples</article-title>. <source>IEEE Transact Reliab</source>. <year>2023</year>;<volume>72</volume>(<issue>3</issue>):<fpage>1029</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2022.3215243</pub-id>.</mixed-citation></ref>
<ref id="ref-109"><label>[109]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>MCBA-MVACGAN: a novel fault diagnosis method for rotating machinery under small sample conditions</article-title>. <source>Machines</source>. <year>2025</year>;<volume>13</volume>(<issue>1</issue>):<fpage>71</fpage>. doi:<pub-id pub-id-type="doi">10.3390/machines13010071</pub-id>.</mixed-citation></ref>
<ref id="ref-110"><label>[110]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A novel intelligent fault diagnosis method of helical gear with multi-channel information fused images under small samples</article-title>. <source>Appl Acoust</source>. <year>2025</year>;<volume>228</volume>(<issue>5</issue>):<fpage>110357</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apacoust.2024.110357</pub-id>.</mixed-citation></ref>
<ref id="ref-111"><label>[111]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational bayes</article-title>. <comment>arXiv:1312.6114. 2013</comment>.</mixed-citation></ref>
<ref id="ref-112"><label>[112]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name>, <string-name><surname>Pouget-Abadie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Warde-Farley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ozair</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generative adversarial networks</article-title>. <source>Commun ACM</source>. <year>2020</year>;<volume>63</volume>(<issue>11</issue>):<fpage>139</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3422622</pub-id>.</mixed-citation></ref>
<ref id="ref-113"><label>[113]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic models</article-title>. <source>Adv Neural Inform Process Syst</source>. <year>2020</year>;<volume>33</volume>:<fpage>6840</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-114"><label>[114]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ellefsen</surname> <given-names>AL</given-names></string-name>, <string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Holmeset</surname> <given-names>FT</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Fault detection with LSTM-based variational autoencoder for maritime components</article-title>. <source>IEEE Sens J</source>. <year>2021</year>;<volume>21</volume>(<issue>19</issue>):<fpage>21903</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2021.3105226</pub-id>.</mixed-citation></ref>
<ref id="ref-115"><label>[115]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Multi-mode non-Gaussian variational autoencoder network with missing sources for anomaly detection of complex electromechanical equipment</article-title>. <source>ISA Transact</source>. <year>2023</year>;<volume>134</volume>(<issue>1</issue>):<fpage>144</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.isatra.2022.09.009</pub-id>; <pub-id pub-id-type="pmid">36150902</pub-id></mixed-citation></ref>
<ref id="ref-116"><label>[116]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A novel deep learning framework for rolling bearing fault diagnosis enhancement using VAE-augmented CNN model</article-title>. <source>Heliyon</source>. <year>2024</year>;<volume>10</volume>(<issue>15</issue>):<fpage>e35407</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.heliyon.2024.e35407</pub-id>; <pub-id pub-id-type="pmid">39166054</pub-id></mixed-citation></ref>
<ref id="ref-117"><label>[117]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Enhancing gas turbine fault diagnosis using a multi-scale dilated graph variational autoencoder model</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>22</issue>):<fpage>104818</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3434708</pub-id>.</mixed-citation></ref>
<ref id="ref-118"><label>[118]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>AHWCN: an interpretable attention-guided hierarchical wavelet convolutional network for rotating machinery intelligent fault diagnosis</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>272</volume>:<fpage>126815</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2025.126815</pub-id>.</mixed-citation></ref>
<ref id="ref-119"><label>[119]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>He</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Deep feature generating network: a new method for intelligent fault detection of mechanical systems under class imbalance</article-title>. <source>IEEE Transact Indus Inform</source>. <year>2020</year>;<volume>17</volume>(<issue>9</issue>):<fpage>6282</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tii.2020.3030967</pub-id>.</mixed-citation></ref>
<ref id="ref-120"><label>[120]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zareapoor</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shamsolmoali</surname> <given-names>P</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Oversampling adversarial network for class-imbalanced fault diagnosis</article-title>. <source>Mech Syst Signal Process</source>. <year>2021</year>;<volume>149</volume>(<issue>1</issue>):<fpage>107175</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2020.107175</pub-id>.</mixed-citation></ref>
<ref id="ref-121"><label>[121]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A multi-module generative adversarial network augmented with adaptive decoupling strategy for intelligent fault diagnosis of machines with small sample</article-title>. <source>Knowl Based Syst</source>. <year>2022</year>;<volume>239</volume>:<fpage>107980</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2021.107980</pub-id>.</mixed-citation></ref>
<ref id="ref-122"><label>[122]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Han</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Intelligent fault diagnosis of bearings under small samples: a mechanism-data fusion approach</article-title>. <source>Eng Appl Artif Intell</source>. <year>2023</year>;<volume>126</volume>:<fpage>107063</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2023.107063</pub-id>.</mixed-citation></ref>
<ref id="ref-123"><label>[123]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Few-Shot GAN: improving the performance of intelligent fault diagnosis in severe data imbalance</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>3516814</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2023.3271746</pub-id>.</mixed-citation></ref>
<ref id="ref-124"><label>[124]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Data augmentation fault diagnosis method based on residual mixed self-attention for rolling bearings under imbalanced samples</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>3528914</fpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2023.3311062</pub-id>.</mixed-citation></ref>
<ref id="ref-125"><label>[125]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Novel imbalanced fault diagnosis method based on generative adversarial networks with balancing serial CNN and Transformer (BCTGAN)</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>258</volume>(<issue>1</issue>):<fpage>125171</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.125171</pub-id>.</mixed-citation></ref>
<ref id="ref-126"><label>[126]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>An intelligent diagnosis scheme based on generative adversarial learning deep neural networks and its application to planetary gearbox fault pattern recognition</article-title>. <source>Neurocomputing</source>. <year>2018</year>;<volume>310</volume>(<issue>1</issue>):<fpage>213</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2018.05.024</pub-id>.</mixed-citation></ref>
<ref id="ref-127"><label>[127]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A generative adversarial network-based intelligent fault diagnosis method for rotating machinery under small sample size conditions</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>149736</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2947194</pub-id>.</mixed-citation></ref>
<ref id="ref-128"><label>[128]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Cross-domain augmentation diagnosis: an adversarial domain-augmented generalization method for fault diagnosis under unseen working conditions</article-title>. <source>Reliab Eng Syst Saf</source>. <year>2023</year>;<volume>234</volume>:<fpage>109171</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ress.2023.109171</pub-id>.</mixed-citation></ref>
<ref id="ref-129"><label>[129]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sel&#x00E7;uk</surname> <given-names>&#x015E;.Y</given-names></string-name>, <string-name><surname>&#x00DC;nal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Albayrak</surname> <given-names>&#x00D6;</given-names></string-name>, <string-name><surname>Jom&#x00E2;a</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A workflow for synthetic data generation and predictive maintenance for vibration data</article-title>. <source>Information</source>. <year>2021</year>;<volume>12</volume>(<issue>10</issue>):<fpage>386</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info12100386</pub-id>.</mixed-citation></ref>
<ref id="ref-130"><label>[130]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Azari</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Santini</surname> <given-names>S</given-names></string-name>, <string-name><surname>Edrisi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Flammini</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Self-adaptive fault diagnosis for unseen working conditions based on digital twins and domain generalization</article-title>. <source>Reliab Eng Syst Saf</source>. <year>2025</year>;<volume>254</volume>(<issue>3</issue>):<fpage>110560</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ress.2024.110560</pub-id>.</mixed-citation></ref>
<ref id="ref-131"><label>[131]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Song</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Novel SVD integrated with GBDT based virtual sample generation and its application in soft sensor</article-title>. <source>IFAC-PapersOnLine</source>. <year>2022</year>;<volume>55</volume>(<issue>7</issue>):<fpage>952</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ifacol.2022.07.567</pub-id>.</mixed-citation></ref>
<ref id="ref-132"><label>[132]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cui</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Virtual sample generation method based on generative adversarial fuzzy neural network</article-title>. <source>Neural Comput Appl</source>. <year>2023</year>;<volume>35</volume>(<issue>9</issue>):<fpage>6979</fpage>&#x2013;<lpage>7001</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-022-08104-5</pub-id>.</mixed-citation></ref>
<ref id="ref-133"><label>[133]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Product quality prediction method in small sample data environment</article-title>. <source>Adv Eng Inform</source>. <year>2023</year>;<volume>56</volume>(<issue>9</issue>):<fpage>101975</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2023.101975</pub-id>.</mixed-citation></ref>
<ref id="ref-134"><label>[134]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>W</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Regression loss-assisted conditional style generative adversarial network for virtual sample generation with small data in soft sensing</article-title>. <source>Eng Appl Artif Intell</source>. <year>2025</year>;<volume>147</volume>(<issue>6</issue>):<fpage>110306</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2025.110306</pub-id>.</mixed-citation></ref>
<ref id="ref-135"><label>[135]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>DC</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Tsai</surname> <given-names>TI</given-names></string-name>, <string-name><surname>Lina</surname> <given-names>YS</given-names></string-name></person-group>. <article-title>Using mega-trend-diffusion and artificial samples in small data set learning for early flexible manufacturing system scheduling knowledge</article-title>. <source>Comput Operat Res</source>. <year>2007</year>;<volume>34</volume>(<issue>4</issue>):<fpage>966</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cor.2005.05.019</pub-id>.</mixed-citation></ref>
<ref id="ref-136"><label>[136]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Principle of information diffusion</article-title>. <source>Fuzzy Sets Syst</source>. <year>1997</year>;<volume>91</volume>(<issue>1</issue>):<fpage>69</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1016/s0165-0114(96)00257-6</pub-id>.</mixed-citation></ref>
<ref id="ref-137"><label>[137]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chawla</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Bowyer</surname> <given-names>KW</given-names></string-name>, <string-name><surname>Hall</surname> <given-names>LO</given-names></string-name>, <string-name><surname>Kegelmeyer</surname> <given-names>WP</given-names></string-name></person-group>. <article-title>SMOTE: synthetic minority over-sampling technique</article-title>. <source>J Artif Intell Res</source>. <year>2002</year>;<volume>16</volume>:<fpage>321</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1613/jair.953</pub-id>.</mixed-citation></ref>
<ref id="ref-138"><label>[138]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>WA</given-names></string-name></person-group>. <article-title>Balanced weighted extreme learning machine for imbalance learning of credit default risk and manufacturing productivity</article-title>. <source>Ann Operat Res</source>. <year>2025</year>;<volume>348</volume>(<issue>2</issue>):<fpage>833</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10479-023-05194-9</pub-id>.</mixed-citation></ref>
<ref id="ref-139"><label>[139]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A two-phase soft sensor modeling framework for quality prediction in industrial processes with missing data</article-title>. <source>J Process Control</source>. <year>2023</year>;<volume>129</volume>(<issue>4</issue>):<fpage>103061</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jprocont.2023.103061</pub-id>.</mixed-citation></ref>
<ref id="ref-140"><label>[140]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ricker</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Optimal steady-state operation of the Tennessee Eastman challenge process</article-title>. <source>Comput Chem Eng</source>. <year>1995</year>;<volume>19</volume>(<issue>9</issue>):<fpage>949</fpage>&#x2013;<lpage>59</lpage>.</mixed-citation></ref>
<ref id="ref-141"><label>[141]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <source>Tennessee Eastman simulation dataset</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>IEEE Dataport</publisher-name>; <year>2019</year>. doi:<pub-id pub-id-type="doi">10.21227/4519-z502</pub-id>.</mixed-citation></ref>
<ref id="ref-142"><label>[142]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zio</surname> <given-names>E</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Artificial intelligence for fault diagnosis of rotating machinery: a review</article-title>. <source>Mechan Syst Signal Process</source>. <year>2018</year>;<volume>108</volume>(<issue>7</issue>):<fpage>33</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2018.02.016</pub-id>.</mixed-citation></ref>
<ref id="ref-143"><label>[143]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Raj</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>RR</given-names></string-name>, <string-name><surname>Andriollo</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Enhanced fault detection in bearings using machine learning and raw accelerometer data: a case study using the Case Western Reserve University dataset</article-title>. <source>Information</source>. <year>2024</year>;<volume>15</volume>(<issue>5</issue>):<fpage>259</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info15050259</pub-id>.</mixed-citation></ref>
<ref id="ref-144"><label>[144]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Smith</surname> <given-names>WA</given-names></string-name>, <string-name><surname>Randall</surname> <given-names>RB</given-names></string-name></person-group>. <article-title>Rolling element bearing diagnostics using the Case Western Reserve University data: a benchmark study</article-title>. <source>Mecha Syst Signal Process</source>. <year>2015</year>;<volume>64</volume>:<fpage>100</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2015.04.021</pub-id>.</mixed-citation></ref>
<ref id="ref-145"><label>[145]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Preprocessing-free gear fault diagnosis using small datasets with deep convolutional neural network-based transfer learning</article-title>. <source>IEEE Access</source>. <year>2018</year>;<volume>6</volume>:<fpage>26241</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2018.2837621</pub-id>.</mixed-citation></ref>
<ref id="ref-146"><label>[146]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xia</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aljerf</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Unveiling dioxin dynamics: a whole-process simulation study of municipal solid waste incineration</article-title>. <source>Sci Total Environ</source>. <year>2024</year>;<volume>954</volume>(<issue>10</issue>):<fpage>176241</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.scitotenv.2024.176241</pub-id>; <pub-id pub-id-type="pmid">39299308</pub-id></mixed-citation></ref>
<ref id="ref-147"><label>[147]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>H</given-names></string-name>, <string-name><surname>Aljerf</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ukaogo</surname> <given-names>PO</given-names></string-name></person-group>. <article-title>Prediction of dioxin emission from municipal solid waste incineration based on expansion, interpolation, and selection for small samples</article-title>. <source>J Environ Chem Eng</source>. <year>2022</year>;<volume>10</volume>(<issue>5</issue>):<fpage>108314</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jece.2022.108314</pub-id>.</mixed-citation></ref>
<ref id="ref-148"><label>[148]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Dioxin emission risk warning model in MSWI process based on GAN with active learning mechanism</article-title>. <source>J Beijing Univ Technol</source>. <year>2023</year>;<volume>49</volume>(<issue>5</issue>):<fpage>507</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>