<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81931</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081931</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Research on Gearbox Fault Diagnosis Method Based on Multi-Dimensional Feature Extraction and Random Forest</article-title>
<alt-title alt-title-type="left-running-head">Research on Gearbox Fault Diagnosis Method Based on Multi-Dimensional Feature Extraction and Random Forest</alt-title>
<alt-title alt-title-type="right-running-head">Research on Gearbox Fault Diagnosis Method Based on Multi-Dimensional Feature Extraction and Random Forest</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Yu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="author-notes" rid="afn1">#</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Tan</surname><given-names>Shihan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="author-notes" rid="afn1">#</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Lian</surname><given-names>Guangyao</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Dun</surname><given-names>Congying</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Hu</surname><given-names>Qiwei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>hu_q_w@163.com</email></contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Guo</surname><given-names>Chiming</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>guochiming@nudt.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>Shijiazhuang Campus, Army Engineering University of PLA</institution>, <addr-line>Shijiazhuang</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>No. 32181 Unit of PLA</institution>, <addr-line>Xian</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Army Command Academy</institution>, <addr-line>Nanjing</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Qiwei Hu. Email: <email>hu_q_w@163.com</email>; Chiming Guo. Email: <email>guochiming@nudt.edu.cn</email></corresp>
<fn id="afn1">
<p><sup>#</sup>These authors contributed equally to this work as the first author</p>
</fn>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>71</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81931.pdf"></self-uri>
<abstract>
<p>Gearboxes are critical components in the transmission systems of various mechanical equipment. Subjected to complex and harsh operating conditions for a long time, they suffer from a high failure rate and potentially severe consequences. Traditional fault diagnosis methods are limited by problems such as noise interference, and can hardly meet the requirements in terms of diagnostic accuracy, generalization ability, and reliability. To tackle the deficiencies of traditional gearbox fault diagnosis methods, including insufficient utilization of features, poor generalization under small-sample conditions, and weak model interpretability, this paper proposes a fault diagnosis method based on multi-dimensional feature extraction and Random Forest (RF). This method integrates intelligent computing, data-driven approaches, and mechanical structural health monitoring. First, fault feature analysis is conducted from multiple dimensions including time domain, frequency domain, and envelope domain, and visualization verification is implemented using Principal Component Analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE). Then, the Random Forest (RF) algorithm is used for dataset training and testing, obtaining a stable diagnostic model with strong generalization ability. Finally, experimental analyses verify the effectiveness and superiority of the proposed method. The research results possess high application potential and practical value in improving the performance of gearbox fault diagnosis.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Fault diagnosis</kwd>
<kwd>feature extraction</kwd>
<kwd>principal component analysis</kwd>
<kwd>gearbox</kwd>
<kwd>random forest</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>As a core power transmission component in the drive systems of mechanical equipment, the gearbox operates for long periods under complex and harsh working conditions such as high rotational speeds, alternating loads, and intense impact vibrations [<xref ref-type="bibr" rid="ref-1">1</xref>], making it prone to failures like tooth surface wear, tooth breakage, and gearbox damage. With a high failure rate and far-reaching impacts, it is the primary source of faults in drive system equipment [<xref ref-type="bibr" rid="ref-2">2</xref>]. Once a gearbox malfunctions, it may cause minor issues such as abnormal power transmission of the equipment and delays, or even lead to serious accidents like drive system failure in severe cases. At present, gearbox condition monitoring mainly relies on characteristic indicators constructed based on expert experience or traditional signal processing technologies [<xref ref-type="bibr" rid="ref-3">3</xref>]. However, with the intelligent upgrading of equipment, existing methods have been difficult to meet the requirements of accurate diagnosis in complex scenarios in terms of diagnostic accuracy, generalization ability, and real-time performance.</p>
<p>In recent years, the vigorous development of big data and artificial intelligence technologies has brought new opportunities to the field of gearbox fault diagnosis. Data-driven intelligent fault diagnosis methods build fault diagnosis models based on massive operational data [<xref ref-type="bibr" rid="ref-4">4</xref>], which can achieve higher-precision fault identification, stronger working condition adaptability and more efficient real-time diagnosis capability. There are relatively many studies on gearbox fault signal processing and feature extraction. For example, Li et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed a coarse-grained lattice feature method. After adaptive filtering, frequency-domain segmentation and other processes, combined with Swin Transformer modeling, the diagnostic accuracy of public datasets and experimental data exceeds 98%. Xiao et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] proposed a combined method of NMD and wavelet threshold denoising. After preprocessing the signal, decomposition and envelope spectrum are performed. Experimental comparison shows that the accuracy of extracting fault characteristic frequencies is better than that of the EMD method. Bie et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] proposed an improved method combining ESMD and SVM, which decomposes vibration signals to select high kurtosis envelope spectrum exponential components and constructs feature vectors using various entropies, thus effectively extracting the impact fault features of gearboxes. Cao et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] proposed a method combining quadratic wavelet packet energy entropy and t-SNE, which decomposes signals to extract entropy values and fuse features, and realizes fault state identification effectively with the help of support vector machine recognition. Zhang and Wang [<xref ref-type="bibr" rid="ref-8">8</xref>] proposed a method combining wavelet packet decomposition and tree-structured pipeline optimization tool, which extracts feature vectors and then uses genetic programming to generate the optimal machine learning pipeline. Experiments confirm that this method has significant advantages. Zhao et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] combined the multi-point optimized minimum entropy deconvolution correction method with improved adaptive noise complete ensemble empirical mode decomposition, and through envelope demodulation, can accurately extract the characteristic frequencies of weak faults. Zhu et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] proposed a transmission path elimination enhanced variational mode decomposition method, combined with spectral editing and whale optimization algorithm for noise reduction, which can accurately extract fault features under complex transmission paths. Zheng et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] proposed a spectral whitening demodulation method for bogie axle box gearboxes, which is adapted to the complex working conditions of high-speed trains, can effectively extract fault features and ensure diagnostic accuracy. Li et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] pointed out that existing signal decomposition algorithms suffer from mode mixing, unstable decomposition, and poor anti-noise performance, which hinder subsequent feature extraction and fault diagnosis. To address these issues, they proposed a Spectral Distribution Decomposition (SDD) algorithm based on the Spectral Probability Density Function (SPDF) for gearbox fault diagnosis. The effectiveness and superiority of the proposed method are verified through comparative simulations with existing decomposition techniques, as well as experimental tests on fault diagnosis for laboratory gearboxes and actual wind turbine gearboxes. Chen et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] aimed to solve the problem of scarce fault data for gearboxes and bearings, as well as the difficulty of effective modeling based on single health state data. They proposed a Convex Optimization Differential Analysis model (CODA model). Within a dual-spectrum framework based on natural frequency and harmonic frequency demodulation, a differential analysis mechanism is introduced to discretize complex modulation features and realize the decoupling of modulation information among multiple components.</p>
<p>There have also been many relevant studies on the construction of gearbox fault diagnosis models. Li et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed a domain-adaptive LSTM-DNN model embedded with the MMD loss function to reduce the distribution discrepancy under variable working conditions. Compared with the single LSTM model, the mean absolute error (MAE) of the proposed model is reduced by 35%, which is suitable for the cross-working condition life prediction of planetary gearboxes. Hogea et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] constructed a LogicLSTM model integrating XAI and logical tensor networks, and realized logical-neural training by virtue of the ApME loss function, achieving accurate and interpretable diagnosis of 9 types of states on the data from the DDS platform. Dou et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] converted one-dimensional signals into two-dimensional graphs via Gramian Angular Field (GAF), built a lightweight network with coordinate attention, and combined it with transfer learning fine-tuning to realize cross-component fault diagnosis of gearboxes under variable working conditions with few samples. Dong et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed a framework based on the nonlinear Wiener process, processed vibration signals using kernel principal component analysis (KPCA), and updated parameters through Bayesian inference. A 954-h experiment verified that the framework can dynamically optimize the maintenance time of gearboxes. Yuan et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] constructed a model integrating improved multi-scale convolutional neural network (CNN) and lightweight convolutional attention, and introduced the parametric rectified linear unit. The diagnostic accuracy of the model exceeds 98.9% under complex working conditions, balancing efficiency and precision. Cheng et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] designed a lightweight channel attention mechanism, obtained the time-frequency distribution of signals through wavelet transform, and combined it with transfer learning to fine-tune the network, efficiently solving the problem of gearbox fault classification under multiple working conditions with few samples. Nguyen et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] combined adaptive noise control with stacked sparse autoencoder to eliminate noise in vibration signals, completed feature extraction and classification in an integrated manner, and improved the diagnostic sensitivity for multi-stage tooth breakage faults under variable rotational speeds. Desai et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] fused SCADA time-series data and physical modeling data as model inputs to predict faults in wind turbine gearboxes, reducing the false alarm rate by 50%, improving the accuracy by 33%, and realizing fault early warning one month in advance. Zhu et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a novel federated learning method based on an incremental generalized federated learning system to address the challenges of long global model training time and model forgetting in deep network federated learning approaches. This method employs random mapping as a bridge between broad learning and federated learning, constructing a federated learning framework built on the broad learning system. Kan et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] developed a fault diagnosis method based on dual-layer adaptive personalized federated learning to tackle the issues of privacy leakage and statistical heterogeneity that often restrict the performance of fault diagnosis in industrial processes. With the aid of federated learning, multiple clients train their models locally, which effectively resolves the problem of privacy leakage.</p>
<p>However, in practical application scenarios, the raw vibration signals collected by sensors are affected by complex operating environments and variable working conditions, and are easily mixed with components such as background noise and interference responses. This greatly weakens the saliency of fault features. Although existing studies have made progress in feature extraction and model construction, limitations still exist: first, most methods rely only on single-dimensional features in the time or frequency domain, and cannot fully characterize fault information from non-stationary and strong-noise signals; second, traditional models have insufficient generalization ability for small-sample and high-dimensional features; third, there is a lack of systematic demonstration of feature dimensions, model adaptability, and hyperparameter selection. To fill the above gaps, this paper proposes a diagnostic framework based on multi-domain feature fusion combined with random forest, to achieve feature complementarity and improved model robustness.</p>
<p>In summary, existing fault diagnosis methods still have obvious shortcomings in the collaborative utilization of multi-domain features, adaptability to small-sample working conditions, model interpretability, and engineering practicability. Therefore, this paper proposes a fault diagnosis method based on multi-dimensional feature extraction and random forest. First, fault feature analysis is performed from multiple dimensions including the time domain, frequency domain, and envelope domain, and visualization verification is carried out using PCA and t-SNE. Then, the RF method is used for dataset training and testing to obtain a stable diagnostic model with strong generalization ability. Finally, experimental comparative analysis verifies the effectiveness and superiority of the proposed method. By systematically fusing time-domain, frequency-domain, and envelope-domain features, and combining an optimized random forest to construct an end-to-end diagnostic framework suitable for small-sample working conditions, multi-dimensional information complementarity and high-precision fault identification are realized.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Fault Feature Extraction Based on Multi-Dimensions</title>
<p>A multi-dimensional feature extraction model is constructed and the optimal feature set parameters are obtained. The diagnostic accuracy is improved through consistency verification of features across different dimensions. This method is suitable for complex non-stationary fault signals of gearboxes, can significantly enhance feature distinguishability and diagnostic robustness, and provides a more comprehensive and reliable basis for subsequent fault identification and classification.</p>
<p>This paper constructs a comprehensive, complementary and highly discriminative feature set from three physical domains: time domain, frequency domain and envelope domain, rather than relying solely on a single domain or a small number of statistics. The extracted features cover four categories: amplitude statistics, distribution morphology, energy distribution and modulation characteristics, which together form a genuine multi-domain feature space. This space can reflect the laws of fault evolution from different perspectives, ensuring the completeness and discriminability of the feature space.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Vibration Signal Preprocessing</title>
<p>Raw vibration signals are susceptible to environmental noise, electromagnetic interference, and baseline drift during acquisition. Direct feature extraction will degrade diagnostic accuracy. Therefore, standardized preprocessing is performed on the signals prior to multi-domain feature extraction, with the procedure as follows:</p>
<p>(1) Wavelet threshold denoising</p>
<p>The db4 wavelet is used for 3-level decomposition, and soft threshold filtering is adopted to remove high-frequency random noise while retaining fault impact features.</p>
<p>(2) Zero-mean normalization</p>
<p>Zero-mean processing is applied to the denoised signal to eliminate the effects of sensor zero offset and amplitude dimensions, with the formula as follows:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi></mml:math></disp-formula>where <italic>&#x03BC;</italic> denotes the mean value of the signal.</p>
<p>(3) Fixed-length segmentation and resampling</p>
<p>Each 6-s signal is segmented into a fixed length of 2048 points to ensure consistent input feature dimensions and avoid errors caused by inconsistent sample lengths.</p>
<p>After the above steps, feature extraction in the time domain, frequency domain, and envelope domain is performed.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Time-Domain Feature Extraction</title>
<p>Time-domain features directly reflect the statistical characteristics of vibration signals. By extracting the statistical and morphological features of signals, they reflect the amplitude distribution, fluctuation law and mutation characteristics of signals, and are sensitive to the transient impact of early faults in gearboxes. The time-domain feature vector is defined to include the following indicators:</p>
<p>(1) Root Mean Square (RMS)
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:msqrt></mml:math></disp-formula>where <italic>N</italic> is the signal length and <italic>x</italic><sub><italic>i</italic></sub> is the <italic>i</italic>-th sampling point. The <italic>RMS</italic> is used to measure the overall energy level of the signal. Gearbox faults will lead to an increase in vibration energy, and the <italic>RMS</italic> value will rise accordingly.</p>
<p>(2) Kurtosis (K)
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:mfrac></mml:mstyle></mml:math></disp-formula>where <italic>&#x03BC;</italic> is the mean value of the signal and <italic>&#x03C3;</italic> is the standard deviation. Kurtosis reflects the sharpness of the signal distribution. When the impact caused by faults thickens the tail of the distribution, the kurtosis value increases.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Frequency-Domain Feature Extraction</title>
<p>Frequency-domain features reveal the periodic components of signals through Fourier transform. As a core mathematical tool in the field of signal processing, Fourier transform can convert signals in the time domain to the frequency domain, revealing the composition of signal frequency components, which breaks the limitation that time-domain analysis is difficult to intuitively present the frequency characteristics of signals.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>j</mml:mi><mml:mi>&#x03C9;</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:math></disp-formula>where <italic>&#x03C9;</italic> is the angular frequency and <italic>j</italic> is the imaginary unit. Through integral operation, the time-domain signal is decomposed into a superposition form of sine/cosine components with different frequencies.</p>
<p>The spectral centroid is an indicator reflecting the concentration trend of signal energy distribution in the frequency domain. The more the energy shifts to high frequencies, the larger the value of the spectral centroid.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>C</mml:mi></mml:math></inline-formula> is the spectral centroid, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>f</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the frequency at the <italic>n</italic>-th point, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is the power spectrum corresponding to the frequency, and the denominator is the total power.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Envelope-Domain Feature Extraction</title>
<p>Envelope analysis is an important method for extracting amplitude variation features in the field of signal processing [<xref ref-type="bibr" rid="ref-24">24</xref>]. Its core lies in separating the slowly varying amplitude envelope that reflects key information from complex signals modulated by high-frequency carriers. This method can effectively eliminate redundant high-frequency components [<xref ref-type="bibr" rid="ref-25">25</xref>] and highlight the core variation trend of signals.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>z</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msqrt></mml:math></disp-formula>where the original signal is denoted as <italic>x</italic>(<italic>t</italic>), the signal obtained via Hilbert transform is denoted as <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and the analytic signal is denoted as <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>z</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>j</mml:mi><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>The envelope root mean square is an indicator reflecting the concentration degree of signal envelope energy, which can effectively highlight the amplitude fluctuations caused by faults and is often used in equipment fault diagnosis.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:msqrt></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the envelope root mean square value, <italic>N</italic> is the number of sampling points of the envelope signal, and B is the sampling value of the <italic>i</italic>-th envelope signal.</p>
<p>In this paper, 12 effective features are extracted from three dimensions: time domain, frequency domain, and envelope domain, and a comprehensive feature matrix is constructed for model training and testing. The dimension of the feature matrix finally input into the random forest model is 32 &#x00D7; 12, where 32 is the total number of samples and 12 is the dimension of the feature vector. The detailed composition is shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Composition of feature vector dimensions.</title>
</caption>
<table>
<colgroup>
<col align="center" width="20mm"/>
<col align="center" width="55mm"/>
<col align="center" width="50mm"/> </colgroup>
<thead>
<tr>
<th>Domain</th>
<th>Feature Name</th>
<th>Physical Meaning</th>
</tr>
</thead>
<tbody>
<tr>
<td>Time Domain</td>
<td>Root Mean Square, Kurtosis, Crest Factor, Root Amplitude, Waveform Index, Impulse Index</td>
<td>Reflects vibration amplitude, impact strength and distribution characteristics</td>
</tr>
<tr>
<td>Frequency Domain</td>
<td>Spectral Centroid, Mean Square Frequency, Frequency Variance</td>
<td>Reflects the distribution law of frequency energy</td>
</tr>
<tr>
<td>Envelope Domain</td>
<td>Envelope RMS, Envelope Kurtosis, Envelope Crest Factor</td>
<td>Reflects fault modulation and impact characteristics</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Fault Diagnosis Model Based on Random Forest</title>
<p>According to the feature extraction results of the fault data mentioned above, the Random Forest (RF) method is adopted to construct an intelligent fault diagnosis model for gearboxes.</p>
<sec id="s3_1">
<label>3.1</label>
<title>RF Model Construction</title>
<p>In this paper, multi-dimensional features are constructed from the time domain, frequency domain, and envelope domain, with weak correlation among these features. Random Forest (RF) is insensitive to high-dimensional features and does not require complex feature selection, enabling direct fusion of multi-domain information. Meanwhile, given the limited sample size in the gearbox experiment, RF reduces the risk of overfitting through ensemble and randomization strategies, and exhibits strong robustness to environmental noise in vibration signals, which conforms to actual industrial field conditions. In contrast, SVM relies on appropriate kernel function selection, and MLP requires a large amount of data and careful parameter tuning. RF imposes no strict assumptions on data distribution, has fewer parameters and is easier to optimize, making it more suitable for engineering applications.</p>
<p>Random Forest (RF) is a powerful ensemble learning algorithm composed of <italic>M</italic> decision trees [<xref ref-type="bibr" rid="ref-26">26</xref>]. Each decision tree is constructed based on a different training subset, which is obtained from the original training set through random sampling with replacement [<xref ref-type="bibr" rid="ref-27">27</xref>]. This sampling method enables each decision tree to be trained on relatively independent and differentiated data, thereby increasing the diversity of the model. The workflow diagram of the RF model is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Workflow of the random forest model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-1.tif"/>
</fig>
<p>For classification problems, each decision tree <italic>T</italic><sub><italic>m</italic></sub> independently performs classification prediction based on the input gearbox feature vector <italic>x</italic> [<xref ref-type="bibr" rid="ref-28">28</xref>]. The final prediction result is determined by a voting mechanism, and the specific formula is as follows:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:munder><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is the final predicted class of the RF model; <italic>c</italic> denotes the fault class; <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the indicator function.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Parameter Optimization</title>
<p>The selection of parameters has a crucial impact on the performance of the RF model. To improve the accuracy and stability of the model, this paper adopts the method of grid search combined with cross-validation to optimize its key parameters.</p>
<p>Number of decision trees: The number of decision trees determines the scale of the ensemble model. A larger number of decision trees can enhance the stability and generalization ability of the model, but it will also increase the computational cost [<xref ref-type="bibr" rid="ref-29">29</xref>]. The value range considered in this paper is {50, 100, 200, 300}.</p>
<p>Maximum depth: The maximum depth limits the growth of decision trees. A smaller maximum depth can prevent overfitting, but may lead to underfitting; a larger maximum depth can enable the model to better fit the training data, but may increase the risk of overfitting. The value range is set as {5, 10, 15, 20, None}, where None means no depth limit, the decision tree will grow until all leaf nodes are pure or other stopping conditions are met.</p>
<p>Number of split features: When splitting at each node, a certain number of features are randomly selected for evaluation. This increases the diversity among decision trees and helps improve the overall performance of the model. The value range is <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mo>{</mml:mo><mml:msqrt><mml:mi>n</mml:mi></mml:msqrt><mml:mo>,</mml:mo><mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>n</italic> is the total number of features.</p>
<p>Minimum number of samples at leaf nodes: This parameter limits the minimum number of samples in leaf nodes and prevents decision trees from being overly complex. The value range is {1, 2, 4}.</p>
<p>The objective function adopts the average accuracy of 5-fold cross-validation, and the calculation formula is as follows:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>S</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>5</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:munderover><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>By traversing all possible hyperparameter combinations and calculating the objective function value corresponding to each combination, the optimal hyperparameter combination of the random forest model is finally determined in this paper as follows:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>100</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mi>n</mml:mi></mml:msqrt></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This set of parameters ensures the classification accuracy of the model while effectively avoiding overfitting, enabling the model to achieve better stability and generalization ability on small-sample datasets.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Model Evaluation Metrics</title>
<p>To comprehensively evaluate the performance of the model on the four-classification problem, the following evaluation metrics are adopted in this paper:</p>
<p>Accuracy: As the most commonly used evaluation metric, accuracy refers to the proportion of correctly predicted samples to the total number of samples. The calculation formula is as follows:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Macro-Precision: Macro-average precision takes into account the precision of each category and avoids the impact of class imbalance on the evaluation results. The calculation formula is as follows:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Macro-Recall: Macro-average recall measures the model&#x2019;s ability to identify faults of each category. The calculation formula is as follows:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>F1-Score: F1-score is the harmonic mean of precision and recall, which comprehensively considers the precision and recall of the model. The calculation formula of macro-average F1-score is as follows:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>macro</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>where <italic>TP</italic>, <italic>TN</italic>, <italic>FP</italic>, and <italic>FN</italic> denote true positive, true negative, false positive, and false negative, respectively; <italic>k</italic> is the number of categories; <italic>P</italic><sub><italic>i</italic></sub> and <italic>R</italic><sub><italic>i</italic></sub> represent the precision and recall of the <italic>i</italic>-th fault category, respectively.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Verification</title>
<p>To verify the effectiveness and superiority of the proposed gearbox fault diagnosis method, vibration signals of the gearbox were collected using a test bench. Time-domain, frequency-domain, and envelope-domain features were extracted and input into the constructed RF model for experimental analysis and comparative verification, thereby providing experimental support for the diagnostic application of gearbox faults.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Platform</title>
<p>The experimental platform consists of two main parts: a gear reducer and a signal acquisition system, whose structure is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The core components include a base, a gear reducer, a three-phase asynchronous motor, an electromagnetic speed-regulating motor, a magnetic powder brake, a speed regulator, and a signal acquisition device. In this experiment, the JZQ250 type two-stage parallel shaft gearbox is taken as the research object. This gear reducer is widely used in various mechanical equipment and has important research significance.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Gearbox experimental platform.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-2.tif"/>
</fig>
<p>For the vibration signal acquisition process, the Donghua DH5981 dynamic signal acquisition instrument was selected, paired with the IEPE low-impedance voltage-output general-purpose vibration sensor (Model 14100), with the sampling frequency set to 20 kHz.</p>
<p>An industrial-standard vibration sensor arrangement is adopted. The sensors are installed vertically on the bearing housing of the gearbox, close to the shortest transmission path of vibration energy, which can effectively capture gear meshing vibration and fault impact signals, ensuring that the acquired vibration data have high signal-to-noise ratio and high representativeness. This installation position follows the general specification layout for rotating machinery fault diagnosis, which can stably obtain fault-sensitive features and meet the requirements of fault feature extraction and diagnostic identification.</p>
<p>The specific connection method is as follows: after installing the vibration sensor at the designated position of the reducer, it is connected to the Donghua DH5981 dynamic signal acquisition instrument via a dedicated connecting cable. The acquisition instrument then transmits the vibration signals to the supporting signal acquisition software of Donghua. This software can implement core functions such as signal storage, format conversion and preliminary analysis, providing fundamental support for subsequent data processing.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Vibration Signal Acquisition Scheme</title>
<p>Common fault modes of gearboxes include gear wear, gear cracks, gearbox tooth breakage, and tooth surface scuffing. This experiment mainly takes gear cracks as an example for research and analysis. First, to explore how to identify gear crack faults of varying degrees, four working conditions were set for the intermediate gear under the same operating conditions, namely normal state, 2 mm crack, 5 mm crack, and 8 mm crack.</p>
<p>Vibration signal acquisition should be initiated after the experimental platform operates stably, so as to avoid the interference of factors such as insufficient equipment preheating and voltage fluctuations in the initial stage of startup on signal quality. The experimental parameters are set as follows: the motor input speed is 1200 r/min, the sampling frequency of the vibration sensor is 20 kHz, the data acquisition duration of each group is 6 s, and the next group of acquisition is carried out after an interval of 5 s between groups, with a total of 8 groups of data obtained. The specific signal acquisition scheme is shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Fault signal acquisition scheme.</title>
</caption>
<table>
<colgroup>
<col align="center" width="25mm"/>
<col align="center" width="25mm"/>
<col align="center" width="20mm"/>
<col align="center" width="14mm"/>
<col align="center" width="14mm"/>
<col align="center" width="20mm"/>
<col align="center" width="25mm"/> </colgroup>
<thead>
<tr>
<th>Working Condition Label</th>
<th>Working Condition Type</th>
<th>Sampling Frequency</th>
<th>Sampling Time</th>
<th>Sampling Interval</th>
<th>Number of Samples</th>
<th>Input Rotational Speed</th>
</tr>
</thead>
<tbody>
<tr>
<td>S1</td>
<td>Normal State</td>
<td>20,000 Hz</td>
<td>6 s</td>
<td>5 s</td>
<td>8</td>
<td>1200 r/min</td>
</tr>
<tr>
<td>S2</td>
<td>2 mm Crack</td>
<td>20,000 Hz</td>
<td>6 s</td>
<td>5 s</td>
<td>8</td>
<td>1200 r/min</td>
</tr>
<tr>
<td>S3</td>
<td>5 mm Crack</td>
<td>20,000 Hz</td>
<td>6 s</td>
<td>5 s</td>
<td>8</td>
<td>1200 r/min</td>
</tr>
<tr>
<td>S4</td>
<td>8 mm Crack</td>
<td>20,000 Hz</td>
<td>6 s</td>
<td>5 s</td>
<td>8</td>
<td>1200 r/min</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To fully verify the diagnostic performance and stability of the proposed method, experiments are carried out in this paper under the premise of consistent controlled variables and strictly identical working conditions. Eight groups of high-quality, high-signal-to-noise ratio valid samples are collected under each working condition to ensure good representativeness and consistency of the samples. A stratified 5-fold cross-validation strategy is adopted in the experiment to comprehensively evaluate the model. Through multiple random divisions of the training set and test set, the value of data is fully exploited, and the reliability and statistical significance of experimental results are effectively improved.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Fault Feature Extraction</title>
<p>For the collected dataset samples, the stratified 5-fold cross-validation strategy is adopted for data partitioning. According to four working condition categories, all samples are randomly divided into five mutually exclusive subsets via proportional stratified sampling. In each iteration, four subsets are used as the training set and one subset as the test set. The process is repeated five times and the average results are obtained.</p>
<p>Afterwards, comparative analysis of signal waveforms is conducted, followed by multi-dimensional feature extraction using the method proposed above. Finally, distribution boxplots are applied, and dimensionality reduction and visualization are performed on the multi-dimensional feature space based on PCA and t-SNE.</p>
<p>As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>: The vibration amplitude of the healthy state S1 remains stable within &#x00B1;0.2 g, with a steady signal and no obvious impacts; the amplitude of the 2 mm crack state S2 expands to &#x00B1;0.3 g, and periodic impacts begin to appear; the amplitude of the 5 mm crack state S3 further increases to &#x00B1;0.4 g, with shortened impact intervals and enhanced energy; the maximum amplitude of the 8 mm crack state S4 exceeds &#x00B1;0.5 g, with both impact intensity and frequency reaching the highest levels. As the crack size increases, the vibration amplitude, impact characteristics, and energy level show a stepwise increasing trend, which is consistent with the gear crack propagation mechanism.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Comparison chart of time-domain waveforms.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-3.tif"/>
</fig>
<p>As illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the energy distribution of the normal state S1 is uniform across the full frequency band without distinct peaks. Obvious prominent peaks emerge at the gear meshing characteristic frequencies for S2, S3 and S4, and the peak amplitudes increase by 2 to 5 times with the expansion of cracks. Meanwhile, the vibration energy gradually concentrates in the high-frequency band, which is positively correlated with the fault severity.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Frequency domain analysis diagram.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-4.tif"/>
</fig>
<p>It can be observed from <xref ref-type="fig" rid="fig-5">Fig. 5</xref> that envelope analysis can effectively separate high-frequency carriers and highlight the slowly varying characteristics of fault impacts. The envelope curve of S1 is stable with slight amplitude fluctuations. For S2&#x2013;S4, the envelope amplitude gradually increases and the fluctuations become more severe as the crack propagates, and the root mean square value of the envelope rises by more than three times, which clearly reflects the evolution law of fault severity.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison chart of envelope analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-5.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents the box plot of feature distribution, a statistical chart for displaying data distribution that enables rapid identification of data features and anomalies. Significant differences exist in the distribution of various statistical features of vibration signals, such as root mean square (RMS), kurtosis, spectral centroid, and envelope root mean square, when the gearbox operates normally and under different fault conditions. Specifically, in the normal state, the RMS values show concentrated distribution with relatively low magnitudes; the kurtosis values are low and centralized; the spectral centroid values are high and stable; and the envelope RMS values are low and concentrated. In contrast, under fault conditions, especially with an 8 mm gear crack, the RMS values exhibit more dispersed distribution with higher magnitudes; the kurtosis values increase sharply with scattered distribution; the spectral centroid values are low with fluctuations; and the envelope RMS values are high and dispersed. These differences indicate that various statistical features can serve as effective indicators for distinguishing between the normal and fault states of the gearbox as well as for identifying different fault types.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Box plot of feature distribution.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-6.tif"/>
</fig>
<p>As can be seen from the multi-dimensional feature space analysis in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, in the feature space distribution of Principal Component Analysis (PCA), the sample points of S1 state (green dots), S2 state (blue dots), S3 state (red dots) and S4 state (orange dots) have a certain degree of separability. The normal sample points are relatively concentrated in a specific area, while the fault sample points are distributed in different ranges, indicating that PCA can partition the feature space of gearbox samples under different states to a certain extent. The principal component contribution chart shows that the explained variance ratio of PC1 reaches 0.400, making it the most dominant principal component. With the increase in the number of principal components, the cumulative contribution degree rises gradually, which indicates that the first several principal components can explain most of the data variance. In the t-SNE feature space distribution, the clustering effect of sample points under different states is more obvious. The normal sample points and various types of fault sample points form relatively independent clusters, respectively, suggesting that t-SNE has a stronger ability to distinguish gearbox samples under different states after dimensionality reduction, which is more conducive to subsequent fault identification and classification.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Multi-dimensional space analysis chart.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-7.tif"/>
</fig>
<p>Finally, the multi-dimensional features including time-domain, frequency-domain and envelope analysis are integrated into a comprehensive feature matrix, forming a complete feature set that can reflect the multi-domain characteristics of gearbox faults.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Comparative Analysis of Fault Diagnosis Models</title>
<p>For the extracted fault feature set, the previously constructed fault diagnosis model based on Random Forest (RF) was trained, and a comprehensive comparative analysis was conducted. The models involved in the comparison included Random Forest (RF), Gradient Boosting (GB), Extremely Randomized Trees (ET), Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), Logistic Regression (LR), K-Nearest Neighbors (KNN), and Gaussian Naive Bayes (GNB). Through the training and evaluation of these models on the same dataset, quantitative analysis was carried out from multi-dimensional metrics such as accuracy, precision, recall, F1-score, and cross-validation stability. Ultimately, a comprehensive performance comparison result was formed, which demonstrated the effectiveness and superiority of the selected fault diagnosis model.</p>
<p>From the perspective of model characteristics, methods such as SVM and MLP usually have stronger nonlinear fitting capabilities in large-scale data scenarios, but their performance is sensitive to sample size. The multi-dimensional feature and random forest framework constructed in this paper is originally designed to adapt to industrial fault diagnosis scenarios with small samples and high-dimensional features. Through ensemble learning and random attribute selection mechanisms, random forest can maintain strong feature learning and classification ability under limited sample conditions. All comparative models in this paper adopt a unified grid search strategy to complete parameter optimization, ensuring consistent comparison conditions. The experimental results can effectively reflect the applicability and performance differences of different methods in gearbox fault diagnosis tasks with small samples.</p>
<p>Based on the comprehensive evaluation results in <xref ref-type="table" rid="table-3">Table 3</xref>, the Random Forest (RF) algorithm exhibited the optimal performance in the fault diagnosis task, achieving an accuracy and F1-score of 95.83% each. It outperformed other algorithms significantly in both the recognition accuracy of various fault categories and overall stability, demonstrating excellent adaptability to the data characteristics of this task. The performance of Gradient Boosting (GB) was extremely close to that of RF, with its core metrics being basically on par. However, due to its serial iterative training mechanism, GB might be slightly less efficient when processing large-scale datasets, resulting in relatively poor adaptability to real-time diagnosis scenarios. Both Extremely Randomized Trees (ET) and Support Vector Machine (SVM) delivered good overall performance, with an accuracy of 91.67% for each. Among them, ET was slightly more sensitive to noise than RF because of its more extreme random feature splitting strategy. Although SVM performed stably on medium-dimensional features, the parameter tuning difficulty and computational complexity of its kernel function would increase significantly when dealing with high-dimensional complex nonlinear relationships. Constrained by the sample size and feature dimensions, the Multi-Layer Perceptron (MLP) failed to give full play to the complex pattern fitting capability of deep learning models, achieving an accuracy of 87.50%, and its generalization performance in small-sample scenarios needs to be improved. Although the K-Nearest Neighbors (KNN) algorithm achieved an accuracy of 91.67%, its prediction speed would decrease significantly with the increase of data volume due to its mechanism relying on inter-sample distance calculation, making it difficult to meet the real-time requirements of large-scale diagnosis tasks. In contrast, Gaussian Naive Bayes (GNB) achieved an accuracy of only 66.67%, performing relatively poorly in this task, because its strict assumption of feature independence was inconsistent with the strong correlation among features in the actual fault data.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Diagnostic performance of different algorithms.</title>
</caption>
<table>
<colgroup>
<col align="center" width="18mm"/>
<col align="center" width="22mm"/>
<col align="center" width="22mm"/>
<col align="center" width="17mm"/>
<col align="center" width="21mm"/> </colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Accuracy (%)</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F1-Score (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>RF</td>
<td>95.83</td>
<td>96.35</td>
<td>95.83</td>
<td>95.83</td>
</tr>
<tr>
<td>GB</td>
<td>95.83</td>
<td>96.30</td>
<td>95.83</td>
<td>95.80</td>
</tr>
<tr>
<td>ET</td>
<td>91.67</td>
<td>93.52</td>
<td>91.67</td>
<td>91.91</td>
</tr>
<tr>
<td>SVM</td>
<td>91.67</td>
<td>93.52</td>
<td>91.67</td>
<td>91.91</td>
</tr>
<tr>
<td>MLP</td>
<td>87.50</td>
<td>88.76</td>
<td>87.50</td>
<td>87.47</td>
</tr>
<tr>
<td>LR</td>
<td>79.17</td>
<td>84.63</td>
<td>91.67</td>
<td>79.16</td>
</tr>
<tr>
<td>KNN</td>
<td>91.67</td>
<td>91.67</td>
<td>91.67</td>
<td>91.67</td>
</tr>
<tr>
<td>GNB</td>
<td>66.67</td>
<td>68.75</td>
<td>66.67</td>
<td>62.97</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Based on the multi-metric performance comparison chart of models in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the Random Forest (RF) and Gradient Boosting (GB) models exhibited excellent performance across accuracy, precision, recall and F1-score, with their metric values remaining close and at high levels, which indicates that these two models have outstanding comprehensive performance in classification tasks. The Extremely Randomized Trees (ET) and Support Vector Machine (SVM) models also achieved relatively high values in all metrics, showing good performance. The Multi-Layer Perceptron (MLP) model performed moderately, while the Logistic Regression (LR) and Naive Bayes (NB) models yielded relatively low metric values. In particular, the NB model showed significantly lower accuracy, precision, recall and F1-score than other models, reflecting poor classification performance. In terms of the comparison of model performance radar charts, the RF and GB models stood out in all metrics, with their radar charts covering a wide area, which demonstrates their strong performance in multiple aspects. The ET model also reached a favorable level with balanced performance across all metrics. The radar charts of models such as SVM and K-Nearest Neighbors (KNN) covered a relatively smaller area, indicating slightly inferior performance. The NB model had the smallest coverage area in its radar chart, which further reflects its weak performance in these evaluation metrics. Overall, the above results are consistent with those in <xref ref-type="table" rid="table-2">Table 2</xref>. Ensemble learning models including RF, GB and ET achieved superior performance in this gearbox fault classification task, whereas simple models such as NB showed relatively insufficient performance.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Multi-metric performance comparison chart of models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-8.tif"/>
</fig>
<p>As can be seen from <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, the Random Forest (RF) model demonstrated remarkable advantages. In terms of performance, its F1-score reached a relatively high level, close to 0.95, indicating that the RF model could accurately identify various types of faults and achieve excellent classification results in the gearbox fault classification task. Meanwhile, in terms of training time, the RF model required a relatively short training duration, which was much lower than that of the Gradient Boosting (GB) model. This characteristic of ensuring high performance while maintaining high training efficiency enables the RF model to not only complete the model training process rapidly in practical applications, but also provide accurate results for gearbox fault diagnosis, thus achieving a favorable balance between training efficiency and classification performance.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Relationship chart between model training time and performance.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-9.tif"/>
</fig>
<p>As can be clearly observed from <xref ref-type="fig" rid="fig-10">Fig. 10</xref>, the Random Forest (RF) model exhibited exceptionally outstanding performance. The narrow box of its box plot indicates that the RF model had minimal performance fluctuations during the cross-validation process, demonstrating excellent stability. This means that the RF model could maintain stable and favorable classification results across different data subsets, with strong generalization ability. Meanwhile, the mean F1-score of the RF model stayed at a high level, which shows that in the gearbox fault classification task, the model could not only perform stably but also ensure high classification accuracy, enabling precise identification of various gearbox faults. Therefore, it serves as an excellent model choice that balances stability and accuracy for gearbox fault diagnosis tasks.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Distribution chart of model cross-validation performance.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-10.tif"/>
</fig>
<p>According to the analysis results of the confusion matrix of the RF model in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>, the overall performance of RF is excellent, with an F1-score of 0.958, indicating good comprehensive classification performance of the model. In the raw count matrix, all 8 samples of state S1 are correctly predicted; all 8 samples of state S2 are accurately classified; and the 8 samples of state S3 are also correctly identified. However, one sample of state S4 is misclassified as S1, with the remaining seven samples correctly predicted. From the row-normalized proportion matrix, the prediction accuracies of classes S1, S2, and S3 are all 1.00, meaning all samples of these three classes are correctly classified. The prediction accuracy of class S4 is 0.88, indicating a certain degree of misclassification for this class, with 12% of S4 samples incorrectly predicted as other categories.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Confusion matrix heatmap of RF.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81931-fig-11.tif"/>
</fig>
<p>Based on the comprehensive evaluation results of all categories, the RF model demonstrated significant superiority in the source domain fault diagnosis task. In terms of performance, its accuracy and F1-score both reached 95.83%. In the multi-metric performance comparison, all core indicators maintained high and balanced levels, leading to remarkable classification results for gearbox faults and high overall classification accuracy. In terms of stability, the RF model showed minimal performance fluctuations in cross-validation, with a narrow box in the box plot, which indicates that it could maintain stable and favorable classification results across different data subsets and possessed strong generalization ability. In terms of efficiency, the RF model required a short training time, which was much lower than that of the GB model. It could complete training rapidly and be put into practical applications, achieving a sound balance between training efficiency and classification performance. Therefore, it is an ideal model choice for gearbox fault diagnosis tasks.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>To address the challenges of insufficient diagnostic accuracy, weak generalization ability, and the difficulty of capturing fault information with single-dimensional features in gearboxes under complex working conditions, this study proposes a gearbox fault diagnosis method based on multi-dimensional feature extraction and Random Forest (RF). The main research conclusions are as follows:<list list-type="simple">
<list-item>
<label>(1)</label>
<p>The multi-dimensional feature extraction strategy significantly improves the completeness and distinguishability of fault information. By extracting key indicators such as root mean square values, this method captures the time-domain statistical characteristics, frequency-domain periodic components, and envelope amplitude variation features of gearbox vibration signals. It effectively overcomes the defect of incomplete characterization of fault information by single-dimensional features, providing more accurate feature support for subsequent diagnostic models.</p></list-item>
<list-item>
<label>(2)</label>
<p>The diagnostic model based on Random Forest achieves a good balance among diagnostic performance, training efficiency, and generalization ability. With strong modeling capability for complex data, the model exhibits excellent performance on the gearbox fault dataset. Compared with traditional models such as Support Vector Machine (SVM) and Multi-Layer Perceptron (MLP), its accuracy and F1-score are both improved by more than 4%. Meanwhile, it features short training time, which can meet the efficiency requirements of practical diagnostic scenarios.</p></list-item>
<list-item>
<label>(3)</label>
<p>The effectiveness and superiority of the proposed method are verified through experimental validation and comparative analysis. The results of feature distribution boxplots and PCA, t-SNE visualization demonstrate that the extracted features can effectively distinguish different fault states. The confusion matrix and multi-metric performance comparison show that the Random Forest model achieves an accuracy of 95.83%, with minimal performance fluctuation in cross-validation and strong generalization ability. Its identification accuracy and stability for gearbox faults are significantly superior to those of other comparative algorithms.</p></list-item>
</list></p>
<p>This paper takes gear crack faults as the typical research object, and focuses on verifying the fault diagnosis performance under different crack severities. The proposed multi-dimensional feature extraction and random forest diagnosis framework has good feature generalization ability and model adaptability. It can not only effectively distinguish the severity of cracks, but is also theoretically applicable to the classification and identification of various typical gearbox fault types such as tooth surface wear, tooth root fracture, and tooth surface scuffing. Future research will further expand the fault types, increase the number of samples and working conditions, carry out diagnosis experiments under multiple fault modes and various operating conditions, and comprehensively verify the universality and engineering applicability of the proposed method.</p>
</sec>
</body>
<back>
<ack>
<p>Author Chiming Guo sincerely acknowledges the financial support from the National Natural Science Foundation of China under contract number 71871220 (Titled: Dynamic Maintenance Optimization of Complex Systems Considering Correlation and Task Constraints).</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Yu Zhang, Shihan Tan; methodology, Qiwei Hu; software, Chiming Guo; validation, Guangyao Lian; formal analysis, Congying Dun; investigation, Shihan Tan, Qiwei Hu; resources, Chiming Guo; data curation, Congying Dun; writing&#x2014;original draft preparation, Yu Zhang; writing&#x2014;review and editing, Shihan Tan, Guangyao Lian, Congying Dun; visualization, Qiwei Hu; supervision, Congying Dun; project administration, Chiming Guo. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on request from the authors. The data that support the findings of this study are available from the Corresponding Author, [Chiming Guo], upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>V</given-names></string-name>, <string-name><surname>Sarangi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>OP</given-names></string-name></person-group>. <article-title>Gearbox fault diagnosis: a higher order moments approach</article-title>. <source>Measurement</source>. <year>2023</year>;<volume>210</volume>:<fpage>112489</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2023.112489</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Damou</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ratni</surname> <given-names>A</given-names></string-name>, <string-name><surname>Benazzouz</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Intelligent multi-fault identification and classification of defective bearings in gearbox</article-title>. <source>Adv Mech Eng</source>. <year>2024</year>;<volume>16</volume>(<issue>4</issue>):<fpage>16878132241246673</fpage>. doi:<pub-id pub-id-type="doi">10.1177/16878132241246673</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A new monitoring technology for bearing fault detection in high-speed trains</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>14</issue>):<fpage>6392</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23146392</pub-id>; <pub-id pub-id-type="pmid">37514687</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Bearing-fault-feature enhancement and diagnosis based on coarse-grained lattice features</article-title>. <source>Sensors</source>. <year>2024</year>;<volume>24</volume>(<issue>11</issue>):<fpage>3540</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s24113540</pub-id>; <pub-id pub-id-type="pmid">38894331</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Research on fault feature extraction method of rolling bearing based on NMD and wavelet threshold denoising</article-title>. <source>Shock Vib</source>. <year>2018</year>;<volume>2018</volume>:<fpage>9495265</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2018/9495265</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bie</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Research on gearbox impact feature extraction method based on improved ESMD</article-title>. <source>Insight</source>. <year>2022</year>;<volume>64</volume>(<issue>1</issue>):<fpage>20</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1784/insi.2022.64.1.20</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>The feature extraction method based on quadratic wavelet packet energy entropy and t-SNE for bearing fault diagnosis</article-title>. <source>Proc Inst Mech Eng Part C J Mech Eng Sci</source>. <year>2025</year>;<volume>239</volume>(<issue>2</issue>):<fpage>520</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1177/09544062241283331</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>The fault diagnosis of rolling bearing based on WPD and TPOT</article-title>. In: <conf-name>Proceedings of the 2019 Chinese Automation Congress (CAC); 2019 Nov 22&#x2013;24</conf-name>; <publisher-loc>Hangzhou, China</publisher-loc>. p. <fpage>1029</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CAC48633.2019.8996312</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Feature extraction for rolling element bearing weak fault based on MOMEDA and ICEEMDAN</article-title>. <source>J Vibroeng</source>. <year>2018</year>;<volume>20</volume>(<issue>6</issue>):<fpage>2352</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.21595/jve.2018.19309</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Fault feature extraction of rolling element bearing based on TPE-EVMD</article-title>. <source>Measurement</source>. <year>2021</year>;<volume>183</volume>:<fpage>109880</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2021.109880</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A fault diagnosis method of bogie axle box bearing based on spectrum whitening demodulation</article-title>. <source>Sensors</source>. <year>2020</year>;<volume>20</volume>(<issue>24</issue>):<fpage>7155</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s20247155</pub-id>; <pub-id pub-id-type="pmid">33327394</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Spectral distribution decomposition and its application in gearbox fault diagnosis</article-title>. <source>Mech Mach Theory</source>. <year>2026</year>;<volume>220</volume>:<fpage>106359</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.mechmachtheory.2026.106359</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A convex optimization difference analysis model for intelligent fault detection and diagnosis of gearboxes</article-title>. <source>Mech Syst Signal Process</source>. <year>2026</year>;<volume>251</volume>:<fpage>114207</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2026.114207</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A study of a domain-adaptive LSTM-DNN-based method for RUL prediction of planetary gearbox</article-title>. <source>Processes</source>. <year>2023</year>;<volume>11</volume>(<issue>8</issue>):<fpage>2245</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.3390/pr11072002</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hogea</surname> <given-names>E</given-names></string-name>, <string-name><surname>Onchi&#x015F;</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>LogicLSTM: logically-driven long short-term memory model for fault diagnosis in gearboxes</article-title>. <source>J Manuf Syst</source>. <year>2024</year>;<volume>77</volume>:<fpage>892</fpage>&#x2013;<lpage>902</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jmsy.2024.10.003</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Gearbox fault diagnosis based on Gramian angular field and TLCA-MobileNetV3 with limited samples</article-title>. <source>Int J Metrol Qual Eng</source>. <year>2024</year>;<volume>15</volume>:<fpage>15</fpage>. doi:<pub-id pub-id-type="doi">10.1051/ijmqe/2024004</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>E</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A novel dynamic predictive maintenance framework for gearboxes utilizing nonlinear Wiener process</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>35</volume>(<issue>12</issue>):<fpage>126210</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad762e</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Efficient gearbox fault diagnosis based on improved multi-scale CNN with lightweight convolutional attention</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>9</issue>):<fpage>2636</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25092636</pub-id>; <pub-id pub-id-type="pmid">40363076</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Gearbox fault diagnosis method based on lightweight channel attention mechanism and transfer learning</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>743</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-023-50826-6</pub-id>; <pub-id pub-id-type="pmid">38185699</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>CD</given-names></string-name>, <string-name><surname>Prosvirin</surname> <given-names>AE</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Construction of a sensitive and speed invariant gearbox fault diagnosis model using an incorporated utilizing adaptive noise control and a stacked sparse autoencoder-based deep neural network</article-title>. <source>Sensors</source>. <year>2021</year>;<volume>21</volume>(<issue>1</issue>):<fpage>18</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s21010018</pub-id>; <pub-id pub-id-type="pmid">33375085</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Desai</surname> <given-names>A</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sheng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Phillips</surname> <given-names>C</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Prognosis of wind turbine gearbox bearing failures using SCADA and modeled data</article-title>. <source>Annu Conf PHM Soc</source>. <year>2020</year>;<volume>12</volume>(<issue>1</issue>):<fpage>10</fpage>. doi:<pub-id pub-id-type="doi">10.36001/phmconf.2020.v12i1.1292</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>A bearing fault diagnosis method based on an incremental broad federated learning system</article-title>. <source>Eng Res Express</source>. <year>2026</year>;<volume>8</volume>(<issue>3</issue>):<fpage>035234</fpage>. doi:<pub-id pub-id-type="doi">10.1088/2631-8695/ae39a5</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A fault diagnosis method based on dual-layer adaptive personalized federated learning in wind turbines</article-title>. <source>IFAC Pap</source>. <year>2025</year>;<volume>59</volume>(<issue>20</issue>):<fpage>1848</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ifacol.2025.11.427</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Na</surname> <given-names>J</given-names></string-name>, <string-name><surname>Fung</surname> <given-names>RF</given-names></string-name></person-group>. <article-title>Fault feature extraction based on combination of envelope order tracking and cICA for rolling element bearings</article-title>. <source>Mech Syst Signal Process</source>. <year>2018</year>;<volume>113</volume>:<fpage>131</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2017.03.050</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Fractional envelope analysis for rolling element bearing weak fault feature extraction</article-title>. <source>IEEE/CAA J Autom Sin</source>. <year>2017</year>;<volume>4</volume>(<issue>2</issue>):<fpage>353</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JAS.2016.7510166</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Random forest grid fault prediction based on genetic algorithm optimization</article-title>. <source>Front Phys</source>. <year>2025</year>;<volume>13</volume>:<fpage>1480749</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fphy.2025.1480749</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cerrada</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zurita</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cabrera</surname> <given-names>D</given-names></string-name>, <string-name><surname>S&#x00E1;nchez</surname> <given-names>RV</given-names></string-name>, <string-name><surname>Art&#x00E9;s</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Fault diagnosis in spur gears based on genetic algorithm and random forest</article-title>. <source>Mech Syst Signal Process</source>. <year>2016</year>;<volume>70</volume>:<fpage>87</fpage>&#x2013;<lpage>103</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2015.08.030</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Quiroz</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Mariun</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mehrjou</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Izadi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Misron</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mohd Radzi</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Fault detection of broken rotor bar in LS-PMSM using random forests</article-title>. <source>Measurement</source>. <year>2018</year>;<volume>116</volume>:<fpage>273</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2017.11.004</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Transformer fault diagnosis based on the improved QPSO and random forest</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>35</volume>(<issue>9</issue>):<fpage>096206</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad574c</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>