<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">57103</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.057103</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Contribution Tracking Feature Selection (CTFS) Based on the Fusion of Sparse Autoencoder and Mutual Information</article-title>
<alt-title alt-title-type="left-running-head">Contribution Tracking Feature Selection (CTFS) Based on the Fusion of Sparse Autoencoder and Mutual Information</alt-title>
<alt-title alt-title-type="right-running-head">Contribution Tracking Feature Selection (CTFS) Based on the Fusion of Sparse Autoencoder and Mutual Information</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Yu</surname><given-names>Yifan</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wang</surname><given-names>Dazhi</given-names></name><email>wangdazhi1@mail.neu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Yanhua</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Hongfeng</given-names></name></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Min</given-names></name></contrib>
<aff><institution>College of Information Science and Engineering, Northeastern University</institution>, <addr-line>Shenyang, 110004</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Dazhi Wang. Email: <email>wangdazhi1@mail.neu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>19</day><month>12</month><year>2024</year>
</pub-date>
<volume>81</volume>
<issue>3</issue>
<fpage>3761</fpage>
<lpage>3780</lpage>
<history>
<date date-type="received">
<day>08</day>
<month>8</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>9</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_57103.pdf"></self-uri>
<abstract>
<p>For data mining tasks on large-scale data, feature selection is a pivotal stage that plays an important role in removing redundant or irrelevant features while improving classifier performance. Traditional wrapper feature selection methodologies typically require extensive model training and evaluation, which cannot deliver desired outcomes within a reasonable computing time. In this paper, an innovative wrapper approach termed Contribution Tracking Feature Selection (CTFS) is proposed for feature selection of large-scale data, which can locate informative features without population-level evolution. In other words, fewer evaluations are needed for CTFS compared to other evolutionary methods. We initially introduce a refined sparse autoencoder to assess the prominence of each feature in the subsequent wrapper method. Subsequently, we utilize an enhanced wrapper feature selection technique that merges Mutual Information (MI) with individual feature contributions. Finally, a fine-tuning contribution tracking mechanism discerns informative features within the optimal feature subset, operating via a dominance accumulation mechanism. Experimental results for multiple classification performance metrics demonstrate that the proposed method effectively yields smaller feature subsets without degrading classification performance in an acceptable runtime compared to state-of-the-art algorithms across most large-scale benchmark datasets.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Feature selection</kwd>
<kwd>contribution tracking</kwd>
<kwd>sparse autoencoders</kwd>
<kwd>mutual information</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2021YFB3300900</award-id>
</award-group>
<award-group id="awg2">
<funding-source>NSFC</funding-source>
<award-id>92267206</award-id>
</award-group>
<award-group id="awg3">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>72201052</award-id>
<award-id>62032013</award-id>
<award-id>62173076</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Fundamental Research Funds for the Central Universities</funding-source>
<award-id>N2204017</award-id>
</award-group>
<award-group id="awg5">
<funding-source>State Key Laboratory of Synthetical Automation for Process Industries</funding-source>
<award-id>2013ZCX11</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid development of information technology, data mining, image processing, bioinformatics, and other fields have generated massive and high-dimensional data. High-dimensional datasets dramatically increase the algorithm&#x2019;s demands in terms of time and space [<xref ref-type="bibr" rid="ref-1">1</xref>], and there is a side effect on the efficiency or effectiveness of the learning system. Moreover, for some learning tasks, such as classification and regression, computational analysis is feasible in low-dimensional space but becomes very difficult in high-dimensional space. To address this problem, one can use either feature extraction or feature selection, two standard dimensionality reduction methods. Feature extraction generally utilizes linear and nonlinear mapping to obtain low-dimensional features. However, for some specific issues, such as genomic datasets, this method may overlook the inherent physical properties of the original genes. Feature selection, also called attribute subset selection [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>], aims to minimize the dimensionality of the feature set while approaching, maintaining, or even improving the accuracy of the classification model compared to the entire feature set. In addition, feature selection removes irrelevant and redundant features, improving the model&#x2019;s generalization ability while avoiding dimensionality catastrophe [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>Feature selection is commonly viewed as a combinatorial optimization problem, mainly involving deciding whether to include each feature from the entire feature set. For a dataset with <italic>m</italic> features, an exhaustive search algorithm requires <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> iterations, representing a non-trivial task categorized as the NP-Hard problem [<xref ref-type="bibr" rid="ref-6">6</xref>]. Over the past few years, due to the widespread application of some effective deep neural networks and classification algorithms [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>], feature selection has become a hot topic in this field.</p>
<p>In general, existing supervised and unsupervised feature selection methods can be classified as filter methods [<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>], wrapper methods [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>], and embedded methods [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. Moreover, researchers have attempted to achieve higher classification performance by putting meta-heuristic algorithms into the feature selection problem to search for optimal or near-optimal feature subsets [<xref ref-type="bibr" rid="ref-18">18</xref>], which serve as a wrapper method for addressing the feature selection problem by necessitating a classifier&#x2019;s specification to assess the solution&#x2019;s quality.</p>
<p>A notable concern in feature selection revolves around selecting features from large sets of high-dimensional datasets. This challenge remains for meta-heuristic algorithms due to their inherent limitations in attaining optimal solutions within a restricted time. Therefore, an applicable and efficient meta-heuristic algorithm must address the feature selection problem with fewer evaluations, higher accuracy, and faster speed. In addition, few studies have simultaneously explored the impact of individual features on the overall performance of a feature subset and non-linear relationships between features. There is a lack of related research on whether this impact helps the heuristic algorithm efficiently discover the optimal feature subset. Based on the above analysis, this paper proposes a novel approach called CTFS. The primary contributions of this paper are outlined below:
<list list-type="bullet">
<list-item>
<p>We employ an unsupervised learning methodology, utilizing an autoencoder network to unveil non-linear relationships between features. Consequently, this approach alleviates missing label information and errors in classification results.</p></list-item>
<list-item>
<p>We adopt the deep learning structure to reduce the extra time and algorithmic complexity for the subsequent wrapper, making it more efficient and less computationally demanding.</p></list-item>
<list-item>
<p>Based on the autoencoder network, sparse training is incorporated to expedite network training and diminish memory occupation.</p></list-item>
<list-item>
<p>Integrating the feature contribution obtained from the network training in the first stage with the relationship between features and labels, a contribution tracking approach is proposed to find the optimal or near-optimal feature subset efficiently.</p></list-item>
</list></p>
<p>The subsequent sections of this paper are organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the meta-heuristic algorithm based on the wrapper method and related works on the sparse autoencoder network. <xref ref-type="sec" rid="s3">Section 3</xref> thoroughly introduces the framework of the sparse autoencoder network. <xref ref-type="sec" rid="s4">Section 4</xref> proposes a novel wrapper feature selection method based on a contribution tracking mechanism. Computational comparison experiments of the proposed algorithm will be reported in <xref ref-type="sec" rid="s5">Section 5</xref>. The last section will conclude the works of this paper and give the subsequent schemes in the research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>The feature selection process consists of four main components: subset generation, subset evaluation, stopping criteria, and result validation. The method proposed in this paper is a wrapper-based technique, so filter-based and embedded feature selection methods are not part of the work in this paper. Different feature selection problems in many domains can be solved using meta-heuristic algorithms based on wrapper methods [<xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-22">22</xref>]. The following is a review of some of the metaheuristics that have been categorized according to the class of algorithms used, and several of the metaheuristics have been used in comparative experiments with our proposed algorithm. In addition, the research in this paper is related to autoencoder networks and sparse training, and an overview of deep neural network methods for feature selection problems is also presented at the end of the section.</p>
<p>PSO algorithms are conceptually simple, easy to implement, and have fast convergence; many PSO-based algorithms have been used for feature selection problems. A novel PSO-based feature selection approach was designed in [<xref ref-type="bibr" rid="ref-23">23</xref>], which can improve the population&#x2019;s quality at each generation to jump out of the local optimum to some extent. The proposed method was combined with the filter Relief algorithm to get the weights of each feature and generate more optimal solutions. In addition, an agent model was used to select the best solution regarding diversity and convergence. Moreover, feature selection based on evolutionary multitasking is also a very hot topic. Wang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed a novel PSO-based multi-task framework to achieve the information shared, which divided the initial population into two subpopulations. Extensive experiments showed the strong competitiveness of the approach compared with other algorithms. An effective FS method based on the idea of multitasking with knowledge transfer between different search space was developed in [<xref ref-type="bibr" rid="ref-25">25</xref>] on high-dimensional data. In addition, promising features were distinguished by a knee point selection scheme, and a novel variable-range strategy effectively reduced the search space of the population.</p>
<p>Harris Hawk optimization is a newly introduced swarm intelligence optimizer that can be used to solve a variety of combinatorial optimization problems. Peng et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] improved the Harris Hawk algorithm by using a hierarchical mechanism that speeded up its operation without increasing its complexity, allowing the algorithm a shorter time and higher accuracy. Two limitations of Harris Hawk optimization were addressed in [<xref ref-type="bibr" rid="ref-27">27</xref>] by introducing chaotic dyadic in the initial stage to improve the population diversity of search agents. Then, a simulated annealing strategy was introduced in each iteration to reduce the risk of falling into a local optimum. Finally, the algorithm was applied to the feature selection problem, and experimental results on some tumor datasets demonstrated the superiority of the algorithm.</p>
<p>In addition to the two aforementioned optimization paradigms, various feature selection problems have been solved using other meta-heuristic algorithms. A modified binary Gray Wolf algorithm with two-stage mutation was developed in [<xref ref-type="bibr" rid="ref-28">28</xref>]. The two-phase mutation enhanced the exploitation capability of the algorithm. The experimental results showed the validity of the algorithm and good performance. Zhou et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] proposed a genetic algorithm combining the correlation between feature-to-feature and feature-to-label, and the algorithm reduced the generation of poor solutions and improved the efficiency of the evolutionary process by utilizing correlations to guide the improvement of the crossover and mutation operators in the genetic algorithms. The algorithm was demonstrated to have good classification accuracy on four artificial datasets and six real datasets.</p>
<p>The main limitation of the aforementioned current wrapper methods is that their score function depends on training and predicting at least one model per iteration. Thus, this approach has a high computational cost. Also, wrappers are classifier-specific. Therefore, they may not obtain satisfactory results if the resulting selection is used with different classifiers other than the one used in training and become very prone to overfit.</p>
<p>On high-dimensional datasets, some fully connected layers of deep neural networks can prominently slow down the training speed and cause too much memory usage. Because of this, Han et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] utilized a three-step method to prune redundant connections and achieved sparsity of weights by learning significant connections without degrading accuracy. Inspired by network science, Mocanu et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] proposed a SET algorithm, which first introduced an Erd&#x00F6;s-R&#x00E9;nyi random graph to initialize the layer-to-layer connections between neurons. A dynamic updating approach was used to increase the generalizability of this structure in the network, and to prove its effectiveness, the article went on to perform supervised and unsupervised learning with different neural networks on multiple datasets. Atashgahi et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] built on the approach of Han et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] and used it for an unsupervised feature selection problem to achieve an optimal trade-off between classification and clustering accuracy, runtime, and maximum memory usage.</p>
<p>Given that the stochastic exploration of sparse topology requires more cycles in the drop-and-grow session [<xref ref-type="bibr" rid="ref-31">31</xref>], Sokar et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] improved the drop-and-grow session. Work was put into the regrowth of the new connections. In this paper, we have improved on the work of Sokar et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] by adjusting its input part. Meanwhile, connections in the input or output layer are re-estimated only by corresponding magnitudes during the drop operation.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Sparse Autoencoder</title>
<p>The autoencoder network employing sparse training is designed to capture the nonlinear relationships between features, which matters for our CTFS approach to quantitatively evaluate the contributions of individual features. In addition, it&#x2019;s noteworthy that this mode of training the network can significantly reduce the runtime as well as the memory occupation [<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
<sec id="s3_1">
<label>3.1</label>
<title>Framework</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Data Input</title>
<p>Assume a dataset with a sample size of <italic>n</italic> and dimension of <italic>m</italic>. The task is to obtain the contribution of each feature in the first stage. These contributions are used to report the importance level of each feature in the subsequent wrapper and prepare for the fusion of mutual information.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Initialization</title>
<p>Sparse Autoencoder (AE) is based on the structure of a neural network. Given the demands of short training time and low memory occupation, this paper adopts the model with the least hidden layer, using sparse connections instead of full ones. The sparse level <italic>&#x03B5;</italic> of the network is obtained from <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>&#x03B5;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (<italic>l</italic> &#x003D; 1, 2) is the sparse weight matrix of hidden layers, <italic>m</italic> is the number of features, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denotes the number of hidden layer neurons. The network is initialized with sparse weights uniformly distributed on the neurons.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Noise Input</title>
<p>To learn more robust features, Gaussian noise is added to perturb the original data:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>, <bold>x</bold> is the initial input vector from the dataset, <italic>&#x03B8;</italic> denotes the noise factor that controls the level of corruption, and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a Gaussian noise.</p>
<p>After obtaining the corrupted input, the autoencoder maps it to the latent representation <bold>e</bold>, followed by mapping <bold>e</bold> back to the reconstructed vector <bold>d</bold>. Here, the hidden representation <bold>e</bold> as well as the reconstructed vector <bold>d</bold> can be calculated from <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext mathvariant="bold">d</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>s</italic> is the neuron activation function and <bold>b</bold>, <bold>b<sup>&#x2032;</sup></bold> are the bias vectors of corresponding layer.</p>
<p>Next, the weight parameters are updated using a stochastic gradient descent algorithm, which minimizes the average reconstruction error:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">d</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Training Process</title>
<p>Since the connections between initialized neurons are randomly generated, to enhance its generalizability, the improved SET approach proposed by Sokar et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] is introduced here, and the overall structure of AE is depicted in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. It proceeds as follows: after updating the weights using each batch of data, we adjust the sparse topology using the drop-and-regrow cycle. Given a score <italic>&#x03B1;</italic>, connections with less than the minimum level <italic>&#x03B1;</italic> are dropped from the sparse weight matrix <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and the same fraction of connections with weight 0 are regrown. This drop-and-regrow approach maintains the overall performance, which was described as well as proved by Han et al. [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Comprehensive overview of the sparse AE network. A sparse autoencoder is initialized with uniformly distributed sparse connections. During sparse training, the connections are redistributed in the most important neurons at iteration <italic>t</italic> during the &#x201C;drop-and-regrow&#x201D; cycle</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-1.tif"/>
</fig>
<p>Especially, new connections are not randomly regrown. Sokar et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] gave a fast focus on new connections to informative features. The gradient of the loss concerning the output neuron <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mrow><mml:mtext mathvariant="bold">d</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula> is used to measure the sensitivity of the reconstruction error to the reconstructed output. Following that, we obtain <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, defining the contribution of the <italic>i</italic>th output neuron, denoted as (<inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>), at iteration <italic>t</italic>:</p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo>|</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mrow><mml:mtext mathvariant="bold">d</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munderover><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Here (<inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>) is initially set to 0.<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the magnitude of the incoming connection between the <italic>j</italic>th hidden neuron and the <italic>i</italic>th output neuron. <italic>&#x03BB;</italic> is a hyperparameter coefficient. Similarly, the same method is applied to the input neurons (<inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) by replacing the magnitude of output weights with input ones <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula>. This distributes the positions of the regrown new connections over these contributing neurons. This modified drop-and-regrow approach dynamically optimizes the sparse topology of the network, ultimately using (<inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) as the contribution of each input neuron. The pseudocode can be found in Algorithm 1:</p>
<fig id="fig-6">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-6.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Wrapper Feature Selection Using a Contribution Tracking Method</title>
<p>Wrapper feature selection methods usually specify a classifier, which evaluates the performance of a candidate feature subset by calculating the predicted values of the classification model. A large number of iterations are required to find the target feature subset that possesses the optimal performance when dealing with high-dimensional datasets. The computational time is considerable. Therefore, the primary task of the wrapper feature selection algorithm is to achieve a better balance between classification performance and running time.</p>
<p>The proposed algorithm employs a contribution tracking mechanism called CTFS, which fully considers these two indicators mentioned above. Specifically, the contribution of each feature incorporates the mutual information metrics from the filter method and the feature importance metrics from the AE network. Following that, a novel wrapper algorithm combining the feature contributions with its corresponding classification accuracy is given. Compared to the traditional wrapper methods, CTFS targets informative features through the contribution tracking strategy instead of finally reaching convergence through population-level evolution, and a classifier is necessary during the training time, which also varies from filter methods. The general framework of the proposed algorithm is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Flowchart of the proposed CTFS algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-2.tif"/>
</fig>
<sec id="s4_1">
<label>4.1</label>
<title>Score Function</title>
<p>A score function is required to evaluate the merit of a candidate feature subset. In this paper, we use some performance metrics in the learning machine as its score function. The specific mathematical description is as follows:</p>
<p>Suppose there are <italic>m</italic> features in the original data, and one of the candidate feature subsets is represented by <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. <italic>x</italic><sub><italic>i</italic></sub> takes the value 0 or 1, which means whether the <italic>i</italic>th feature is selected or not. The best feature subset can be discovered by finding the highest score in the subset of candidate features, which can be denoted by the mathematical <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>:</p>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Here <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is related to the performance metrics of the classifier.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Generation of Initial Feature Subset</title>
<p>In this paper, many initial feature subsets are required for subsequent calculation. For some of the commonly used methods of generating initial feature subsets, such as the one-dimensional chaotic mapping method, taking into account that it is more sensitive to the initial conditions as well as the parameters, which will lead to a decrease in the uniformity and stochasticity of the sequences; in other words, each feature does not appear at the same frequency, and some features appear repeatedly in the subset while others never appear, which result in a negative influence on the acquisition of the target feature subset. This research develops a method to initialize the generation of the feature subset, which ensures that each feature has the same frequency of appearing in the initial randomly generated feature subset.</p>
<p>In Algorithm 2, the parameters <italic>LL</italic> and <italic>UL</italic> denote the upper and lower limits for dividing the set of features; the closer to the upper limit, the smaller the number of features in the initial feature subsets generated. <italic>C</italic> represents the number of repetitions. The parameter <italic>&#x0394;</italic> is the tolerance for accepting poor initial feature subsets; when the classification accuracy of an initial feature subset differs too much from the best, then the corresponding initial feature subset will be deleted, and this parameter controls the number of initial feature subsets generated at the end.</p>
<fig id="fig-7">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-7.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Contribution Tracking Strategy</title>
<p>Current wrappers often do not adequately consider the influence of a single feature on the accuracy of a classification model. As a result, an improved wrapper feature selection algorithm has been developed to address this limitation, which is inspired by the fact that after obtaining the optimal or suboptimal feature subsets, we can know exactly which features work for classification accuracy. Accordingly, we will accumulate the dominance corresponding to the features that appear in the feature subsets with high classification accuracy. Given that unsupervised learning does not involve label information, this paper fuses the concept of mutual information to increase the information between features and labels.</p>
<p>We first introduce the concept of mutual information [<xref ref-type="bibr" rid="ref-10">10</xref>]. Suppose <italic>X</italic> and <italic>Y</italic> are two discrete random variables, and their mutual information <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> can be defined as:</p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the marginal probability distribution function of <italic>X</italic> and <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the joint probability distribution function of <italic>X</italic> and <italic>Y</italic>. We can find the relationship <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>M</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> between features and labels through the method above.</p>
<p>To facilitate the description, the symbol <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is employed to indicate the joint action between the contribution (<inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) of the feature <italic>i</italic> and the relationship <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>M</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with labels, as shown in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>:</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mi>M</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Since the same feature in two different initial feature subsets shares an equal contribution, we adjust the original <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref> to reflect the variations of the same feature in different feature subsets by introducing the target value <italic>F</italic>, that is:</p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>l</mml:mi><mml:mi>n</mml:mi><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p><p>We quantitatively evaluate the possibility that subset <italic>I</italic> contains the feature <italic>i</italic> by thoroughly combining both the joint action <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the <italic>i</italic>th feature and the objective value of the feature subset. This formula also accounts for the fact that various candidate feature subsets share different <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> values for the same feature condition. That is to say, different candidate feature subsets affect the inclusion of feature <italic>i</italic> in the optimal feature subset.</p>
<p>Typically, if multiple local optimal solutions with good performance have been achieved, the optimal feature subsets will be easier to obtain after certain processing operations. The specific processing operations based on the contribution tracking method will be given subsequently.</p>
<p>This research generates quantities of initial feature subsets. For simplicity of description, suppose these subsets are <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, respectively, then the probability of the feature <italic>i</italic> appearing in the <italic>j</italic>th subset is <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The cumulative effect of feature <italic>i</italic> in the <italic>q</italic> candidate feature subsets can be obtained in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:</p>
<p><disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where the symbol <italic>C</italic><sub><italic>i</italic></sub> is used to denote this cumulative effect, which is the aforementioned contribution tracking mechanism. Then, it is necessary to find a mean value for this cumulative impact, which is used to quantitatively characterize the probability of the feature <italic>i</italic> appearing in the target feature subset. <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> can be achieved by <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>:</p>
<p><disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <italic>O</italic><sub><italic>i</italic></sub> denotes the cumulative occurrences of feature <italic>i</italic> in <italic>q</italic> candidate feature subsets. From <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>, we obtain the <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of each feature in the target feature subset, but the number of features cannot be determined. To address this problem, we design a greedy approach. Initially, the dominance of the features is sorted in descending order, and then the features are added sequentially in descending feature order. Since the dimension of the original dataset is <italic>m</italic>, <italic>m</italic> additions in total are required. Every time a feature is added, the value of the corresponding objective function is calculated. Finally, we select the feature subset with the largest objective function value as the target one. The pseudocode for obtaining the target feature subset is shown in Algorithm 3:</p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-8.tif"/>
</fig>
<p>The objective feature subset is finally obtained through the above approach, and the overall framework structure of the proposed CTFS algorithm is depicted in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Structure of the proposed CTFS algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-3.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experiments and Analysis</title>
<sec id="s5_1">
<label>5.1</label>
<title>Test Datasets</title>
<p>To verify the performance of the proposed algorithm, eight datasets from the UCI repository and two datasets from the literature [<xref ref-type="bibr" rid="ref-34">34</xref>] are selected. Given the huge amount of data in real life, the characteristics of the selected datasets are also large sample sizes and high dimensions. These datasets involve life, microarrays, computer science, business, and other fields. <xref ref-type="table" rid="table-1">Table 1</xref> lists their basic information: the name of datasets (Dataset), the number of features (#of features), the number of instances (#of instances), and the number of classes (#of classes).</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Basic information on test datasets</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>#of features</th>
<th>#of instances</th>
<th>#of classes</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>36</td>
<td>3196</td>
<td>2</td>
</tr>
<tr>
<td>Spambase</td>
<td>57</td>
<td>4601</td>
<td>2</td>
</tr>
<tr>
<td>Isolet</td>
<td>617</td>
<td>1560</td>
<td>26</td>
</tr>
<tr>
<td>CNAE9</td>
<td>856</td>
<td>1080</td>
<td>9</td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>241</td>
<td>4464</td>
<td>2</td>
</tr>
<tr>
<td>Letter</td>
<td>16</td>
<td>20,000</td>
<td>26</td>
</tr>
<tr>
<td>Lung2</td>
<td>3312</td>
<td>203</td>
<td>5</td>
</tr>
<tr>
<td>Prostate1</td>
<td>5966</td>
<td>102</td>
<td>2</td>
</tr>
<tr>
<td>Gutenberg</td>
<td>8266</td>
<td>590</td>
<td>8</td>
</tr>
<tr>
<td>Ovarian</td>
<td>15,154</td>
<td>253</td>
<td>2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Comparative Methods and Parameter Settings</title>
<p>The binary HHO, GWO, and PSO algorithms perform well in solving the feature selection problem. This research selects the following algorithms for comparison with the proposed CTFS algorithm: all available features (Full), the binary HHO (BHHO), the correlation-guided updating strategy and surrogate-assisted PSO (CUS-SPSO) [<xref ref-type="bibr" rid="ref-23">23</xref>], and the Grey-Wolf algorithm integrating a two-phase mutation (TMGWO) [<xref ref-type="bibr" rid="ref-28">28</xref>]. These comparison algorithms are re-programmed based on the algorithm flow and pseudocode in the literature.</p>
<p>To illustrate the generalizability of the algorithm and analyze the effect of the classifier on performance, this research selects two classifiers, KNN and Bayesian, where K is set to 5. For each dataset, 80% of the samples are used as training sets randomly, and the remaining 20% are selected as test sets. During the training process, k-fold cross-validation is employed, where k is set to 5. The specific procedure is to iteratively rotate the training and test sets k times in sequence. After the training, the selected features will be evaluated on the test set samples to obtain the corresponding accuracy, recall rate, and F1 score.</p>
<p>The parameters of the algorithms proposed in this paper, as well as the feature selection methods used for comparison, are given in <xref ref-type="table" rid="table-2">Table 2</xref>. For the swarm intelligence algorithms, we set the maximum number of iterations to 100 and the population size to 30. The settings of other parameter values are adjusted according to the characteristics of each algorithm.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Parameter settings</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Population size</th>
<th>MaxIter</th>
<th>Parameter values</th>
</tr>
</thead>
<tbody>
<tr>
<td>BHHO</td>
<td>30</td>
<td>100</td>
<td><bold><italic>&#x2014;</italic></bold></td>
</tr>
<tr>
<td>TMGWO</td>
<td>30</td>
<td>100</td>
<td><italic>a</italic> &#x003D; 2 &#x2212; 2 <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:math></inline-formula> (<italic>iter</italic>/<italic>MaxIter</italic>), <italic>A</italic> &#x003D; [0,2], <italic>M</italic><sub><italic>p</italic></sub> &#x003D; 0.5</td>
</tr>
<tr>
<td>CUS-SPSO</td>
<td>30</td>
<td>100</td>
<td><italic>c</italic><sub><italic>1</italic></sub> &#x003D; <italic>c</italic><sub><italic>2</italic></sub> &#x003D; 1.5, <italic>&#x03C9;</italic> &#x003D; 0.9 - 0.5 <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:math></inline-formula> (<italic>iter</italic>/<italic>MaxIter</italic>),<break/><italic>n</italic><sub><italic>c</italic></sub> &#x003D; 2, <italic>A</italic> &#x003D; 0.15, <italic>B</italic> &#x003D; 0.05</td>
</tr>
<tr>
<td>CTFS</td>
<td><bold><italic>&#x2014;</italic></bold></td>
<td><bold><italic>&#x2014;</italic></bold></td>
<td><italic>&#x03B5;</italic> &#x003D; 0.8, <italic>&#x03B8;</italic> &#x003D; 0.2, <italic>n</italic><sup><italic>h</italic></sup> &#x003D; 200, <italic>&#x03BB;</italic> &#x003D; 0.9, <italic>epoch</italic> &#x003D; 10<break/><italic>LL</italic>&#x003D; 2, <italic>UL</italic> &#x003D; 10, <italic>C</italic> &#x003D; 5, <italic>&#x0394;</italic> &#x003D; 2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Results and Discussions</title>
<p><xref ref-type="table" rid="table-3">Tables 3</xref> and <xref ref-type="table" rid="table-4">4</xref> list the average classification accuracies of the proposed CTFS method and the other four comparison algorithms based on ten independent runs on the test sets. The epoch of the SAE network is also set to 10, which is a balance between the reconstruction error of the test set and the training time of the network. Moreover, the average rank approach is adopted to express the average ranking performance of the methods, which is denoted as AR.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Classification accuracy obtained by KNN</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>94.79</td>
<td>95.62</td>
<td>96.32</td>
<td><bold>97.13</bold></td>
<td>95.33</td>
</tr>
<tr>
<td>Spambase</td>
<td>89.59</td>
<td>90.20</td>
<td><bold>90.41</bold></td>
<td>89.91</td>
<td>89.93</td>
</tr>
<tr>
<td>Isolet</td>
<td>79.32</td>
<td>81.28</td>
<td>82.98</td>
<td>83.14</td>
<td><bold>85.71</bold></td>
</tr>
<tr>
<td>CNAE9</td>
<td>83.56</td>
<td>83.51</td>
<td>82.66</td>
<td>84.26</td>
<td><bold>87.82</bold></td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>98.14</td>
<td>98.45</td>
<td>98.47</td>
<td><bold>98.69</bold></td>
<td>98.66</td>
</tr>
<tr>
<td>Letter</td>
<td>95.24</td>
<td>95.32</td>
<td>95.80</td>
<td><bold>95.86</bold></td>
<td>95.58</td>
</tr>
<tr>
<td>Lung2</td>
<td>93.41</td>
<td>93.76</td>
<td>91.71</td>
<td>95.12</td>
<td><bold>96.10</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>83.33</td>
<td>82.85</td>
<td>86.19</td>
<td>84.76</td>
<td><bold>90.48</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>28.55</td>
<td>30.51</td>
<td>54.66</td>
<td>27.29</td>
<td><bold>66.28</bold></td>
</tr>
<tr>
<td>Ovarian</td>
<td>91.37</td>
<td>92.54</td>
<td>98.03</td>
<td>94.71</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.5</td>
<td>3.6</td>
<td>2.7</td>
<td>2.4</td>
<td>1.8</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>3</td>
<td>2</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Classification accuracy obtained by NB</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>62.34</td>
<td>91.50</td>
<td>93.26</td>
<td><bold>93.66</bold></td>
<td>77.93</td>
</tr>
<tr>
<td>Spambase</td>
<td>81.61</td>
<td>88.03</td>
<td>89.76</td>
<td>89.51</td>
<td><bold>90.05</bold></td>
</tr>
<tr>
<td>Isolet</td>
<td>78.46</td>
<td>82.46</td>
<td>82.60</td>
<td><bold>83.07</bold></td>
<td>80.80</td>
</tr>
<tr>
<td>CNAE9</td>
<td>86.94</td>
<td>82.40</td>
<td><bold>89.39</bold></td>
<td>88.19</td>
<td>82.59</td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>24.37</td>
<td>30.18</td>
<td>90.33</td>
<td>39.94</td>
<td><bold>95.82</bold></td>
</tr>
<tr>
<td>Letter</td>
<td>64.12</td>
<td>66.02</td>
<td>66.10</td>
<td><bold>66.20</bold></td>
<td>66.13</td>
</tr>
<tr>
<td>Lung2</td>
<td>80.97</td>
<td>84.15</td>
<td>89.02</td>
<td>85.61</td>
<td><bold>96.10</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>64.76</td>
<td>60.00</td>
<td>60.95</td>
<td>68.57</td>
<td><bold>95.24</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>76.45</td>
<td>75.17</td>
<td><bold>78.51</bold></td>
<td>76.10</td>
<td>77.04</td>
</tr>
<tr>
<td>Ovarian</td>
<td>90.19</td>
<td>91.37</td>
<td>98.03</td>
<td>88.82</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.3</td>
<td>4.0</td>
<td>2.1</td>
<td>2.5</td>
<td>2.1</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>1</td>
<td>3</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="table-3">Table 3</xref>, the CTFS algorithm shows better performance on multiple test datasets compared to other stochastic approaches when KNN is used as a classifier. Specifically, the CTFS method achieves the best classification accuracy on 6 out of 10 datasets with a large sample size and high dimensions. Meanwhile, for the remaining 4 test datasets, the relative errors between the classification accuracy obtained by the CTFS method and the highest acquired by other comparison algorithms range from a maximum of 1.8% to a minimum of 0.03%, which is a relatively small gap. Therefore, it is concluded that the CTFS algorithm has a prominent advantage over other compared wrapper algorithms in terms of classification accuracy when handling test datasets with large sample sizes and high dimensions.</p>

<p>Different from the KNN classifier, the performance of the CTFS algorithm varies with the Bayesian classifier. <xref ref-type="table" rid="table-4">Table 4</xref> demonstrates that the classification accuracy of the CTFS algorithm is worse than that of other comparison algorithms on the test datasets with large sample sizes and medium dimensions. However, for test datasets with a moderate sample size and high dimensions, such as Lung2, Prostate1, etc., the CTFS algorithm achieves the best classification accuracy.</p>

<p>From <xref ref-type="table" rid="table-3">Tables 3</xref> and <xref ref-type="table" rid="table-4">4</xref>, as the number of features in the dataset grows, the CTFS algorithm gradually demonstrates its capability to efficiently discern the objective feature sets. When the number of features in the test datasets reaches 15,154, the classification accuracy of the CTFS method is 100%, while the classification accuracies of the other four compared algorithms are 91.37%, 92.54%, 98.03%, and 94.71%, respectively.</p>

<p>To evaluate the comprehensive classification performance of the proposed CTFS algorithm, we also analyze the recall rate and F1 score on various test datasets. <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> record the recall rate using KNN and Bayesian as classifiers, respectively. From <xref ref-type="table" rid="table-5">Table 5</xref>, the CTFS approach achieves a remarkable recall rate on 7 out of 10 datasets. Similarly, in <xref ref-type="table" rid="table-6">Table 6</xref>, the CTFS algorithm obtains the highest recall rate on 6 out of all datasets with an average recall rate of 84.5%, while the TMGWO algorithm has an average recall rate of 79.28%, which still holds the highest value among the four comparison algorithms.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Recall rate obtained by KNN</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>94.72</td>
<td>95.64</td>
<td>96.32</td>
<td><bold>97.11</bold></td>
<td>95.31</td>
</tr>
<tr>
<td>Spambase</td>
<td>88.72</td>
<td>89.58</td>
<td><bold>90.33</bold></td>
<td>89.30</td>
<td>89.13</td>
</tr>
<tr>
<td>Isolet</td>
<td>79.08</td>
<td>81.75</td>
<td>83.44</td>
<td>83.25</td>
<td><bold>85.83</bold></td>
</tr>
<tr>
<td>CNAE9</td>
<td>83.82</td>
<td>83.59</td>
<td>82.92</td>
<td>84.30</td>
<td><bold>88.77</bold></td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>96.49</td>
<td>97.06</td>
<td>97.15</td>
<td>97.99</td>
<td><bold>98.26</bold></td>
</tr>
<tr>
<td>Letter</td>
<td>95.24</td>
<td>95.29</td>
<td>95.78</td>
<td><bold>95.86</bold></td>
<td>95.56</td>
</tr>
<tr>
<td>Lung2</td>
<td>88.47</td>
<td>85.63</td>
<td>84.90</td>
<td>91.82</td>
<td><bold>92.00</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>83.90</td>
<td>83.66</td>
<td>86.28</td>
<td>85.13</td>
<td><bold>89.29</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>13.27</td>
<td>14.41</td>
<td>34.07</td>
<td>12.60</td>
<td><bold>47.96</bold></td>
</tr>
<tr>
<td>Ovarian</td>
<td>89.62</td>
<td>90.18</td>
<td>97.28</td>
<td>92.92</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.4</td>
<td>3.7</td>
<td>2.6</td>
<td>2.5</td>
<td>1.8</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>3</td>
<td>2</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Recall rate obtained by NB</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>60.60</td>
<td>91.43</td>
<td>93.15</td>
<td><bold>93.60</bold></td>
<td>76.28</td>
</tr>
<tr>
<td>Spambase</td>
<td>83.88</td>
<td>87.47</td>
<td>89.13</td>
<td>89.37</td>
<td><bold>89.69</bold></td>
</tr>
<tr>
<td>Isolet</td>
<td>79.01</td>
<td>82.42</td>
<td>82.65</td>
<td><bold>83.45</bold></td>
<td>80.54</td>
</tr>
<tr>
<td>CNAE9</td>
<td>86.53</td>
<td>82.09</td>
<td><bold>89.12</bold></td>
<td>88.24</td>
<td>83.40</td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>52.04</td>
<td>55.52</td>
<td>80.97</td>
<td>61.53</td>
<td><bold>96.05</bold></td>
</tr>
<tr>
<td>Letter</td>
<td>64.12</td>
<td>65.86</td>
<td><bold>66.04</bold></td>
<td>65.87</td>
<td>65.94</td>
</tr>
<tr>
<td>Lung2</td>
<td>52.55</td>
<td>62.99</td>
<td>71.03</td>
<td>64.55</td>
<td><bold>93.94</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>64.17</td>
<td>60.76</td>
<td>61.21</td>
<td>69.84</td>
<td><bold>93.93</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>56.77</td>
<td>56.63</td>
<td>60.97</td>
<td>56.74</td>
<td><bold>65.21</bold></td>
</tr>
<tr>
<td>Ovarian</td>
<td>89.79</td>
<td>91.21</td>
<td>98.57</td>
<td>88.71</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.3</td>
<td>4</td>
<td>2.1</td>
<td>2.6</td>
<td>2.0</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>2</td>
<td>3</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-7">Tables 7</xref> and <xref ref-type="table" rid="table-8">8</xref> illustrate the F1 score on ten datasets using KNN and Bayesian as classifiers. From <xref ref-type="table" rid="table-7">Table 7</xref>, the CTFS method achieves the highest F1 score on 6 out of 10 datasets. For the rest of the test datasets, the relative errors between the F1 score obtained by the CTFS method and the highest acquired by other comparison algorithms range from a maximum of 1.89% to a minimum of 0.14%, which is a relatively small difference. <xref ref-type="table" rid="table-8">Table 8</xref> shows the good performance of the CTFS algorithm using Bayesian as a classifier, which reaches the highest F1 score on 6 out of all test datasets with an average F1 score of 84.07%. From <xref ref-type="table" rid="table-5">Tables 5</xref>&#x2013;<xref ref-type="table" rid="table-8">8</xref>, it can be summarized that, in addition to the classification accuracy metric, the CTFS algorithm still maintains prominent performance advantages.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>F1 score obtained by KNN</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>94.77</td>
<td>95.60</td>
<td>96.31</td>
<td><bold>97.12</bold></td>
<td>95.23</td>
</tr>
<tr>
<td>Spambase</td>
<td>88.98</td>
<td>89.69</td>
<td><bold>90.30</bold></td>
<td>89.39</td>
<td>89.34</td>
</tr>
<tr>
<td>Isolet</td>
<td>78.43</td>
<td>81.00</td>
<td>82.43</td>
<td>82.38</td>
<td><bold>84.62</bold></td>
</tr>
<tr>
<td>CNAE9</td>
<td>83.10</td>
<td>83.36</td>
<td>82.19</td>
<td>83.93</td>
<td><bold>88.17</bold></td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>97.06</td>
<td>97.57</td>
<td>97.67</td>
<td><bold>97.94</bold></td>
<td>97.83</td>
</tr>
<tr>
<td>Letter</td>
<td>95.22</td>
<td>95.32</td>
<td>95.78</td>
<td><bold>95.85</bold></td>
<td>95.57</td>
</tr>
<tr>
<td>Lung2</td>
<td>90.31</td>
<td>87.37</td>
<td>85.43</td>
<td>90.97</td>
<td><bold>93.98</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>82.68</td>
<td>82.55</td>
<td>85.58</td>
<td>84.44</td>
<td><bold>89.20</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>6.19</td>
<td>9.03</td>
<td>34.37</td>
<td>5.44</td>
<td><bold>47.94</bold></td>
</tr>
<tr>
<td>Ovarian</td>
<td>90.43</td>
<td>91.33</td>
<td>97.63</td>
<td>94.01</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.5</td>
<td>3.6</td>
<td>2.6</td>
<td>2.4</td>
<td>1.9</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>3</td>
<td>2</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>F1 score obtained by NB</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Full</th>
<th>BHHO</th>
<th>TMGWO</th>
<th>CUS-SPSO</th>
<th>CTFS</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>54.53</td>
<td>91.46</td>
<td>93.22</td>
<td><bold>93.64</bold></td>
<td>76.32</td>
</tr>
<tr>
<td>Spambase</td>
<td>81.53</td>
<td>87.49</td>
<td>89.25</td>
<td>89.05</td>
<td><bold>89.65</bold></td>
</tr>
<tr>
<td>Isolet</td>
<td>78.38</td>
<td>81.95</td>
<td>82.40</td>
<td><bold>83.01</bold></td>
<td>80.07</td>
</tr>
<tr>
<td>CNAE9</td>
<td>86.21</td>
<td>81.51</td>
<td><bold>88.86</bold></td>
<td>87.88</td>
<td>82.87</td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>22.55</td>
<td>29.39</td>
<td>82.59</td>
<td>38.38</td>
<td><bold>93.46</bold></td>
</tr>
<tr>
<td>Letter</td>
<td>63.66</td>
<td>65.45</td>
<td><bold>65.73</bold></td>
<td>65.63</td>
<td>65.72</td>
</tr>
<tr>
<td>Lung2</td>
<td>55.71</td>
<td>65.12</td>
<td>70.93</td>
<td>67.49</td>
<td><bold>95.36</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>63.01</td>
<td>59.10</td>
<td>59.99</td>
<td>67.94</td>
<td><bold>94.55</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>54.87</td>
<td>54.66</td>
<td>58.16</td>
<td>55.20</td>
<td><bold>62.72</bold></td>
</tr>
<tr>
<td>Ovarian</td>
<td>88.83</td>
<td>90.53</td>
<td>97.76</td>
<td>88.15</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>AR</td>
<td>4.4</td>
<td>4</td>
<td>2.0</td>
<td>2.6</td>
<td>2.0</td>
</tr>
<tr>
<td>Rank</td>
<td>5</td>
<td>4</td>
<td>1</td>
<td>3</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition to the classification performance metrics, the number of selected features is also an important reference indicator. Generally, the lower the dimension of the candidate feature subset, the stronger its generalization performance and the better the dimensionality reduction effect of the algorithm. <xref ref-type="table" rid="table-9">Table 9</xref> shows the number of features selected by the five algorithms on each dataset. The CTFS algorithm achieves the minimum number of features when using KNN as a classifier on 8 out of all test datasets. Based on <xref ref-type="table" rid="table-3">Tables 3</xref>&#x2013;<xref ref-type="table" rid="table-8">8</xref>, it is concluded that the CTFS algorithm can achieve remarkable classification performance with a relatively small number of features.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>The average number of selected features achieved by five algorithms</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th align="center" colspan="2">Full</th>
<th align="center" colspan="2">BHHO</th>
<th align="center" colspan="2">TMGWO</th>
<th align="center" colspan="2">CUS-SPSO</th>
<th align="center" colspan="2">CTFS</th>
</tr>
<tr>
<th>Dataset</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>36.0</td>
<td>36.0</td>
<td>21.5</td>
<td>16.1</td>
<td>23.1</td>
<td><bold>8.1</bold></td>
<td>22.7</td>
<td>16.6</td>
<td><bold>11.9</bold></td>
<td>18.6</td>
</tr>
<tr>
<td>Spambase</td>
<td>57.0</td>
<td>57.0</td>
<td>30.5</td>
<td>27.8</td>
<td><bold>18.3</bold></td>
<td>34.0</td>
<td>43.9</td>
<td>36.6</td>
<td>50.6</td>
<td><bold>22.8</bold></td>
</tr>
<tr>
<td>Isolet</td>
<td>617.0</td>
<td>617.0</td>
<td>303.5</td>
<td>311.0</td>
<td>212.8</td>
<td><bold>168.6</bold></td>
<td>424.0</td>
<td>437.3</td>
<td><bold>191.0</bold></td>
<td>519.3</td>
</tr>
<tr>
<td>CNAE9</td>
<td>856.0</td>
<td>856.0</td>
<td>451.1</td>
<td><bold>476.4</bold></td>
<td>823.7</td>
<td>856.0</td>
<td>696.4</td>
<td>737.6</td>
<td><bold>342.8</bold></td>
<td>556.3</td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>241.0</td>
<td>241.0</td>
<td>122.5</td>
<td>120.0</td>
<td><bold>53.3</bold></td>
<td><bold>21.8</bold></td>
<td>181.1</td>
<td>152.3</td>
<td>104.7</td>
<td>22.3</td>
</tr>
<tr>
<td>Letter</td>
<td>16.0</td>
<td>16.0</td>
<td>12.7</td>
<td>11.7</td>
<td><bold>11.0</bold></td>
<td><bold>11.0</bold></td>
<td><bold>11.0</bold></td>
<td>11.6</td>
<td><bold>11.0</bold></td>
<td><bold>11.0</bold></td>
</tr>
<tr>
<td>Lung2</td>
<td>3312.0</td>
<td>3312.0</td>
<td>1624.0</td>
<td>1620.4</td>
<td>2687.6</td>
<td>139.7</td>
<td>2306.7</td>
<td>2195.9</td>
<td><bold>348.7</bold></td>
<td><bold>126.6</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>5966.0</td>
<td>5966.0</td>
<td>3041.9</td>
<td>3013.8</td>
<td>5519.1</td>
<td>4947.8</td>
<td>3666.5</td>
<td>3988.0</td>
<td><bold>17.1</bold></td>
<td><bold>3.5</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>8266.0</td>
<td>8266.0</td>
<td>4063.0</td>
<td><bold>4169.8</bold></td>
<td>182.5</td>
<td>6420.0</td>
<td>4892.0</td>
<td>6085.9</td>
<td><bold>55.4</bold></td>
<td>7258.5</td>
</tr>
<tr>
<td>Ovarian</td>
<td>15154.0</td>
<td>15154.0</td>
<td>7584.5</td>
<td>7621.1</td>
<td>632.4</td>
<td>117.0</td>
<td>7622.6</td>
<td>8381.3</td>
<td><bold>88.3</bold></td>
<td><bold>50.2</bold></td>
</tr>
<tr>
<td>Average</td>
<td>3452.1</td>
<td>3452.1</td>
<td>1725.5</td>
<td>1738.8</td>
<td>1016.4</td>
<td>1272.4</td>
<td>1986.7</td>
<td>2204.3</td>
<td><bold>122.2</bold></td>
<td><bold>858.9</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the comparison of the number of selected features for four algorithms on ten test datasets. Specifically, we define <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>N</mml:mi><mml:mi>u</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to normalize the number of features selected by each algorithm, where <italic>i</italic> denotes the number of test datasets, and <italic>j</italic> is the sequence of the algorithms in the table. <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>N</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the number of features selected by the <italic>j</italic>th algorithm on the <italic>i</italic>th dataset. <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>N</mml:mi><mml:mi>u</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the number of the initial features on the <italic>i</italic>th dataset. In the spider web diagram, the red lines denote the CTFS algorithm. The red lines are distributed mostly inside the entire spider web, indicating that the CTFS algorithm selects the fewest features. The red lines are primarily at the center of the spider web on Prostate1 and Gutenberg when Bayesian is used as a classifier, which shows the extremely marked dimensionality reduction effect of the CTFS algorithm on these two datasets.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Number of selected features of different methods on ten datasets (a) KNN classifier (b) Bayesian classifier</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-4.tif"/>
</fig>
<p>Besides the indicators mentioned above to verify the efficiency of each method, running time is also another metric that cannot be ignored. The purpose of feature selection is to select key information attributes with less time and cost to achieve higher classification performance. <xref ref-type="table" rid="table-10">Table 10</xref> lists the average running time consumed by these algorithms. Compared with other comparison algorithms, the CTFS algorithm requires the least running time on most datasets.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Running time consumed by the four algorithms on the ten datasets (unit: s)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th align="center" colspan="2">BHHO</th>
<th align="center" colspan="2">TMGWO</th>
<th align="center" colspan="2">CUS-SPSO</th>
<th align="center" colspan="2">CTFS</th>
</tr>
<tr>
<th>Dataset</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
<th>KNN</th>
<th>NB</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chess</td>
<td>286.9</td>
<td>45.3</td>
<td>183.4</td>
<td>39.6</td>
<td>531.2</td>
<td>28.8</td>
<td><bold>32.3</bold></td>
<td><bold>8.5</bold></td>
</tr>
<tr>
<td>Spambase</td>
<td>153.4</td>
<td>63.9</td>
<td>334.9</td>
<td>70.6</td>
<td>96.4</td>
<td>42.8</td>
<td><bold>47.9</bold></td>
<td><bold>11.0</bold></td>
</tr>
<tr>
<td>Isolet</td>
<td>115.4</td>
<td>294.6</td>
<td>758.8</td>
<td>1640.8</td>
<td>83.4</td>
<td>220.6</td>
<td><bold>47.5</bold></td>
<td><bold>59.2</bold></td>
</tr>
<tr>
<td>CNAE9</td>
<td>108.3</td>
<td>159.3</td>
<td>998.2</td>
<td>1883.5</td>
<td>187.1</td>
<td>174.4</td>
<td><bold>55.4</bold></td>
<td><bold>42.2</bold></td>
</tr>
<tr>
<td>TUANDROMD</td>
<td>152.8</td>
<td>128.7</td>
<td>1259.3</td>
<td>312.2</td>
<td>971.8</td>
<td>108.1</td>
<td><bold>25.6</bold></td>
<td><bold>16.9</bold></td>
</tr>
<tr>
<td>Letter</td>
<td>5371.5</td>
<td>245.9</td>
<td>1512.9</td>
<td>147.0</td>
<td>4369.7</td>
<td>177.2</td>
<td><bold>139.7</bold></td>
<td><bold>21.1</bold></td>
</tr>
<tr>
<td>Lung2</td>
<td>83.8</td>
<td>119.5</td>
<td>3451.1</td>
<td>2479.2</td>
<td><bold>65.6</bold></td>
<td>93.7</td>
<td>1189.5</td>
<td><bold>76.7</bold></td>
</tr>
<tr>
<td>Prostate1</td>
<td>88.3</td>
<td>111.4</td>
<td>6523.2</td>
<td>7703.9</td>
<td><bold>72.2</bold></td>
<td>97.6</td>
<td>1888.3</td>
<td><bold>92.5</bold></td>
</tr>
<tr>
<td>Gutenberg</td>
<td>520.1</td>
<td>1068.9</td>
<td>16990.8</td>
<td>101827.6</td>
<td><bold>422.3</bold></td>
<td><bold>997.9</bold></td>
<td>8249.2</td>
<td>1258.9</td>
</tr>
<tr>
<td>Ovarian</td>
<td>377.1</td>
<td>743.3</td>
<td>38984.7</td>
<td>48328.9</td>
<td><bold>264.7</bold></td>
<td><bold>519.1</bold></td>
<td>26272.9</td>
<td>1200.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> intuitively displays the running time of algorithms on ten datasets when using KNN and Bayesian as classifiers. The length of the orange rectangle represents the proportional running time of the CTFS algorithm. Because of the long running time of the TMGWO algorithm on high-dimensional datasets, we convert the running time by using the proportional running time of a single algorithm in the total running time of all algorithms as the value of the length. The algorithm proposed in this paper has a shorter length on most datasets, especially using Bayesian as a classifier; that is, the running time is relatively short.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Proportional running time of different algorithms on ten datasets (a) KNN classifier (b) Bayesian classifier</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57103-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions</title>
<p>This paper introduces a novel wrapper feature selection algorithm called CTFS. Initially, the modified sparse autoencoder network is integrated with mutual information to achieve dominance over each feature. Following that, the informative attributes can be located in the optimal feature subset by a designed contribution tracking strategy, which performs in a dominance accumulation form. The computational results on multiple test datasets illustrate that the CTFS algorithm outperforms advanced algorithms in terms of classification accuracy, recall rate, F1 score, and dimension reduction capability. However, compared to other algorithms, the experimental results of runtime for high-dimensional datasets are not significantly superior. Subsequent research endeavors will focus on further reducing the training time in addressing ultra-high-dimensional feature selection problems.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their gratitude to the editor and the anonymous reviewers for their insightful suggestions, which significantly raised the caliber of this work.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported in part by the National Key Research and Development Program of China under Grant (No. 2021YFB3300900); the NSFC Key Supported Project of the Major Research Plan under Grant (No. 92267206); the National Natural Science Foundation of China under Grant (Nos. 72201052, 62032013, 62173076); the Fundamental Research Funds for the Central Universities under Grant (No. N2204017); the Fundamental Research Funds for State Key Laboratory of Synthetical Automation for Process Industries under Grant (No. 2013ZCX11).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Study conception, design and draft manuscript preparation: Yifan Yu; data collection and editing: Dazhi Wang; analysis and interpretation of results: Yanhua Chen; validation: Hongfeng Wang; project administration and supervision: Min Huang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data sets applied in the paper are available in the UCI repository at <ext-link ext-link-type="uri" xlink:href="https://archive.ics.uci.edu/datasets">https://archive.ics.uci.edu/datasets</ext-link> (accessed on 05 April 2024) and the scikit-feature feature selection repository at <ext-link ext-link-type="uri" xlink:href="https://jundongl.github.io/scikit-feature/datasets">https://jundongl.github.io/scikit-feature/datasets</ext-link> (accessed on 11 April 2024).</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Gui</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ji</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Tao</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<article-title>Feature selection based on structured sparsity: A comprehensive study</article-title>,&#x201D; <source>IEEE Trans. Neural Netw. Learn. Syst.</source>, vol. <volume>28</volume>, no. <issue>7</issue>, pp. <fpage>1490</fpage>&#x2013;<lpage>1507</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1109/TNNLS.2016.2551724</pub-id>; <pub-id pub-id-type="pmid">28287983</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. S.</given-names> <surname>Gandhi</surname></string-name> and <string-name><given-names>S. S.</given-names> <surname>Prabhune</surname></string-name></person-group>, &#x201C;<article-title>Overview of feature subset selection algorithm for high dimensional data</article-title>,&#x201D; in <conf-name>2017 Proc. Int. Conf. Inventive Syst. Control (ICISC)</conf-name>, <publisher-loc>Coimbatore, India</publisher-loc>, <year>2017</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Fast attribute reduction via inconsistent equivalence classes for large-scale data</article-title>,&#x201D; <source>Int. J. Approximate Reasoning</source>, vol. <volume>163</volume>, <year>2023</year>, Art. no. 109039. doi: <pub-id pub-id-type="doi">10.1016/j.ijar.2023.109039</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. A.</given-names> <surname>Alhussan</surname></string-name> and <string-name><given-names>S. K.</given-names> <surname>Towfek</surname></string-name></person-group>, &#x201C;<article-title>5G resource allocation using feature selection and greylag goose optimization algorithm</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>80</volume>, no. <issue>1</issue>, pp. <fpage>1180</fpage>&#x2013;<lpage>1181</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2024.049874</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yi</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yuan</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>A heuristic radiomics feature selection method based on frequency iteration and multi-supervised training mode</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>79</volume>, no. <issue>2</issue>, pp. <fpage>2278</fpage>&#x2013;<lpage>2279</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2024.047989</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Braik</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hammouri</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Alzoubi</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Sheta</surname></string-name></person-group>, &#x201C;<article-title>Feature selection based nature inspired capuchin search algorithm for solving classification problems</article-title>,&#x201D; <source>Expert Syst. Appl.</source>, vol. <volume>235</volume>, <year>2024</year>, Art. no. 121128. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2023.121128</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>J. M.</given-names> <surname>Zurada</surname></string-name>, and <string-name><given-names>N. R.</given-names> <surname>Pal</surname></string-name></person-group>, &#x201C;<article-title>Feature selection for neural networks using group lasso regularization</article-title>,&#x201D; <source>IEEE Trans. Knowl. Data Eng.</source>, vol. <volume>32</volume>, no. <issue>4</issue>, pp. <fpage>659</fpage>&#x2013;<lpage>673</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TKDE.2019.2893266</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Manifold regularized stacked denoising autoencoders with feature selection</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>358</volume>, pp. <fpage>235</fpage>&#x2013;<lpage>245</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2019.05.050</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. J.</given-names> <surname>Urbanowicz</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Meeker</surname></string-name>, <string-name><given-names>W. La</given-names> <surname>Cava</surname></string-name>, <string-name><given-names>R. S.</given-names> <surname>Olson</surname></string-name>, and <string-name><given-names>J. H.</given-names> <surname>Moore</surname></string-name></person-group>, &#x201C;<article-title>Relief-based feature selection: Introduction and review</article-title>,&#x201D; <source>J. Biomed. Inform.</source>, vol. <volume>85</volume>, pp. <fpage>189</fpage>&#x2013;<lpage>203</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jbi.2018.07.014</pub-id>; <pub-id pub-id-type="pmid">30031057</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Long</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Ding</surname></string-name></person-group>, &#x201C;<article-title>Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>27</volume>, no. <issue>8</issue>, pp. <fpage>1226</fpage>&#x2013;<lpage>1238</lpage>, <year>2005</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2005.159</pub-id>; <pub-id pub-id-type="pmid">16119262</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Roffo</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Melzi</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Castellani</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Vinciarelli</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Cristani</surname></string-name></person-group>, &#x201C;<article-title>Infinite feature selection: A graph-based feature filtering approach</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>43</volume>, no. <issue>12</issue>, pp. <fpage>4396</fpage>&#x2013;<lpage>4410</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2020.3002843</pub-id>; <pub-id pub-id-type="pmid">32750789</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Hancer</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xue</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Differential evolution for filter feature selection based on information theory and feature ranking</article-title>,&#x201D; <source>Knowl.-Based Syst.</source>, vol. <volume>140</volume>, no. <issue>10</issue>, pp. <fpage>103</fpage>&#x2013;<lpage>119</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2017.10.028</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Yan</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Feature selection and analysis on correlated gas sensor data with recursive feature elimination</article-title>,&#x201D; <source>Sens. Actuators B: Chem.</source>, vol. <volume>212</volume>, pp. <fpage>353</fpage>&#x2013;<lpage>363</lpage>, <year>2015</year>. doi: <pub-id pub-id-type="doi">10.1016/j.snb.2015.02.025</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Recursive elimination-election algorithms for wrapper feature selection</article-title>,&#x201D; <source>Appl. Soft Comput.</source>, vol. <volume>113</volume>, <year>2021</year>, Art. no. 107956. doi: <pub-id pub-id-type="doi">10.1016/j.asoc.2021.107956</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. B.</given-names> <surname>Kursa</surname></string-name> and <string-name><given-names>W. R.</given-names> <surname>Rudnicki</surname></string-name></person-group>, &#x201C;<article-title>Feature selection with the Boruta package</article-title>,&#x201D; <source>J. Stat. Softw.</source>, vol. <volume>36</volume>, no. <issue>11</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.18637/jss.v036.i11</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. K.</given-names> <surname>Naik</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Kuppili</surname></string-name></person-group>, &#x201C;<article-title>An embedded feature selection method based on generalized classifier neural network for cancer classification</article-title>,&#x201D; <source>Comput. Biol. Med.</source>, vol. <volume>168</volume>, <year>2024</year>, Art. no. 107677. doi: <pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.107677</pub-id>; <pub-id pub-id-type="pmid">37988786</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Jim&#x00E9;nez-Cordero</surname></string-name>, <string-name><given-names>J. M.</given-names> <surname>Morales</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Pineda</surname></string-name></person-group>, &#x201C;<article-title>A novel embedded min-max approach for feature selection in nonlinear support vector machine classification</article-title>,&#x201D; <source>Eur. J. Oper. Res.</source>, vol. <volume>293</volume>, no. <issue>1</issue>, pp. <fpage>24</fpage>&#x2013;<lpage>35</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ejor.2020.12.009</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Braik</surname></string-name></person-group>, &#x201C;<article-title>Enhanced Ali Baba and the forty thieves algorithm for feature selection</article-title>,&#x201D; <source>Neural Comput. Appl.</source>, vol. <volume>35</volume>, no. <issue>8</issue>, pp. <fpage>6153</fpage>&#x2013;<lpage>6184</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1007/s00521-022-08015-5</pub-id>; <pub-id pub-id-type="pmid">36408290</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Sun</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Variable-size cooperative coevolutionary particle swarm optimization for feature selection on high-dimensional data</article-title>,&#x201D; <source>IEEE Trans. Evol. Comput.</source>, vol. <volume>24</volume>, no. <issue>5</issue>, pp. <fpage>882</fpage>&#x2013;<lpage>895</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TEVC.2020.2968743</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Qu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>He</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Xie</surname></string-name></person-group>, &#x201C;<article-title>Explicit and size-adaptive PSO-based feature selection for classification</article-title>,&#x201D; <source>Swarm Evol. Comput.</source>, vol. <volume>77</volume>, <year>2023</year>, Art. no. 101249. doi: <pub-id pub-id-type="doi">10.1016/j.swevo.2023.101249</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Xiong</surname></string-name></person-group>, &#x201C;<article-title>A high-dimensional feature selection method based on modified gray wolf optimization</article-title>,&#x201D; <source>Appl. Soft Comput.</source>, vol. <volume>135</volume>, <year>2023</year>, Art. no. 110031. doi: <pub-id pub-id-type="doi">10.1016/j.asoc.2023.110031</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. G.</given-names> <surname>Mohammadi</surname></string-name> and <string-name><given-names>M. S.</given-names> <surname>Abadeh</surname></string-name></person-group>, &#x201C;<article-title>Image steganalysis using a bee colony based feature selection algorithm</article-title>,&#x201D; <source>Eng. Appl. Artif. Intell.</source>, vol. <volume>31</volume>, pp. <fpage>35</fpage>&#x2013;<lpage>43</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1016/j.engappai.2013.09.016</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xue</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>Correlation-guided updating strategy for feature selection in classification with surrogate-assisted particle swarm optimization</article-title>,&#x201D; <source>IEEE Trans. Evol. Comput.</source>, vol. <volume>26</volume>, no. <issue>5</issue>, pp. <fpage>1015</fpage>&#x2013;<lpage>1029</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TEVC.2021.3134804</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shangguan</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wu</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>MEL: Efficient multi-task evolutionary learning for high-dimensional feature selection</article-title>,&#x201D; <source>IEEE Trans. Knowl. Data Eng.</source>, vol. <volume>36</volume>, no. <issue>8</issue>, pp. <fpage>4020</fpage>&#x2013;<lpage>4033</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1109/TKDE.2024.3366333</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xue</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>An evolutionary multitasking-based feature selection method for high-dimensional classification</article-title>,&#x201D; <source>IEEE Trans. Cybern.</source>, vol. <volume>52</volume>, no. <issue>7</issue>, pp. <fpage>7172</fpage>&#x2013;<lpage>7186</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TCYB.2020.3042243</pub-id>; <pub-id pub-id-type="pmid">33382668</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Heidari</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Hierarchical Harris hawks optimizer for feature selection</article-title>,&#x201D; <source>J. Adv. Res.</source>, vol. <volume>53</volume>, pp. <fpage>261</fpage>&#x2013;<lpage>278</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jare.2023.01.014</pub-id>; <pub-id pub-id-type="pmid">36690206</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Lahmar</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Zaier</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Yahia</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Boaullegue</surname></string-name></person-group>, &#x201C;<article-title>A novel improved binary harris hawks optimization for high dimensionality feature selection</article-title>,&#x201D; <source>Pattern Recognit. Lett.</source>, vol. <volume>171</volume>, no. <issue>6</issue>, pp. <fpage>170</fpage>&#x2013;<lpage>176</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.patrec.2023.05.007</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Abdel-Basset</surname></string-name>, <string-name><given-names>D.</given-names> <surname>El-Shahat</surname></string-name>, <string-name><given-names>I.</given-names> <surname>El-Henawy</surname></string-name>, <string-name><given-names>V. H. C. De</given-names> <surname>Albuquerque</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Mirjalili</surname></string-name></person-group>, &#x201C;<article-title>A new fusion of grey wolf optimizer algorithm with a two-phase mutation for feature selection</article-title>,&#x201D; <source>Expert Syst. Appl.</source>, vol. <volume>139</volume>, no. <issue>3</issue>, pp. <fpage>432</fpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2019.112824</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhou</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Hua</surname></string-name></person-group>, &#x201C;<article-title>A correlation guided genetic algorithm and its application to feature selection</article-title>,&#x201D; <source>Appl. Soft Comput.</source>, vol. <volume>123</volume>, <year>2022</year>, Art. no. 108964. doi: <pub-id pub-id-type="doi">10.1016/j.asoc.2022.108964</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Pool</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tran</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Dally</surname></string-name></person-group>, &#x201C;<article-title>Learning both weights and connections for efficient neural networks</article-title>,&#x201D; in <conf-name>2015 Adv. Neural Inf. Proces. Syst. (NIPS)</conf-name>, <publisher-loc>Montreal, QC, Canada</publisher-loc>, <year>2015</year>, pp. <fpage>1135</fpage>&#x2013;<lpage>1143</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. C.</given-names> <surname>Mocanu</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Mocanu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Stone</surname></string-name>, <string-name><given-names>P. H.</given-names> <surname>Nguyen</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Gibescu</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Liotta</surname></string-name></person-group>, &#x201C;<article-title>Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science</article-title>,&#x201D; <source>Nat. Commun.</source>, vol. <volume>9</volume>, no. <issue>1</issue>, <year>2018</year>, Art. no. 2383. doi: <pub-id pub-id-type="doi">10.1038/s41467-018-04316-3</pub-id>; <pub-id pub-id-type="pmid">29921910</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Atashgahi</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Quick and robust feature selection: The strength of energy-efficient sparse training for autoencoders</article-title>,&#x201D; <source>Mach. Learn.</source>, vol. <volume>111</volume>, no. <issue>1</issue>, pp. <fpage>377</fpage>&#x2013;<lpage>414</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s10994-021-06063-x</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Sokar</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Atashgahi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Pechenizkiy</surname></string-name>, and <string-name><given-names>D. C.</given-names> <surname>Mocanu</surname></string-name></person-group>, &#x201C;<article-title>Where to pay attention in sparse training for feature selection?</article-title>&#x201D; in <conf-name>2022 Adv. Neural Inf. Process. Syst (NIPS)</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <year>2022</year>, pp. <fpage>1627</fpage>&#x2013;<lpage>1642</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Feature selection: A data perspective</article-title>,&#x201D; <source>ACM Comput. Surv.</source>, vol. <volume>50</volume>, no. <issue>6</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>45</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1145/3136625</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>