<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">33603</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.033603</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Feature-Limited Prediction on the UCI Heart Disease Dataset</article-title>
<alt-title alt-title-type="left-running-head">Feature-Limited Prediction on the UCI Heart Disease Dataset</alt-title>
<alt-title alt-title-type="right-running-head">Feature-Limited Prediction on the UCI Heart Disease Dataset</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Alfadli</surname><given-names>Khadijah Mohammad</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Almagrabi</surname><given-names>Alaa Omran</given-names></name><email>aalmagrabi3@kau.edu.sa</email></contrib>
<aff id="aff-1"><institution>Department of Information Systems, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Alaa Omran Almagrabi. Email: <email>aalmagrabi3@kau.edu.sa</email></corresp>
</author-notes>
<pub-date publication-format="print" date-type="pub" iso-8601-date="2022-12-15"><day>15</day>
<month>12</month>
<year>2022</year></pub-date>
<volume>74</volume>
<issue>3</issue>
<fpage>5871</fpage>
<lpage>5883</lpage>
<history>
<date date-type="received"><day>22</day><month>6</month><year>2022</year></date>
<date date-type="accepted"><day>11</day><month>10</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Alfadli and Almagrabi</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Alfadli and Almagrabi</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_33603.pdf"></self-uri>
<abstract>
<p>Heart diseases are the undisputed leading causes of death globally. Unfortunately, the conventional approach of relying solely on the patient&#x2019;s medical history is not enough to reliably diagnose heart issues. Several potentially indicative factors exist, such as abnormal pulse rate, high blood pressure, diabetes, high cholesterol, etc. Manually analyzing these health signals&#x2019; interactions is challenging and requires years of medical training and experience. Therefore, this work aims to harness machine learning techniques that have proved helpful for data-driven applications in the rise of the artificial intelligence era. More specifically, this paper builds a hybrid model as a tool for data mining algorithms like feature selection. The goal is to determine the most critical factors that play a role in discriminating patients with heart illnesses from healthy individuals. The contribution in this field is to provide the patients with accurate and timely tentative results to help prevent further complications and heart attacks using minimum information. The developed model achieves 84.24&#x0025; accuracy, 89.22&#x0025; Recall, and 83.49&#x0025; Precision using only a subset of the features.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Machine learning</kwd>
<kwd>feature selection</kwd>
<kwd>heart disease</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>According to the World Health Organization (WHO), cardiovascular diseases (CVDs), commonly known as heart diseases, are the leading causes of death globally. In 2016, the total death count reached about 58 million people, 31&#x0025; of whom died due to CVDs. Most of these deaths, around 85&#x0025;, were from heart attacks and strokes [<xref ref-type="bibr" rid="ref-1">1</xref>]. WHO has put a worldwide action plan spanning 2013 to 2020 in response to CVDs and cancer, diabetes, and chronic respiratory diseases, collectively known as Noncommunicable diseases (NCDs). The goal is to attain a 25&#x0025; relative reduction in premature death from NCDs by 2025. These efforts are necessary steps toward fighting this on a global scale. However, humanity needs awareness on a more individual level. For example, the American Health Association has reported several behavioral risk factors that can be regulated to prevent CVDs, such as smoking cigarettes, eating unhealthy food, and not exercising regularly. High cardiovascular risk patients suffering from hypertension, diabetes, and/or hyperlipidemia should be closely monitored as early detection of CVDs can prevent premature deaths [<xref ref-type="bibr" rid="ref-2">2</xref>]. Many of these risk factors can be easily measured using accessible tools that might be part of any modern household. Moreover, with the advancement of technology, there are even now smartwatches and wearables equipped with health-tracking sensors. Every factor alone might not be a good indicator of heart disease, but their interaction can provide a clearer signal to the health counselor or the doctor [<xref ref-type="bibr" rid="ref-3">3</xref>]. Developing systems that can assist human professionals in monitoring high-risk patients is a good strategy for performing widespread testing for CVDs and devising proactive measures [<xref ref-type="bibr" rid="ref-4">4</xref>]. To enable such application without losing utility, it is imperative to use as less information about the patient as possible.</p>
<p>With the rise of the Artificial Intelligence (AI) era, many data-driven problems have become possible to solve with expert accuracy. Most of the recent success can be attributed to the advances in Machine Learning (ML), a subfield of AI that relies heavily on abundant data. To that end, using datasets that contain patients&#x2019; information with and without CVDs, such as the UCI Heart Disease Dataset [<xref ref-type="bibr" rid="ref-5">5</xref>], is essential to applying ML algorithms. However, analyzing these datasets requires cleaning and preprocessing. This work proposes different approaches to classify early whether a patient has heart disease using classical and modern ML methods with the help of some Data Mining (DM) techniques. Finally, it will perform ablation studies to determine the most distinctive features of CVDs.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>To reliably diagnose heart diseases in a patient, a doctor needs to ask some questions and run a few tests. The goal is to identify important attributes as the basis for the final diagnosis. Examples include the patient&#x2019;s age, sex, type of chest pain, and resting blood pressure. In the ML community, these attributes are referred to as features. One can formulate the problem as an ML problem (precisely, a classification problem); given the input features (i.e., patient information), the goal is to predict whether the patient has cardiovascular disease (CVD) or not. The proposed system attempts to solve this problem to prevent further complications that might lead to heart failures like heart attacks and strokes [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<sec id="s2_1"><label>2.1</label><title>Algorithms</title>
<p>From the DM field [<xref ref-type="bibr" rid="ref-6">6</xref>], it is known that some features are more important than others for classification. However, sometimes the combination of two weak features can be more critical than a stronger feature. All of this led to the study of feature selection methods [<xref ref-type="bibr" rid="ref-7">7</xref>]. Examples of such methods include the Relief method, Minimal-Redundancy Maximal-Relevance Algorithm, Least Absolute Shrinkage, and the Selection Operator, all of which were studied for CVDs in [<xref ref-type="bibr" rid="ref-8">8</xref>]. This research will leverage feature selection to its advantage for two main reasons. The first reason is to improve the predictive power of the proposed classifier. The second reason is to train multiple models that rely on less information which helps when specific values are hard to attain (e.g., blood pressure is unknown).</p>
<p>ML classifiers can be divided into classical and modern [<xref ref-type="bibr" rid="ref-9">9</xref>]. Heart disease prediction systems were developed using both methods. Examples of classical methods include K-Nearest Neighbor (KNN), Support Vector Machine (SVM) [<xref ref-type="bibr" rid="ref-10">10</xref>], and Naive Bayes (NB) [<xref ref-type="bibr" rid="ref-11">11</xref>], all of which were studied in [<xref ref-type="bibr" rid="ref-12">12</xref>]. Other classical approaches include Logistic Regression (LR) [<xref ref-type="bibr" rid="ref-13">13</xref>], Ridge Classifier (RC) [<xref ref-type="bibr" rid="ref-14">14</xref>], Linear Discriminant Analysis (LDA) [<xref ref-type="bibr" rid="ref-15">15</xref>], Gaussian Process (GP) [<xref ref-type="bibr" rid="ref-16">16</xref>], Decision Tree (DT) [<xref ref-type="bibr" rid="ref-17">17</xref>], and Random Forest (RF) [<xref ref-type="bibr" rid="ref-18">18</xref>]. Modern ML methods focus on Deep Learning (DL), the study of deep Artificial Neural Networks (ANN). Examples of ANNs include Multi-Layer Perceptron (MLP) [<xref ref-type="bibr" rid="ref-19">19</xref>] and Recurrent Neural Network (RNN) [<xref ref-type="bibr" rid="ref-20">20</xref>]. This paper will compare a few classical and modern ML methods and build a hybrid model combining multiple models, also known as the ensemble model, as in [<xref ref-type="bibr" rid="ref-21">21</xref>]. Ensembles are better since two minds are always better than one (e.g., the wisdom of the crowd). The biggest hurdle to this work is the availability of data. Since health records are considered private information, coming across useful data for research is not as easy as in other fields. Up to our knowledge, the only publicly available dataset for CVD was collected three decades ago [<xref ref-type="bibr" rid="ref-5">5</xref>]. Other datasets exist, but they require signing NDAs because of their sensitive nature. Hence, most cited work use only this dataset [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>Contributions</title>
<p>The main contributions can be summarized as follows: (1) Providing exploratory data analysis on the UCI Heart Disease Dataset to study its features. (2) Following proper ML workflow to train on the entire dataset without removing patients with missing values. (3) Determining the most discriminative features of CVDs using feature selection on an ensemble model. (4) Performing a comparative study of multiple ML models and releasing a competitive model using only a few selected features. (5) Open-sourcing reproducible code for all the experiments in the supplementary material. Contemporary arts exist, such as [<xref ref-type="bibr" rid="ref-19">19</xref>] and [<xref ref-type="bibr" rid="ref-23">23</xref>]. Nevertheless, they do not train on the entire dataset and do not perform feature selection.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>UCI Heart Disease Dataset</title>
<p>This dataset was collected in 1988 from four cities: Cleveland, Hungary, Switzerland, and Long Beach. It has 920 cases of people with and without CVDs with 76 attributes each. However, only 13 are used in practice, as seen in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The top five plots illustrate the histograms for the numerical features in the dataset. The count of patients missing the value for a particular feature is presented in the legend. The bottom plots show the categorical features in pie charts (missing values are labeled as &#x201C;Unknown&#x201D;).</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>UCI heart disease dataset features&#x2019; distributions</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-1.png"/></fig>
<p>After the Extract-Transform-Load (ETL) step comes performing Exploratory Data Analysis (EDA). The first thing to note here is that two-thirds of the cases have missing values. Removing them as commonly practiced is unadvisable since the dataset is already too small. In addition, the data shows five different CVDs severity levels ranging from healthy to Severe. These class labels are imbalanced, but it is possible to balance them out by changing the problem into binary classifications (two class labels: healthy and unhealthy). From this point onward, this assumption will be held to simplify the analysis. Lastly, it is essential to mention that about 80&#x0025; of the patients are males. This is unlikely a truly representative sample of the real world, which might indicate a bias in the dataset. It is paramount to keep this in mind as it might have a detrimental effect on the predictive power and reliability of the trained models.</p>
<p>Nevertheless, one needs to see how the data is distributed given the class label to get a deeper insight into interpreting these features. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> plots the numerical features against each other in pairs while color coding the points by whether the patients suffer from CVDs or not (missing values are ignored). The plots on the diagonal are simply the histograms of the features since the scatter plot of any feature will result in a degenerate line. It can be observed that the most discriminative features are &#x201C;Max Heart Rate&#x201D; and &#x201C;ST Depression Peak&#x201D;. In addition, there is no strong correlation between the features, which means that they encode different information and are not replaceable (no multicollinearity).</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Numerical features&#x2019; correlations for patients with and without CVDs</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-2.png"/></fig>
<p>The same analysis can be applied to the categorical features, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. From these plots, it can be observed that the ratio of healthy to sick people is almost five times more in males than females. This could be true globally, but one cannot be confident of this since the data is not statistically significant. Furthermore, most patients with heart disease appear to have no chest pain, &#x201C;Asymptomatic&#x201D;, which shows the importance of this research. Cases like this can easily go unnoticed and undiagnosed.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Categorical features&#x2019; statistics for patients with and without CVDs</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-3a.png"/><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-3b.png"/></fig>
</sec>
<sec id="s4"><label>4</label><title>Methodology</title>
<p>The proposed workflow is outlined in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> and explained in more detail in this section.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>The proposed system pipeline</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-4.png"/></fig>
<sec id="s4_1"><label>4.1</label><title>Data Preprocessing</title>
<sec id="s4_1_1"><label>4.1.1</label><title>Numerical Features</title>
<p>Since every feature has different ranges, like age and heart rate, this work will apply normalization to them. Normalization can significantly impact the trained model as it avoids preempting it to think that heart rate is more important than the person&#x2019;s age. If such a relationship exists, the model should learn it on its own. To that end, the experiments will normalize the features to be zero-centered with unit variance. This is done by taking the mean <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> and dividing by the standard deviation <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mfrac></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_1_2"><label>4.1.2</label><title>Categorical Features</title>
<p>Most ML models work strictly with numerical features. It is possible to convert categorical features into numerical features. The trick is to use one-hot encoding, an all-zeros vector with a single element being one corresponding to the index of the category.</p>
</sec>
<sec id="s4_1_3"><label>4.1.3</label><title>Missing Values</title>
<p>This work replaces any missing value with the mean if it was numerical or a new class label &#x201C;Unknown&#x201D; if it was categorical, and the target classes are balanced by repeating randomly selected cases.</p>
</sec>
<sec id="s4_1_4"><label>4.1.4</label><title>Data Splits</title>
<p>Training a complex model on simple data can result in overfitting. Informally, it is when the model can memorize the dataset entirely without learning how to classify it correctly. It is the model&#x2019;s inability to detect the underlying patterns in the data. Whereas training a simple model on complex data might result in underfitting (learning trivial rules). For example, a model classifies patients based on age only (sick if old and healthy otherwise). To avoid both problematic outcomes, the data is split into two chunks. The first split will be used to train the model, and the second to test it. Both splits should be representative enough of the entire dataset (the same ratios of healthy to sick cases; stratified). A trained model is overfitting if its performance in training surpasses the testing and underfitting if it could not improve over a fixed classifier; it always predicts the same thing (healthy or sick) regardless of the input.</p>
</sec>
</sec>
<sec id="s4_2"><label>4.2</label><title>Training</title>
<sec id="s4_2_1"><label>4.2.1</label><title>Hyperparameter Tuning</title>
<p>Each ML model has a few configurable hyperparameters, a set of properties that changes its training behavior and final performance. Usually, they depend on each other (e.g., a particular hyperparameter setting has a different meaning and effect if the value of another hyperparameter is changed). So, to achieve the best results for a model, it needs to be trained under all combinations of possible assigned values for its hyperparameters if feasible. This is what is known as hyperparameter tuning through grid-search. <xref ref-type="table" rid="table-1">Table 1</xref> lists the models and their grid-search values. However, it is possible to accidentally face overfitting on the test set during hyperparameter tuning (hold-out set leakage). Therefore, it is a widespread practice to tune on a small chunk of the training set, usually called the validation set.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>ML models and their grid-search values for hyperparameter tuning</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Hyperparameter</th>
<th align="left">Grid-Search Values</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">KNN</td>
<td align="left">n_neighbors</td>
<td align="left">1, 2, 3, 4</td>
</tr>
<tr>
<td/>
<td align="left">weights</td>
<td align="left">&#x2018;uniform&#x2019;, &#x2018;distance&#x2019;</td>
</tr>
<tr>
<td/>
<td align="left">p</td>
<td align="left">1, 2, 3</td>
</tr>
<tr>
<td align="left">LR</td>
<td align="left">penalty</td>
<td align="left">&#x2018;l1&#x2019;, &#x2018;l2&#x2019;</td>
</tr>
<tr>
<td/>
<td align="left">C</td>
<td align="left">100, 10, 1, 0.1, 0.01, 0.001</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">kernel</td>
<td align="left">&#x2018;linear&#x2019;, &#x2018;rbf&#x2019;</td>
</tr>
<tr>
<td align="left">NB</td>
<td align="left">alpha</td>
<td align="left">1, 0.1, 0.01, 0.001, 0.0001, 0.00001</td>
</tr>
<tr>
<td align="left">DT</td>
<td align="left">max_features</td>
<td align="left">&#x2018;auto&#x2019;, &#x2018;sqrt&#x2019;, &#x2018;log2&#x2019;, 5, 10, 30</td>
</tr>
<tr>
<td/>
<td align="left">max_depth</td>
<td align="left">2, 8, 16, 32, 64, 128</td>
</tr>
<tr>
<td/>
<td align="left">min_samples_split</td>
<td align="left">1, 2, 4, 8, 16, 24</td>
</tr>
<tr>
<td/>
<td align="left">min_samples_leaf</td>
<td align="left">1, 2, 5, 10, 15, 30</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="left">n_estimators</td>
<td align="left">10, 50, 100, 200, 500</td>
</tr>
<tr>
<td/>
<td align="left">max_features</td>
<td align="left">&#x2018;auto&#x2019;, &#x2018;sqrt&#x2019;, &#x2018;log2&#x2019;, 5, 10, 30</td>
</tr>
<tr>
<td/>
<td align="left">max_depth</td>
<td align="left">2, 8, 16, 32, 64, 128</td>
</tr>
<tr>
<td/>
<td align="left">min_samples_split</td>
<td align="left">1, 2, 4, 8, 16, 24</td>
</tr>
<tr>
<td/>
<td align="left">min_samples_leaf</td>
<td align="left">1, 2, 5, 10, 15, 30</td>
</tr>
<tr>
<td align="left">RC</td>
<td align="left">alpha</td>
<td align="left">1, 2, 3, 4, 5, 6, 7, 8, 9, 10</td>
</tr>
<tr>
<td align="left">GP</td>
<td align="left">max_iter_predict</td>
<td align="left">100, 200, 300, 400, 500, 600, 700, 800, 900, 1000</td>
</tr>
<tr>
<td align="left">LDA</td>
<td align="left">solver</td>
<td align="left">&#x2018;lsqr&#x2019;, &#x2018;eigen&#x2019;</td>
</tr>
<tr>
<td/>
<td align="left">shrinkage</td>
<td align="left">&#x2018;empirical&#x2019;, &#x2018;auto&#x2019;, 0.0001, 0.001, 0.01, 0.1, 0.5, 1</td>
</tr>
<tr>
<td align="left">MLP<sup>1</sup></td>
<td align="left">num_layers</td>
<td align="left">1, 3, 5</td>
</tr>
<tr>
<td/>
<td align="left">max_units</td>
<td align="left">30, 50, 100</td>
</tr>
<tr>
<td/>
<td align="left">batchnorm</td>
<td align="left">False, True</td>
</tr>
<tr>
<td/>
<td align="left">dropout</td>
<td align="left">0, 0.5</td>
</tr>
<tr>
<td align="left">RNN<sup>2</sup></td>
<td align="left">hidden_size</td>
<td align="left">30, 50, 100</td>
</tr>
<tr>
<td/>
<td align="left">num_layers</td>
<td align="left">1, 2</td>
</tr>
<tr>
<td/>
<td align="left">bidirectional</td>
<td align="left">False, True</td>
</tr>
<tr>
<td/>
<td align="left">dropout</td>
<td align="left">0, 0.3</td>
</tr>
<tr>
<td/>
<td align="left">model</td>
<td align="left">&#x2018;GRU&#x2019;, &#x2018;LSTM&#x2019;, &#x2018;RNN&#x2019;</td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p><sup>1</sup>DL models (MLP and RNN) are trained using the default Adam optimizer.</p></fn><fn><p><sup>2</sup>RNN treats the features as a fixed sequence processing them one-by-one followed by a linear classifier head.</p></fn></table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_2_2"><label>4.2.2</label><title>K-Fold Cross-Validation</title>
<p>Sometimes it is not clear what constitutes a representative training-validation split. One technique is randomly splitting the training data into <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow></mml:math></inline-formula> chunks of equal sizes. Then, pick a single chunk for validation, and the rest <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> chunks are combined as the training set. After that, repeat the same thing for all <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow></mml:math></inline-formula> chunks (every time, the validation set is a different chunk). The final score is averaged over all the <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow></mml:math></inline-formula> trials. This score will reflect the true strength of the model and training procedure better than an arbitrary splitting. This method is what is known as k-fold cross-validation in the ML field. It is necessary to use this score for hyperparameter tuning and only report the performance on the untouched testing set.</p>
</sec>
</sec>
<sec id="s4_3"><label>4.3</label><title>Model Ensemble</title>
<p>After building and training multiple ML models, one can compare their classifying power and see which one performs better. However, some classifiers may better classify certain types of patients than others. Therefore, choosing which classifier works best in every situation is possible. This method is an example of a model ensemble, which combines the prediction of multiple classifiers to get at least a better accuracy than all constituent classifiers. This paper will combine the top five scoring classifiers in the experiments to build a powerful ensemble model.</p>
</sec>
<sec id="s4_4"><label>4.4</label><title>Feature Selection</title>
<p>Some features are more relevant than others. For instance, the number of languages a patient can speak has no relation to whether they have heart disease or not. In addition, ML models can be susceptible to noise. Moreover, some features might not be easily attainable or measured. For instance, not every patient is willing to spend money on an MRI scan. A simple feature selection technique can be used; select the most correlated features to the output. One well-known scoring function for the features is the <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msup><mml:mi>&#x03C7;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> Test (pronounced as chi-square). It can measure how likely two random variables are to be independent. It is a univariate feature selection method and does not consider the cumulative effect of all the features. For example, two weak features together can be better than a single strong feature. It also ignores correlated features. Therefore, this work will also apply feature selection with multiple variables using Importance Permutation [<xref ref-type="bibr" rid="ref-24">24</xref>]. The earlier method tests the dependence between every feature and the target, while the latter uses a trained model to evaluate the effect of introducing randomness to every feature.</p>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Experiments</title>
<sec id="s5_1"><label>5.1</label><title>Setup</title>
<p>The classical ML models will be trained using the PyCaret package [<xref ref-type="bibr" rid="ref-25">25</xref>], and the DL models will be implemented using the PyTorch package [<xref ref-type="bibr" rid="ref-26">26</xref>]. Moreover, this work has developed an interface between the two packages to facilitate the training and the analysis using the convenient features of PyCaret. It is worth noting that the code is modular and can be easily adapted to work with any other PyTorch classifier. The experiments will be done under Python 3.9 in Google Colaboratory [<xref ref-type="bibr" rid="ref-27">27</xref>]. Under the hood, NumPy [<xref ref-type="bibr" rid="ref-28">28</xref>], Pandas [<xref ref-type="bibr" rid="ref-29">29</xref>], Scikit-Learn [<xref ref-type="bibr" rid="ref-30">30</xref>], Matplotlib [<xref ref-type="bibr" rid="ref-31">31</xref>], and Seaborn [<xref ref-type="bibr" rid="ref-32">32</xref>] are used as supporting libraries. The dataset will be divided into 80&#x0025; as the training set and 20&#x0025; as the testing set. Finally, seven evaluation metrics will be reported per model on the testing set<xref ref-type="fn" rid="fn3"><sup>3</sup></xref><fn id="fn3"><label>3</label><p>The description of the models and evaluation metrics were omitted since they are not part of the contribution.</p></fn>. These exact experiments will be repeated twice, once on the complete feature set and again on the selected subset of features.</p>
</sec>
<sec id="s5_2"><label>5.2</label><title>Results on the Complete Features</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> only represents the top five machine learning models, followed by the top two deep learning models and the model ensemble. It can be noticed that the models with the best overall performance on the complete feature set are the more complex models (the best value for each metric is in boldface font). Here, the ensemble model is the most accurate, achieving 83.15&#x0025; accuracy on the testing set. It is essential to mention that this number cannot be compared directly with other sources as this work uses the entire dataset here, including the patients with missing values. Finally, the MLP model performs the best in the AUC metric over all the other models.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Models&#x2019; test comparisons trained using the complete set of features</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Accuracy</th>
<th align="left">AUC</th>
<th align="left">Recall</th>
<th align="left">Precision</th>
<th align="left">F1</th>
<th align="left">Kappa</th>
<th align="left">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">SVM</td>
<td align="left">83.15</td>
<td align="left">89.47</td>
<td align="left"><bold>87.25</bold></td>
<td align="left">83.18</td>
<td align="left"><bold>85.17</bold></td>
<td align="left"><bold>65.70</bold></td>
<td align="left">65.80</td>
</tr>
<tr>
<td align="left">GP</td>
<td align="left">80.43</td>
<td align="left">86.24</td>
<td align="left">85.29</td>
<td align="left">80.56</td>
<td align="left">82.86</td>
<td align="left">60.12</td>
<td align="left">60.25</td>
</tr>
<tr>
<td align="left">LR</td>
<td align="left">80.43</td>
<td align="left">88.85</td>
<td align="left">83.33</td>
<td align="left">81.73</td>
<td align="left">82.52</td>
<td align="left">60.31</td>
<td align="left">60.32</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="left">79.35</td>
<td align="left">88.83</td>
<td align="left">83.33</td>
<td align="left">80.19</td>
<td align="left">81.73</td>
<td align="left">58.00</td>
<td align="left">58.06</td>
</tr>
<tr>
<td align="left">RC</td>
<td align="left">78.80</td>
<td align="left">00.00</td>
<td align="left">79.41</td>
<td align="left">81.82</td>
<td align="left">80.60</td>
<td align="left">57.26</td>
<td align="left">57.29</td>
</tr>
<tr>
<td align="left">MLP</td>
<td align="left">82.61</td>
<td align="left"><bold>90.58</bold></td>
<td align="left">85.29</td>
<td align="left">83.65</td>
<td align="left">84.47</td>
<td align="left">64.72</td>
<td align="left">64.73</td>
</tr>
<tr>
<td align="left">RNN</td>
<td align="left">80.98</td>
<td align="left">88.30</td>
<td align="left">82.35</td>
<td align="left">83.17</td>
<td align="left">82.76</td>
<td align="left">61.55</td>
<td align="left">61.55</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left"><bold>83.15</bold></td>
<td align="left">89.97</td>
<td align="left">86.00</td>
<td align="left"><bold>83.97</bold></td>
<td align="left">84.95</td>
<td align="left">65.83</td>
<td align="left"><bold>65.90</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3"><label>5.3</label><title>Feature Selection</title>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> compares both feature selection methods, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msup><mml:mi>&#x03C7;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> Test and Importance Permutation, where the latter was computed using the best model from the previous experiments. Despite their distinct ordering, one can see that the top seven features are shared among the two rankings. The same goes for the bottom features, where both methods agree that they are mostly irrelevant to the task at hand. Following best intuition, the top seven features will be selected for the ablation study of feature selection [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Feature importance rankings using two different methods</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_33603-fig-5.png"/></fig>
</sec>
<sec id="s5_4"><label>5.4</label><title>Results on the Selected Features</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> shows that the best model for the selected features is the GP model. It is even better than all other models, including the deep learning and the ensemble model. More interestingly, the achieved accuracy is better than using the full features. Two things can explain this; the feature selection done on the best model on the complete feature set did its job ideally, and the GP model is more suitable when using fewer data under more uncertainty. The final observation is that the RNN model was not better in both experiments. However, its striking consistency despite using vastly different features demonstrates its robustness. This can mean that it has played a vital role in the model ensemble in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Models&#x2019; test comparisons trained using the selected set of features</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Accuracy</th>
<th align="left">AUC</th>
<th align="left">Recall</th>
<th align="left">Precision</th>
<th align="left">F1</th>
<th align="left">Kappa</th>
<th align="left">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">GP</td>
<td align="left"><bold>84.24</bold></td>
<td align="left">87.45</td>
<td align="left"><bold>89.22</bold></td>
<td align="left">83.49</td>
<td align="left"><bold>86.26</bold></td>
<td align="left"><bold>67.83</bold></td>
<td align="left"><bold>68.04</bold></td>
</tr>
<tr>
<td align="left">LR</td>
<td align="left">80.43</td>
<td align="left">88.49</td>
<td align="left">82.35</td>
<td align="left">82.35</td>
<td align="left">82.35</td>
<td align="left">60.40</td>
<td align="left">60.40</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">80.43</td>
<td align="left">88.70</td>
<td align="left">87.25</td>
<td align="left">79.46</td>
<td align="left">83.18</td>
<td align="left">59.92</td>
<td align="left">60.30</td>
</tr>
<tr>
<td align="left">RC</td>
<td align="left">79.89</td>
<td align="left">00.00</td>
<td align="left">80.39</td>
<td align="left">82.83</td>
<td align="left">81.59</td>
<td align="left">59.45</td>
<td align="left">59.48</td>
</tr>
<tr>
<td align="left">LDA</td>
<td align="left">79.89</td>
<td align="left">88.40</td>
<td align="left">81.37</td>
<td align="left">82.18</td>
<td align="left">81.77</td>
<td align="left">59.35</td>
<td align="left">59.35</td>
</tr>
<tr>
<td align="left">MLP</td>
<td align="left">81.52</td>
<td align="left">88.73</td>
<td align="left">85.29</td>
<td align="left">82.08</td>
<td align="left">83.65</td>
<td align="left">62.42</td>
<td align="left">62.48</td>
</tr>
<tr>
<td align="left">RNN</td>
<td align="left">80.98</td>
<td align="left">88.07</td>
<td align="left"><bold>89.22</bold></td>
<td align="left">79.13</td>
<td align="left">83.87</td>
<td align="left">60.89</td>
<td align="left">61.55</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left">81.12</td>
<td align="left"><bold>89.00</bold></td>
<td align="left">81.81</td>
<td align="left"><bold>83.85</bold></td>
<td align="left">82.75</td>
<td align="left">61.89</td>
<td align="left">62.03</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_5"><label>5.5</label><title>Limitations and Future Work</title>
<p>The most significant limitation of this work is the non-sufficient data to draw statistically significant conclusions. This research can benefit greatly from more rich and diverse datasets publicly available for general use. To build on this work, one can collect more data and introduce new features that can improve the discriminative power of these classifiers. For example, the developed system can be used to collect anonymized data for similar future applications. Moreover, it might be interesting to predict the hardest available features using other more attainable information, extending the usability to a broader audience.</p>
</sec>
</sec>
<sec id="s6"><label>6</label><title>Conclusion</title>
<p>This research tackled the prediction problem of the UCI heart disease dataset in a feature-limited setting. It followed a proper data science workflow from data analysis and preprocessing to model building, training, and evaluation. In particular, this work trained multiple classical machine learning and deep learning models including a hybrid model of all the top performing models. Each model was tested under different hyperparameter configurations using a validation data split. Then, this paper applied feature selection and repeated the same process to get a model that uses only a subset of the features with competitive performance. This makes it easier for patients with limited access to benefit from the system while achieving an 84.24&#x0025; accuracy, 89.22&#x0025; Recall, and 83.49&#x0025; Precision. As a result, this effort satisfied two critical goals of machine learning: interpretation and prediction.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>World Health Organization</collab></person-group>, &#x201C;<article-title>The top 10 causes of death</article-title>,&#x201D; <year>2020</year>. [Online]. Available: <uri xlink:href="https://www.who.int/news-room/fact-sheets/detail/the-top-10-causes-of-death">https://www.who.int/news-room/fact-sheets/detail/the-top-10-causes-of-death</uri>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. J.</given-names> <surname>Benjamin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Muntner</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alonso</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Bittencourt</surname></string-name>, <string-name><given-names>C. W.</given-names> <surname>Callaway</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Heart disease and stroke statistics&#x2014;2019 update: A report from the American heart association</article-title>,&#x201D; <source>Circulation</source>, vol. <volume>139</volume>, no. <issue>10</issue>, pp. <fpage>56</fpage>&#x2013;<lpage>528</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Bashir</surname></string-name>, <string-name><given-names>Z. S.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>F. H.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Anjum</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Bashir</surname></string-name></person-group>, &#x201C;<article-title>Improving heart disease prediction using feature selection approaches</article-title>,&#x201D; in <conf-name>Proc. of IBCAST</conf-name>, <conf-loc>Islamabad, Pakistan</conf-loc>, pp. <fpage>619</fpage>&#x2013;<lpage>623</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A. H.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>S. Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>P. S.</given-names> <surname>Hong</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>Cheng</surname></string-name> and <string-name><given-names>E. J.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<chapter-title>HDPS: Heart disease prediction system</chapter-title>,&#x201D; in <source>2011 Computing in Cardiology</source>, <publisher-name>IEEE</publisher-name>, <publisher-loc>Hangzhou, China</publisher-loc>, pp. <fpage>557</fpage>&#x2013;<lpage>560</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Center for Machine Learning and Intelligent Systems</collab></person-group>, &#x201C;<article-title>UCI machine learning repository</article-title>,&#x201D; <source>Heart Disease Data Set</source>, <year>1988</year>. [Online]. Available: <uri xlink:href="https://archive.ics.uci.edu/ml/datasets/heart+disease">https://archive.ics.uci.edu/ml/datasets/heart+disease</uri>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. P.</given-names> <surname>Alex </surname></string-name></person-group> and <person-group person-group-type="author"><string-name><given-names>S. P.</given-names> <surname>Shaji</surname></string-name></person-group>, &#x201C;<article-title>Prediction and diagnosis of heart disease patients using data mining technique</article-title>,&#x201D; in <conf-name>2019 Int. Conf. on Communication and Signal Processing (ICCSP)</conf-name>, <conf-loc>Dalian, China</conf-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>848</fpage>&#x2013;<lpage>852</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. B.</given-names> <surname>Gokulnath</surname></string-name> and <string-name><given-names>S. P.</given-names> <surname>Shantharajah</surname></string-name></person-group>, &#x201C;<article-title>An optimized feature selection based on genetic approach and support vector machine for heart disease</article-title>,&#x201D; <source>Cluster Computing</source>, vol. <volume>22</volume>, no. <issue>S6</issue>, pp. <fpage>14777</fpage>&#x2013;<lpage>14787</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. U.</given-names> <surname>Haq</surname></string-name>, <string-name><given-names>J. P.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Memon</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nazir</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>A hybrid intelligent system framework for the prediction of heart disease using machine learning algorithms</article-title>,&#x201D; <source>Mobile Information Systems</source>, vol. <volume>2018</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>21</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Courville</surname></string-name></person-group>, <source>Deep Learning</source>. <publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>26</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W. S.</given-names> <surname>Noble</surname></string-name></person-group>, &#x201C;<article-title>What is a support vector machine?,</article-title>&#x201D; <source>Nature Biotechnology</source>, vol. <volume>24</volume>, no. <issue>12</issue>, pp. <fpage>1565</fpage>&#x2013;<lpage>1567</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H. -T.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Tang</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Enhancing ontology-driven diagnostic reasoning with a symptom-dependency-aware Na&#x00EF;ve Bayes classifier</article-title>,&#x201D; <source>BMC Bioinformatics</source>, vol. <volume>20</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Jain</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Nagrath</surname></string-name></person-group>, &#x201C;<article-title>Heart disease prediction using classification (Naive Bayes)</article-title>,&#x201D; in <conf-name>Int. Conf. on Computing, Communications, and Cyber-Security</conf-name>, <conf-loc>Springer, Singapore</conf-loc>, pp. <fpage>561</fpage>&#x2013;<lpage>573</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>C. R.</given-names> <surname>Shalizi</surname></string-name></person-group>, &#x201C;<source>Advanced Data Analysis from an Elementary Point of View</source>,&#x201D; <publisher-loc>Pittsburgh, Pennsylvania, USA</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>, pp. <fpage>234</fpage>&#x2013;<lpage>260</lpage>, <year>2019</year>. [Online]. Available: <uri xlink:href="https://www.stat.cmu.edu/~cshalizi/ADAfaEPoV/ADAfaEPoV.pdf">https://www.stat.cmu.edu/~cshalizi/ADAfaEPoV/ADAfaEPoV.pdf</uri>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Scikit-Learn</collab></person-group>, &#x201C;<article-title>Ridge regression and classification, linear models</article-title>,&#x201D; <year>2022</year>. [Online]. Available: <uri xlink:href="https://scikit-learn.org/stable/modules/linear_model.html#ridge-regression-and-classification">https://scikit-learn.org/stable/modules/linear_model.html#ridge-regression-and-classification</uri>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>G. J.</given-names> <surname>McLachlan</surname></string-name></person-group>, &#x201C;<chapter-title>Logistic discrimination</chapter-title>,&#x201D; in <source>Discriminant Analysis and Statistical Pattern Recognition</source>, <edition>1<sup>st</sup> ed.</edition>, vol. <volume>1</volume>. <publisher-loc>Queensland, Australia</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>, pp. <fpage>255</fpage>&#x2013;<lpage>282</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>D. J. C.</given-names> <surname>MacKay</surname></string-name></person-group>, &#x201C;<article-title>Introduction to Gaussian processes</article-title>,&#x201D; Cambridge, United Kingdom, <year>1998</year>. [Online]. Available: <uri xlink:href="https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.81.1927&amp;rep=rep1&amp;type=pdf">https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.81.1927&amp;rep=rep1&amp;type=pdf</uri>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Friedl</surname></string-name> and <string-name><given-names>C. E.</given-names> <surname>Brodley</surname></string-name></person-group>, &#x201C;<article-title>Decision tree classification of land cover from remotely sensed data</article-title>,&#x201D; <source>Remote Sensing of Environment</source>, vol. <volume>61</volume>, no. <issue>3</issue>, pp. <fpage>399</fpage>&#x2013;<lpage>409</lpage>, <year>1997</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Liaw</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Wiener</surname></string-name></person-group>, &#x201C;<article-title>Classification and regression by randomForest</article-title>,&#x201D; <source>R News</source>, vol. <volume>2</volume>, pp. <fpage>18</fpage>&#x2013;<lpage>22</lpage>, <year>2002</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. S.</given-names> <surname>Yadav</surname></string-name>, <string-name><given-names>S. M.</given-names> <surname>Jadhav</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nagrale</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Patil</surname></string-name></person-group>, &#x201C;<article-title>Application of machine learning for the detection of heart disease</article-title>,&#x201D; in <conf-name>ICIMI</conf-name>, <conf-loc>Bangalore, India</conf-loc>, pp. <fpage>165</fpage>&#x2013;<lpage>172</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Medsker</surname></string-name> and <string-name><given-names>L. C.</given-names> <surname>Jain</surname></string-name></person-group>, &#x201C;<chapter-title>Recurrent neural networks: Design and applications</chapter-title>,&#x201D; in <source>International Series on Computational Intelligence</source>, <edition>1<sup>st</sup> ed.</edition>, vol. <volume>1</volume>. <publisher-loc>Washington D.C., USA</publisher-loc>: <publisher-name>CRC Press</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>1999</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Mohan</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Thirumalai</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Srivastava</surname></string-name></person-group>, &#x201C;<article-title>Effective heart disease prediction using hybrid machine learning techniques</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>81542</fpage>&#x2013;<lpage>81554</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Katarya</surname></string-name> and <string-name><given-names>S. K.</given-names> <surname>Meena</surname></string-name></person-group>, &#x201C;<article-title>Machine learning techniques for heart disease prediction: A comparative study and analysis</article-title>,&#x201D; <source>Health and Technology</source>, vol. <volume>11</volume>, no. <issue>1</issue>, pp. <fpage>87</fpage>&#x2013;<lpage>97</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Dileep</surname></string-name>, <string-name><given-names>K. N.</given-names> <surname>Rao</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Bodapati</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Gokuruboyina</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Peddi</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>An automatic heart disease prediction using cluster-based bi-directional LSTM (C-BiLSTM) algorithm</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>34</volume>, no. <issue>9</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Scikit-Learn</collab></person-group>, &#x201C;<article-title>Chi-squared statistic test feature selection</article-title>,&#x201D; <year>2022</year>. [Online]. Available: <uri xlink:href="https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.chi2.html">https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.chi2.html</uri>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ali</surname></string-name></person-group>, &#x201C;<article-title>PyCaret: An open source, low-code machine learning library in python</article-title>,&#x201D; <year>2020</year>. [Online]. Available: <uri xlink:href="https://www.pycaret.org">https://www.pycaret.org</uri>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Paszke</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Gross</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Massa</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Lerer</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Bradbury</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>PyTorch: An imperative style, high-performance deep learning library</article-title>,&#x201D; in <conf-name>Conf. on Neural Information Processing Systems</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>, pp. <fpage>8024</fpage>&#x2013;<lpage>8035</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Bisong</surname></string-name></person-group>, &#x201C;<chapter-title>Google Colaboratory</chapter-title>,&#x201D; in <source>Building Machine Learning and Deep Learning Models on Google Cloud Platform</source>, <edition>1<sup>st</sup> ed.</edition>, vol. <volume>1</volume>. <publisher-loc>Berkeley, CA, USA</publisher-loc>: <publisher-name>Apress</publisher-name>, pp. <fpage>59</fpage>&#x2013;<lpage>64</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. R.</given-names> <surname>Harris</surname></string-name>, <string-name><given-names>K. J.</given-names> <surname>Millman</surname></string-name>, <string-name><given-names>S. J.</given-names> <surname>Van Der Walt</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Gommers</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Virtanen</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Array programming with NumPy</article-title>,&#x201D; <source>Nature</source>, vol. <volume>585</volume>, no. <issue>7825</issue>, pp. <fpage>357</fpage>&#x2013;<lpage>362</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>The Pandas Development Team</collab></person-group>, &#x201C;<article-title>Pandas 1.4.2</article-title>.,&#x201D; <year>2022</year>. [Online]. Available: <uri xlink:href="https://pandas.pydata.org/">https://pandas.pydata.org/</uri>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Pedregosa</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Varoquaux</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Gramfort</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Michel</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Thirion</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Scikit-learn: Machine learning in python</article-title>,&#x201D; <source>Journal of Machine Learning Research</source>, vol. <volume>12</volume>, pp. <fpage>2825</fpage>&#x2013;<lpage>2830</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. D.</given-names> <surname>Hunter</surname></string-name></person-group>, &#x201C;<article-title>Matplotlib: A 2D graphics environment</article-title>,&#x201D; <source>Computing in Science &#x0026; Engineering</source>, vol. <volume>9</volume>, no. <issue>3</issue>, pp. <fpage>90</fpage>&#x2013;<lpage>95</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Waskom</surname></string-name></person-group>, &#x201C;<article-title>Seaborn: Statistical data visualization</article-title>,&#x201D; <source>Journal of Open Source Software</source>, vol. <volume>6</volume>, no. <issue>60</issue>, pp.&#x00A0;<fpage>3021</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Scikit-Learn</collab></person-group>, &#x201C;<article-title>User guide to feature selection using scikit-learn</article-title>,&#x201D; <year>2022</year>. [Online]. Available: <uri xlink:href="https://scikit-learn.org/stable/modules/feature_selection.html">https://scikit-learn.org/stable/modules/feature_selection.html</uri>.</mixed-citation></ref>
</ref-list>
</back>
</article>







