<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CSSE</journal-id>
<journal-id journal-id-type="nlm-ta">CSSE</journal-id>
<journal-id journal-id-type="publisher-id">CSSE</journal-id>
<journal-title-group>
<journal-title>Computer Systems Science &#x0026; Engineering</journal-title>
</journal-title-group>
<issn pub-type="ppub">0267-6192</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">22152</article-id>
<article-id pub-id-type="doi">10.32604/csse.2022.022152</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Towards Improving Predictive Statistical Learning Model Accuracy by Enhancing Learning Technique</article-title><alt-title alt-title-type="left-running-head">Towards Improving Predictive Statistical Learning Model Accuracy by Enhancing Learning Technique</alt-title><alt-title alt-title-type="right-running-head">Towards Improving Predictive Statistical Learning Model Accuracy by Enhancing Learning Technique</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Algarni</surname><given-names>Ali</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ragab</surname><given-names>Mahmoud</given-names></name>
<xref ref-type="aff" rid="aff-2">2</xref>
<xref ref-type="aff" rid="aff-3">3</xref>
<xref ref-type="aff" rid="aff-4">4</xref>
<email>samih_montser@sci.svu.edu.eg</email>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Alamri</surname><given-names>Wardah</given-names></name>
<xref ref-type="aff" rid="aff-5">5</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Mostafa</surname><given-names>Samih M.</given-names></name>
<xref ref-type="aff" rid="aff-6">6</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>Statistics Department, Faculty of Science, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Information Technology Department, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-3"><label>3</label><institution>Mathematics Department, Faculty of Science, Al-Azhar University</institution>, <addr-line>Naser City, 11884</addr-line>, <country>Egypt</country></aff>
<aff id="aff-4"><label>4</label><institution>Centre of Artificial Intelligence for Precision Medicines, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-5"><label>5</label><institution>Educational Technology Department, Educational Graduate Studies Faculty, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-6"><label>6</label><institution>Computer Science Department, Faculty of Computers and Information, South Valley University</institution>, <addr-line>Qena, 83523</addr-line>, <country>Egypt</country></aff>
</contrib-group><author-notes><corresp id="cor1">&#x002A;Corresponding Author: Mahmoud Ragab. Email: <email>mragab@kau.edu.sa</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-11-23"><day>23</day>
<month>11</month>
<year>2021</year></pub-date>
<volume>42</volume>
<issue>1</issue>
<fpage>303</fpage>
<lpage>318</lpage>
<history>
<date date-type="received"><day>29</day><month>7</month><year>2021</year></date>
<date date-type="accepted"><day>30</day><month>8</month><year>2021</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Algarni et al.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Algarni et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CSSE_22152.pdf"></self-uri>
<abstract>
<p>The accuracy of the statistical learning model depends on the learning technique used which in turn depends on the dataset&#x2019;s values. In most research studies, the existence of missing values (MVs) is a vital problem. In addition, any dataset with MVs cannot be used for further analysis or with any data driven tool especially when the percentage of MVs are high. In this paper, the authors propose a novel algorithm for dealing with MVs depending on the feature selection (FS) of similarity classifier with fuzzy entropy measure. The proposed algorithm imputes MVs in cumulative order. The candidate feature to be manipulated is selected using similarity classifier with Parkash&#x2019;s fuzzy entropy measure. The predictive model to predict MVs within the candidate feature is the Bayesian Ridge Regression (BRR) technique. Furthermore, any imputed features will be incorporated within the BRR equation to impute the MVs in the next chosen incomplete feature. The proposed algorithm was compared against some practical state-of-the-art imputation methods by conducting an experiment on four medical datasets which were gathered from several databases repository with MVs generated from the three missingness mechanisms. The evaluation metrics of mean absolute error (MAE), root mean square error (RMSE) and coefficient of determination (<italic>R</italic><sup>2</sup> score) were used to measure the performance. The results exhibited that performance vary depending on the size of the dataset, amount of MVs and the missingness mechanism type. Moreover, compared to other methods, the results showed that the proposed method gives better accuracy and less error in most cases.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Bayesian ridge regression</kwd>
<kwd>fuzzy entropy measure</kwd>
<kwd>feature selection</kwd>
<kwd>imputation</kwd>
<kwd>missing values</kwd>
<kwd>missingness mechanisms</kwd>
<kwd>similarity classifier</kwd>
<kwd>medical dataset</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>MVs are considered a critical problem that can occur in many scientific areas such as biological, psychological, or medical [<xref ref-type="bibr" rid="ref-1">1</xref>]. Commonly, many reasons may lead to the occurrence of MVs, for instance, wrong data entry, improper data collection, management of similar but not identical datasets and malfunctioning measurement equipment [<xref ref-type="bibr" rid="ref-2">2</xref>]. Machine learning (ML), big data and any data driven tool require high data quality which results in good analysis and outcomes. The existence of MVs within a dataset can result in problems, for instance, bad data analysis, reducing the research results obtained from such dataset and presenting amount of bias [<xref ref-type="bibr" rid="ref-3">3</xref>]. To this end, significant information is incorporated within MVs which should be manipulated before using the incomplete dataset with any data driven tool. Furthermore, many researches were done and novel algorithms were proposed to solve the problem of MVs, especially in medical data [<xref ref-type="bibr" rid="ref-4">4</xref>]. Nevertheless, several imputation algorithms may result in poor imputation and may fail in handling all MVs in the dataset. In addition, they may not deal with all missingness mechanisms. These shortcomings of these algorithms encouraged the authors to propose a novel algorithm introduced in this paper. The proposed algorithm utilizes the most significant feature to impute MVs in cumulative order. Besides MVs, FS also affects the ML model performance.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Feature Selection</title>
<p>High dimensionality data is problematic especially in circumstances when a dataset contains a few numbers of training instances and a large number of features. This type of data commonly exists in medicine where cost and time problems may limit the number of training observations, while the number of diseases increases through the years [<xref ref-type="bibr" rid="ref-5">5</xref>]. FS helps to overcome the problem of high dimensionality by selecting a subset of features that have a strong relationship with the target feature. In addition, in the existence of MVs the FS is considered a vital preprocessing step such as correlation, mutual information and fuzzy FS. Dropping features that hold a large number of MVs (<italic>e.g</italic>., &#x003E;50%) is an easy solution. But such a solution may result in bad analysis, losing the ability to recognize statistically significant variations and may also generates bias. Missingness mechanisms have a large effect on FS that&#x2019;s why before applying any FS technique missingness mechanisms need to be taken into consideration [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Missingness Mechanisms</title>
<p>Before introducing different methods for handling MVs, it is essential to present the different types of missingness mechanisms (<italic>i.e</italic>., the reason for the occurrence of MVs in data). MVs are commonly classified to one of three MVs mechanisms [<xref ref-type="bibr" rid="ref-6">6</xref>]:<list list-type="bullet"><list-item>
<p><italic>Missing Completely at Random (MCAR):</italic> This type of MVs mechanisms happens when the probability of the existence of MVs is independent from any other features in the data. From statistical perspective, MCAR can be stated as in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> [<xref ref-type="bibr" rid="ref-1">1</xref>].</p></list-item></list></p>
<p><disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">&#x2205;</mml:mi><mml:mspace width="thickmathspace" /></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-1">
<mml:math id="mml-ieqn-1"><mml:mi>M</mml:mi></mml:math>
</inline-formula> and <inline-formula id="ieqn-2">
<mml:math id="mml-ieqn-2"><mml:mi>Y</mml:mi></mml:math>
</inline-formula> represent the missing and observed data respectively. The conditional probability is denoted by <inline-formula id="ieqn-3">
<mml:math id="mml-ieqn-3"><mml:mi>f</mml:mi></mml:math>
</inline-formula> and <inline-formula id="ieqn-4">
<mml:math id="mml-ieqn-4"><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:math>
</inline-formula> represents the unknown parameter. An example of MCAR MVs occur when the measuring equipment stops working correctly [<xref ref-type="bibr" rid="ref-7">7</xref>].<list list-type="bullet"><list-item>
<p><italic>Missing at Random (MAR):</italic> In this mechanism the relationship between MVs and other features existed in the dataset is a dependent relationship. In other words, the probability of the occurrence of MVs depends on observed values in other features and not on other MVs in the target feature [<xref ref-type="bibr" rid="ref-8">8</xref>]. MAR can be represented as in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> [<xref ref-type="bibr" rid="ref-1">1</xref>].</p></list-item></list></p>
<p><disp-formula id="eqn-2"><label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-5">
<mml:math id="mml-ieqn-5"><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> and <inline-formula id="ieqn-6">
<mml:math id="mml-ieqn-6"><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> are the missing and observed values from <inline-formula id="ieqn-7">
<mml:math id="mml-ieqn-7"><mml:mi>Y</mml:mi></mml:math>
</inline-formula> respectively. From a medical perspective, this situation may occur when an experiment has not been accomplished because a feature within the dataset shows that the patient is a woman for example [<xref ref-type="bibr" rid="ref-7">7</xref>].<list list-type="bullet"><list-item>
<p><italic>Missing Not at Random (MNAR):</italic> For this type, there is a dependent relationship between the MVs and the observed data. MNAR can be expressed using <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> [<xref ref-type="bibr" rid="ref-1">1</xref>].</p></list-item></list></p>
<p><disp-formula id="eqn-3"><label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>M</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-8">
<mml:math id="mml-ieqn-8"><mml:mi>&#x03B8;</mml:mi></mml:math>
</inline-formula> (<italic>i.e</italic>., parameter of the distribution <inline-formula id="ieqn-9">
<mml:math id="mml-ieqn-9"><mml:mi>Y</mml:mi></mml:math>
</inline-formula>). is estimated from the detected data The distribution of the missingness is denoted by <inline-formula id="ieqn-10">
<mml:math id="mml-ieqn-10"><mml:mi mathvariant="normal">&#x2205;</mml:mi><mml:mo>.</mml:mo></mml:math>
</inline-formula> An example of MNAR MVs, when people having a too high or too low income reject to reveal it [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Handling Missing Data</title>
<p>The simplest methods for handling MVs are traditional methods. Traditional methods can be deletion (<italic>i.e</italic>., delete instances that hold MVs), mean or median (<italic>i.e</italic>., replace the MVs with the mean or median of the feature that holds MVs) substitution [<xref ref-type="bibr" rid="ref-9">9</xref>]. Deletion can be case deletion or pairwise deletion. Case deletion (a.k.a., listwise) in which any instance holds MVs is dropped from analysis. In many statistical packages listwise is the default choice [<xref ref-type="bibr" rid="ref-10">10</xref>]. Pairwise deletion is considered a selective method, which tries to minimize the lost amount of data instances that occur in case of using the listwise method by including into the analysis the instances with M<italic>Vs</italic>. In other words, pairwise deletion will drop only particular features with MVs from the analysis and use the remainder features with no M<italic>Vs</italic>. The selection of features varies from analysis to another depending on the missingness. Using deletion methods results in reducing the data size [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
<p>The other methods that overcome the defects of deletion methods are called imputation methods. In imputation methods, predefined (mean, median, etc.) or estimated (using statistical methods, ML algorithms, etc.) value is used instead of MVs [<xref ref-type="bibr" rid="ref-12">12</xref>]. Imputation is classified into single and multiple imputation. In single imputation, MVs are imputed by a value one time. Though, single imputation does not require computational resources it can result in biased results [<xref ref-type="bibr" rid="ref-3">3</xref>]. In multiple imputation, <inline-formula id="ieqn-11">
<mml:math id="mml-ieqn-11"><mml:mi>m</mml:mi></mml:math>
</inline-formula> copies from the original dataset are generated. In each generated dataset, MVs are imputed using single imputation techniques. The final imputed dataset is the average analysis of the <inline-formula id="ieqn-12">
<mml:math id="mml-ieqn-12"><mml:mi>m</mml:mi></mml:math>
</inline-formula> imputed datasets [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]. ML algorithms can also be used to predict MVs depending on using the available information within the given dataset. Some examples of ML techniques that are used to predict MVs include linear regression, k-nearest neighbour (KNN), decision trees [<xref ref-type="bibr" rid="ref-1">1</xref>] and BRR. BRR is the predictive model used within the proposed algorithm in this paper to predict MVs, which can be expressed using <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p><disp-formula id="eqn-4"><label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mi>y</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>where:</p>
<p><disp-formula id="eqn-5">
<mml:math id="mml-eqn-5" display="block"><mml:mi>&#x03BC;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-6">
<mml:math id="mml-eqn-6" display="block"><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-7">
<mml:math id="mml-eqn-7" display="block"><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-8">
<mml:math id="mml-eqn-8" display="block"><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>The target feature is denoted by <inline-formula id="ieqn-13">
<mml:math id="mml-ieqn-13"><mml:mi>y</mml:mi></mml:math>
</inline-formula> which is distributed as a normal distribution characterized by mean <inline-formula id="ieqn-14">
<mml:math id="mml-ieqn-14"><mml:mi>&#x03BC;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mi>X</mml:mi></mml:math>
</inline-formula> and variance <inline-formula id="ieqn-15">
<mml:math id="mml-ieqn-15"><mml:mi>&#x03B1;</mml:mi></mml:math>
</inline-formula>. <inline-formula id="ieqn-16">
<mml:math id="mml-ieqn-16"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">q</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</inline-formula> denotes the unknown parameters and <inline-formula id="ieqn-17">
<mml:math id="mml-ieqn-17"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">q</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</inline-formula> denotes the independent features. The number of independent features is represented by <inline-formula id="ieqn-18">
<mml:math id="mml-ieqn-18"><mml:mi>q</mml:mi></mml:math>
</inline-formula>. <inline-formula id="ieqn-19">
<mml:math id="mml-ieqn-19"><mml:mi>&#x03B1;</mml:mi></mml:math>
</inline-formula> and <inline-formula id="ieqn-20">
<mml:math id="mml-ieqn-20"><mml:mi>&#x03BB;</mml:mi></mml:math>
</inline-formula> represent the regularization parameters which are assessed jointly while fitting the model through maximizing the log marginal likelihood and both of them are assumed to be distributed as gamma distribution. <inline-formula id="ieqn-21">
<mml:math id="mml-ieqn-21"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</inline-formula> are hyper-parameters of the gamma prior distributions.</p>
<p>The rest of the paper is organized as follows: Section 2 presents a brief literature review about analysis of M<italic>Vs</italic>. Sections 3 and 4 reveal the proposed algorithm and explains in detail the experimental setup, respectively. Section 5 is devoted to the presentation of the results and discussion while section 6 concludes this paper and exhibits some perspectives of future work.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p>Hot-Deck (HD) imputation is a popular choice for manipulating MVs in survey research. Hot deck technique finds a similar dataset and imputes MVs by substituting MVs with an observed value from this dataset. Although, this technique is easy to implement but it may be computationally [<xref ref-type="bibr" rid="ref-16">16</xref>]. The method that looks like hot-deck imputation but the data source and current data set must be different from each other is known as Cold-Deck imputation [<xref ref-type="bibr" rid="ref-17">17</xref>]. In many time-series and longitudinal data, one of the most common and used imputation methods is the Last Observation Carried Forward (LOCF). This method imputes each missing value using the last observed value from the same data [<xref ref-type="bibr" rid="ref-18">18</xref>]. The maximum likelihood method can also be used to manipulate M<italic>Vs</italic>. The maximum likelihood assumes that the detected data is a sample taken from a multivariate normal distribution. After the estimation of the parameters using the available information, the MVs are imputed depending on the estimated parameters [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. In regression imputation, the complete features are used to predict the MVs within the features that contain M<italic>Vs</italic>. The predicted values is used to impute the M<italic>Vs</italic>. Regression imputation keeps all data and hence overcomes the pairwise or listwise deletion and does not change the shape of the distribution. In regression imputation, no information is changed or added and the standard error is reduced, hence, little or no biased predictions are generated from the imputation stage [<xref ref-type="bibr" rid="ref-21">21</xref>]. Expectation-Maximization Imputation (EMI) is a kind of the maximum likelihood technique that can be used to manipulate M<italic>Vs</italic>. EMI uses the values assessed by the use of maximum likelihood methods to impute MVs [<xref ref-type="bibr" rid="ref-22">22</xref>]. This method begins with the expectation step, through which the parameters (<italic>e.g</italic>., means, covariances, and variances) are assessed, possibly by the use of listwise deletion. Predicting MVs is implemented after creation of a regression equation by the use of the estimated parameters. In the maximization step, the regression equations are used to impute M<italic>Vs</italic>. By repeating the expectation and maximization steps until the covariance matrix for the successive iteration is almost the same as that for the previous one. When there is large amount of MVs, EMI method require long time to converge. EMI can result in biased parameter assessments, hence, the standard error is underestimated [<xref ref-type="bibr" rid="ref-21">21</xref>]. KNN imputation technique is considered as one of the most commonly used imputation techniques KNN detects between the complete instances the k most nearest neighbors of a missing data point. The MVs are then imputed with an average of the values of its neighbors in this point. The performance of KNN is extremely bounded especially when the percentage of MVs is high. A simple improvement for manipulating MVs using KNN lies in looking for incomplete neighbors (<italic>i.e</italic>., act as donors) of an instance given that these neighbors are detected for the features missing instances. This method is known as incomplete case k-nearest neighbors imputation (ICkNNI). ICkNNI in somewhat considered a complex method [<xref ref-type="bibr" rid="ref-23">23</xref>]. Methods that manipulate MVs problems directly without the need of any deletion or imputation step have been developed. For example, logistic regression with MVs by using a Gaussian mixture model to assess the conditional density functions was performed by the authors in [<xref ref-type="bibr" rid="ref-24">24</xref>]. For clustering purposes, the Kernel Spectral Clustering (KSC) algorithm was proposed, which encodes as a set of supplemental soft constraints the partially detected features [<xref ref-type="bibr" rid="ref-25">25</xref>]. MLPimpute is a novel algorithm for handling MVs depending multilayer perceptron (MLP) networks was proposed. Although MLP exhibits a good accuracy the relationship between data genes is not sufficient for the method [<xref ref-type="bibr" rid="ref-26">26</xref>]. An iterative learning method consists of fuzzy k-means and decision trees was used to manipulate M<italic>Vs</italic>. When this iterative learning compared with KNNimpute it exhibits a better accuracy [<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Algorithm</title>
<p>This section aims to introduce and elaborate the proposed algorithm in details. The next procedural steps help in clarifying the proposed algorithm.<list list-type="bullet"><list-item>
<p><bold>Splitting Dataset:</bold> The proposed algorithm gets a dataset <inline-formula id="ieqn-22">
<mml:math id="mml-ieqn-22"><mml:mi>D</mml:mi></mml:math>
</inline-formula> as input which incorporates MVs, then creates from <inline-formula id="ieqn-23">
<mml:math id="mml-ieqn-23"><mml:mi>D</mml:mi></mml:math>
</inline-formula> two subsets. The first set <inline-formula id="ieqn-24">
<mml:math id="mml-ieqn-24"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> holds all features with no MVs and the second set <inline-formula id="ieqn-25">
<mml:math id="mml-ieqn-25"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> holds all features with M<italic>Vs</italic>. The target feature was assumed to be perfect feature (<italic>i.e</italic>., does not hold MVs), thus <inline-formula id="ieqn-26">
<mml:math id="mml-ieqn-26"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> holds all perfect features besides the target feature <inline-formula id="ieqn-27">
<mml:math id="mml-ieqn-27"><mml:mi>y</mml:mi></mml:math>
</inline-formula>.</p></list-item><list-item>
<p><bold>Feature Selection:</bold> The proposed algorithm uses the FS of fuzzy entropy measure introduced by Parkash et al. given by <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> with the similarity-based classification [<xref ref-type="bibr" rid="ref-28">28</xref>]</p></list-item></list></p>
<p><disp-formula id="eqn-19"><label>(5)</label>
<mml:math id="mml-eqn-19" display="block"><mml:mi>S</mml:mi><mml:mrow><mml:mo>&#x27E8;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x27E9;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>t</mml:mi></mml:mfrac></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>t</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mi>p</mml:mi></mml:msup></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>v</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mi>p</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mfrac><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mfrac></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">v</mml:mi></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>t</mml:mi></mml:msup></mml:mrow></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-28">
<mml:math id="mml-ieqn-28"><mml:mi>t</mml:mi></mml:math>
</inline-formula> represents the number of features of varied types <inline-formula id="ieqn-29">
<mml:math id="mml-ieqn-29"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> that can be observed from the objects, the ideal vector <inline-formula id="ieqn-30">
<mml:math id="mml-ieqn-30"><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math>
</inline-formula> should be determined for every class <inline-formula id="ieqn-31">
<mml:math id="mml-ieqn-31"><mml:mi>i</mml:mi></mml:math>
</inline-formula>. <inline-formula id="ieqn-32">
<mml:math id="mml-ieqn-32"><mml:mrow><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math>
</inline-formula> represents vectors which belong to known class. <inline-formula id="ieqn-33">
<mml:math id="mml-ieqn-33"><mml:mi>m</mml:mi></mml:math>
</inline-formula> is the power value that is obtained from the generalized mean from the generalized Lukasiewicz structure. The parameter <inline-formula id="ieqn-34">
<mml:math id="mml-ieqn-34"><mml:mi>p</mml:mi></mml:math>
</inline-formula> can be detected from the generalized Lukasiewicz structure. <inline-formula id="ieqn-35">
<mml:math id="mml-ieqn-35"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> is a weight parameter. The weights were set as one.</p>
<p><disp-formula id="eqn-9"><label>(6)</label>
<mml:math id="mml-eqn-9" display="block"><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>A</mml:mi><mml:mo>;</mml:mo><mml:mi>w</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:mrow><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-36">
<mml:math id="mml-ieqn-36"><mml:mi>H</mml:mi></mml:math>
</inline-formula> represents the fuzzy entropy. <inline-formula id="ieqn-37">
<mml:math id="mml-ieqn-37"><mml:mi>j</mml:mi></mml:math>
</inline-formula> represents the number of features and the fuzzy values are denoted by<inline-formula id="ieqn-38">
<mml:math id="mml-ieqn-38"><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</inline-formula>. <inline-formula id="ieqn-39">
<mml:math id="mml-ieqn-39"><mml:mi>A</mml:mi></mml:math>
</inline-formula> denotes the fuzzy set which is the maximum element of the ordering specified by <inline-formula id="ieqn-40">
<mml:math id="mml-ieqn-40"><mml:mi>H</mml:mi></mml:math>
</inline-formula> when <inline-formula id="ieqn-41">
<mml:math id="mml-ieqn-41"><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math>
</inline-formula> &#x003D; 0.5.</p>
<p>The proposed algorithm chooses the feature that exhibits the lowest fuzzy entropy, which gives a strong relationship with the output feature.<list list-type="bullet"><list-item>
<p><bold>Imputation</bold>: After the candidate feature <inline-formula id="ieqn-42">
<mml:math id="mml-ieqn-42"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula>is being selected, the model is fitted using <inline-formula id="ieqn-43">
<mml:math id="mml-ieqn-43"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> as the input features and the candidate feature as target with the cumulative formula described in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>.</p></list-item></list></p>
<p><disp-formula id="eqn-10"><label>(7)</label>
<mml:math id="mml-eqn-10" display="block"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where:</p>
<p><disp-formula id="eqn-11">
<mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>c</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:msubsup><mml:mi>X</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-12">
<mml:math id="mml-eqn-12" display="block"><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03BB;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-13">
<mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-14">
<mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-44">
<mml:math id="mml-ieqn-44"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:math>
</inline-formula> <inline-formula id="ieqn-45">
<mml:math id="mml-ieqn-45"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:math>
</inline-formula> is the number of features holding MVs and <inline-formula id="ieqn-46">
<mml:math id="mml-ieqn-46"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi></mml:mrow></mml:mrow></mml:math>
</inline-formula> is the number of perfect features.<list list-type="bullet"><list-item>
<p><bold>Update datasets:</bold> The selected feature dropped from <inline-formula id="ieqn-47">
<mml:math id="mml-ieqn-47"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> and after imputing MVs within the <inline-formula id="ieqn-48">
<mml:math id="mml-ieqn-48"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula>, the imputed feature <inline-formula id="ieqn-49">
<mml:math id="mml-ieqn-49"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula> is added to <inline-formula id="ieqn-50">
<mml:math id="mml-ieqn-50"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula>. Now <inline-formula id="ieqn-51">
<mml:math id="mml-ieqn-51"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> holds all perfect features, <inline-formula id="ieqn-52">
<mml:math id="mml-ieqn-52"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula> and <inline-formula id="ieqn-53">
<mml:math id="mml-ieqn-53"><mml:mi>y</mml:mi></mml:math>
</inline-formula>. A new <inline-formula id="ieqn-54">
<mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula> from <inline-formula id="ieqn-55">
<mml:math id="mml-ieqn-55"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> is chosen. The model is fitted with the cumulative formula with <inline-formula id="ieqn-56">
<mml:math id="mml-ieqn-56"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> as an input features and the new <inline-formula id="ieqn-57">
<mml:math id="mml-ieqn-57"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula>as the target feature.</p></list-item></list></p>
<p>Repeat from step 2 of feature selection until <inline-formula id="ieqn-58">
<mml:math id="mml-ieqn-58"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> holds no features, at that moment return (<inline-formula id="ieqn-59">
<mml:math id="mml-ieqn-59"><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula>) as the imputed dataset as described in the following algorithm.</p>
<fig id="fig-4">
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_22152-fig-4.png"/>
</fig>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Setup</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>Usually, applying and comparing several imputation algorithms on diverse datasets <italic>versus</italic> the proposed algorithm will result in different imputation performances. Furthermore, this difference in imputation helps in judgement about the compared algorithms and the proposed one and also gives an insight about how the proposed algorithm will perform in future and in different situations. The focus in this paper is on medical datasets. The used datasets in this experiment were obtained from several data repositories and are freely access. <xref ref-type="table" rid="table-1">Tab. 1</xref> gives an overview about the specifications of the datasets used in the experiment. In each dataset, the generation of MVs proportions, 10%, 20%, 30%, 40% and 50%, were performed using the ampute function from the R environment [<xref ref-type="bibr" rid="ref-29">29</xref>] for every missingness mechanism, MAR, MCAR and MNAR.</p>
<table-wrap id="table-1"><label>Table 1</label>
<caption>
<title>The fundamental specifications the used datasets</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Dataset name</th>
<th rowspan="2">#Ins.</th>
<th rowspan="2">#Features</th>
<th rowspan="2">#Class</th>
<th colspan="3">Missingness mechanism</th>
</tr>
<tr>
<th>MAR</th>
<th>MCAR</th>
<th>MNAR</th>
</tr>
</thead>
<tbody>
<tr>
<td>Dermatology [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>366</td>
<td>33</td>
<td>6</td>
<td rowspan="4" colspan="3">10%, 20%, 30%, 40% and 50%</td>
</tr>
<tr>
<td>Breast Cancer [<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>699</td>
<td>10</td>
<td>2</td>
</tr>
<tr>
<td>Parkinsons [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>197</td>
<td>23</td>
<td>2</td>
</tr>
<tr>
<td>Pima Indians Diabetes [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>768</td>
<td>8</td>
<td>2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Five practical imputation algorithms were used in the experiment against the proposed algorithm. <xref ref-type="table" rid="table-2">Tab. 2</xref> describes briefly the compared algorithms used in the experiment.</p>
<table-wrap id="table-2"><label>Table 2</label>
<caption>
<title>The algorithms used in the comparison</title></caption>
<table><colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Package (function name)</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>autoimpute (stochastic) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>imputes MVs using the least squares methodology, then adds to the imputations a stochastic element.</td>
</tr>
<tr>
<td>autoimpute (nocb) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>imputes MVs by carrying next observation moving backward.</td>
</tr>
<tr>
<td>SimpleImputer (median) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>imputes MVs using the median for each feature.</td>
</tr>
<tr>
<td>impyute (EMI) [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>handles MVs using Expectation Maximization Imputation.</td>
</tr>
<tr>
<td>impyute (random) [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>imputes MVs using a randomly selected value from the same feature.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The experiments were conducted using a laptop with the following specification: Windows 10 OS, 4 GB memory, AMD A4-6210 APU with AMD Radeon R3 Graphics (1.80 GHz) processor, 500 GB HDD and Python (version 3.7) programming language and R (version 3.5.2).</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Evaluation Metrics</title>
<p>Imputation performance can be measured using various metrics. This section exhibits an overview of most metrics used in the experimental implementation to measure the imputation performance; these metrics include MAE, RMSE, and <italic>R</italic><sup><italic>2</italic></sup> score.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>MAE and RMSE</title>
<p>MAE is used to calculate the average of the absolute differences between the predicted and true values. It gives an intuition about the magnitudes (absolute values) of the error in prediction, but does not offer any idea about the direction of the prediction (<italic>i.e</italic>., under or over predicting) [<xref ref-type="bibr" rid="ref-36">36</xref>]. RMSE is much like the MAE in that it gives an idea of the magnitude of error. Furthermore, as the variance related to the error magnitudes distribution increases RMSE also increases and MAE is steady. <xref ref-type="disp-formula" rid="eqn-8">Eqs. (8)</xref> and <xref ref-type="disp-formula" rid="eqn-9">(9)</xref> describes MAE and RMSE respectively [<xref ref-type="bibr" rid="ref-3">3</xref>]</p>
<p><disp-formula id="eqn-15"><label>(8)</label>
<mml:math id="mml-eqn-15" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-16"><label>(9)</label>
<mml:math id="mml-eqn-16" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:msqrt></mml:math>
</disp-formula></p>
<p>where the real and predicted values are denoted as <inline-formula id="ieqn-86">
<mml:math id="mml-ieqn-86"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> and <inline-formula id="ieqn-87">
<mml:math id="mml-ieqn-87"><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math>
</inline-formula> of the <inline-formula id="ieqn-88">
<mml:math id="mml-ieqn-88"><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:math>
</inline-formula> observation respectively and the number of the observations is denoted as <inline-formula id="ieqn-89">
<mml:math id="mml-ieqn-89"><mml:mi>n</mml:mi></mml:math>
</inline-formula>.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>R<sup>2</sup> Score</title>
<p><italic>R</italic><sup><italic>2</italic></sup> score, given by <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>, gives an indication of the prediction&#x2019;s goodness of fit to the true values. From a statistical perspective, the <italic>R</italic><sup><italic>2</italic></sup> score has been dubbed as the coefficient of determination [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p><disp-formula id="eqn-17"><label>(10)</label>
<mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math>
</disp-formula></p>
<p>where:</p>
<p><disp-formula id="eqn-18">
<mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math>
</disp-formula></p>
<p><inline-formula id="ieqn-90">
<mml:math id="mml-ieqn-90"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:math>
</inline-formula> represents the mean of the detected data.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Results and Discussion</title>
<p><xref ref-type="fig" rid="fig-1">Figs. 1</xref> to <xref ref-type="fig" rid="fig-3">3</xref> present the improvement in performance, using RMSE, MAE and <italic>R</italic><sup><italic>2</italic></sup> score, of the compared algorithms <italic>versus</italic> the proposed one. The performance evaluation of the proposed algorithm against the compared algorithms for each MVs percentage, 10%, 20%, 30%, 40% and 50%, generated from the missingness mechanisms, MAR, MCAR and MNAR, is presented in more details in <xref ref-type="table" rid="table-3">Tabs. 3</xref> to <xref ref-type="table" rid="table-6">6</xref>. The results exhibit that the performance differs from one algorithm to another depending on the dimension of the dataset, the missingness mechanism type, and the amount of MVs in the dataset. The computational complexity of both CBRL and CBRC is <italic>O(n)</italic>.</p>
<p>This section is subdivided into two subsections. The first section explains the accuracy analysis evaluated using <italic>R</italic><sup><italic>2</italic></sup> score (higher values is better) and the second represents the error analysis evaluated using RMSE and MAE metrics (lower value is better).</p>
<sec id="s5_1">
<label>5.1</label>
<title>Accuracy Analysis</title>
<p>This subsection exhibits that the proposed algorithm offers better accuracy <italic>versus</italic> the compared algorithms in many cases. The accuracy analysis is represented by calculating <italic>R</italic><sup><italic>2</italic></sup> score. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> exhibits the improvement percentage of <italic>R</italic><sup><italic>2</italic></sup> score which is given by <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> for the proposed algorithm <italic>versus</italic> the compared algorithms. In what follows, the comparison of <italic>R</italic><sup><italic>2</italic></sup> score is discussed in detail. In all missigness mechanisms, <italic>R</italic><sup><italic>2</italic></sup> score of the proposed algorithm is better than nocb, median, EMI and random when applied on all datasets used in the experiment. In addition, <italic>R</italic><sup><italic>2</italic></sup> score of the proposed algorithm is better than stochastic when applied on all used datasets but worse than stochastic when applied on parkinsons dataset in all missigness mechanisms.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Error Analysis</title>
<p>This subsection exhibits that the proposed algorithm gives lower error <italic>versus</italic> the compared algorithms in many cases. MAE and RMSE, given by <xref ref-type="disp-formula" rid="eqn-8">Eqs. (8)</xref> and <xref ref-type="disp-formula" rid="eqn-9">(9)</xref> respectively, are the metrics used for assessing the error in imputation. <xref ref-type="fig" rid="fig-2">Figs. 2</xref> and <xref ref-type="fig" rid="fig-3">3</xref> exhibit the improvement percentage in evaluating both MAE and RMSE respectively.</p>
<p>In all missigness mechanisms, MAE given by the proposed algorithm is lower than MAE given by EMI and random when applied on all datasets used in the experiment. When the proposed algorithm is compared with stochastic in MAR and MCAR, it was observed that MAE granted by the proposed algorithm is the lowest in all used datasets. In MNAR, MAE of the proposed algorithm is better than stochastic in all used datasets except when applied on the parkinsons dataset. In MAR and MCAR, MAE of the proposed algorithm is better than nocb in all used datasets except when applied on the breast cancer dataset. In MNAR, MAE of the proposed algorithm is better than nocb in all used datasets except when applied on the breast cancer and parkinsons datasets. When the proposed algorithm is compared with median in MCAR it was observed that MAE of the proposed algorithm is better in all used datasets. In MAR, MAE of the proposed algorithm is better than median in all used datasets except when applied on the parkinsons dataset. In MNAR, MAE of the proposed algorithm is better than median in all used datasets except when applied on the breast cancer dataset.</p>
<p>In MAR and MCAR, RMSE of the proposed algorithm is better than stochastic in all used datasets. Also in MNAR, RMSE of the proposed algorithm is better than stochastic in all used datasets except when applied on the parkinsons dataset. When the proposed algorithm is compared with nocb in MAR and MNAR, it was noticed that RMSE of the proposed algorithm is better than nocb in all used datasets except when applied on the parkinsons and breast cancer datasets. When the proposed algorithm is compared with median, EMI and random in MAR and MNAR, it was observed that RMSE of the proposed algorithm is better in all used datasets. Also in MCAR, RMSE of the proposed algorithm is better than RMSE of median, EMI and random when applied in all datasets used in the experiment except in Pima Indians Diabetes and breast cancer datasets.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title><italic>R</italic><sup><italic>2</italic></sup> score improvement percentage of the proposed algorithm <italic>versus</italic> the compared algorithms</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_22152-fig-1.png"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>MAE improvement percentage of the proposed algorithm <italic>versus</italic> the compared algorithms</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_22152-fig-2.png"/>
</fig>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>RMSE improvement percentage of the proposed algorithm <italic>versus</italic> the compared algorithms</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_22152-fig-3.png"/>
</fig>
<table-wrap id="table-3"><label>Table 3</label>
<caption>
<title>MAE, RMSE and R<sup>2</sup> score of the proposed algorithm against the compared algorithms (<bold>breast cancer</bold> dataset)</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<tbody><tr>
<td colspan="14"><bold>Breast cancer</bold></td>
</tr>
<tr>
<td rowspan="2"><bold>Mech</bold></td>
<td rowspan="2"><bold>% missing value</bold></td>
<td colspan="2"><bold>Proposed algorithm</bold></td>
<td colspan="2"><bold>Stochastic</bold></td>
<td colspan="2"><bold>NOCB</bold></td>
<td colspan="2"><bold>Median</bold></td>
<td colspan="2"><bold>EMI</bold></td>
<td colspan="2"><bold>Random</bold></td>
</tr><tr>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td>320.64</td>
<td>14398.76</td>
<td>407.99</td>
<td>16689.91</td>
<td>17.316</td>
<td>1215.105</td>
<td>373.904</td>
<td>16117.67</td>
<td>373.9</td>
<td>19940.49</td>
<td>373.90</td>
<td>16754.00</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>443.09</td>
<td>14370.89</td>
<td>678.96</td>
<td>21339.27</td>
<td>235.09</td>
<td>12320.38</td>
<td>447.678</td>
<td>16186.82</td>
<td>447.68</td>
<td>41898.87</td>
<td>447.68</td>
<td>88603.04</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>394.73</td>
<td>12750.13</td>
<td>973.25</td>
<td>25093.19</td>
<td>31.25</td>
<td>1475.361</td>
<td>421.783</td>
<td>14278.08</td>
<td>421.783</td>
<td>22493.84</td>
<td>421.78</td>
<td>15263.37</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>983.17</td>
<td>21635.06</td>
<td>2532.58</td>
<td>50233.29</td>
<td>404.21</td>
<td>17943.89</td>
<td>901.744</td>
<td>23224.22</td>
<td>901.744</td>
<td>49081.14</td>
<td>901.74</td>
<td>22287.51</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>834.06</td>
<td>17070.85</td>
<td>1792.33</td>
<td>33483.05</td>
<td>280.10</td>
<td>15684.06</td>
<td>670.810</td>
<td>17701.69</td>
<td>670.810</td>
<td>40735.57</td>
<td>670.81</td>
<td>27890.12</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>595.14</bold></td>
<td><bold>16045.14</bold></td>
<td><bold>1277.02</bold></td>
<td><bold>29367.74</bold></td>
<td><bold>193.59</bold></td>
<td><bold>9727.76</bold></td>
<td><bold>563.18</bold></td>
<td><bold>17501.69</bold></td>
<td><bold>563.18</bold></td>
<td><bold>34829.98</bold></td>
<td><bold>563.18</bold></td>
<td><bold>34159.61</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.53</bold></td>
<td><bold>0.45</bold></td>
<td><bold>&#x2013;2.07</bold></td>
<td><bold>&#x2013;0.65</bold></td>
<td><bold>&#x2013;0.06</bold></td>
<td><bold>0.08</bold></td>
<td><bold>&#x2013;0.06</bold></td>
<td><bold>0.54</bold></td>
<td><bold>&#x2013;0.06</bold></td>
<td><bold>0.53</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td>522.92</td>
<td>15759.95</td>
<td>1335.46</td>
<td>40018.23</td>
<td>197.18</td>
<td>14155.30</td>
<td>509.28</td>
<td>17330.07</td>
<td>509.28</td>
<td>37906.66</td>
<td>509.28</td>
<td>17228.65</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>541.59</td>
<td>11785.89</td>
<td>1335.61</td>
<td>32870.79</td>
<td>256.52</td>
<td>14398.2</td>
<td>378.15</td>
<td>10291.57</td>
<td>378.15</td>
<td>31369.78</td>
<td>378.15</td>
<td>27283.8</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>791.10</td>
<td>18334.46</td>
<td>1979.26</td>
<td>46012.12</td>
<td>79.82</td>
<td>3571.15</td>
<td>808.66</td>
<td>20314.66</td>
<td>808.65</td>
<td>50625.01</td>
<td>808.66</td>
<td>18119.15</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>1039.62</td>
<td>19614.44</td>
<td>1531.01</td>
<td>27097.39</td>
<td>354.61</td>
<td>17217.28</td>
<td>905.83</td>
<td>20613.91</td>
<td>905.83</td>
<td>55373.14</td>
<td>905.83</td>
<td>31896.57</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>1296.27</td>
<td>24990.94</td>
<td>3480.9</td>
<td>57980.1</td>
<td>255.97</td>
<td>11428.92</td>
<td>1289.99</td>
<td>27442.8</td>
<td>1289.99</td>
<td>50641.55</td>
<td>1289.98</td>
<td>33227.33</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>838.3</bold></td>
<td><bold>18097.14</bold></td>
<td><bold>1932.45</bold></td>
<td><bold>40795.91</bold></td>
<td><bold>228.82</bold></td>
<td><bold>12154.17</bold></td>
<td><bold>778.38</bold></td>
<td><bold>19198.60</bold></td>
<td><bold>778.38</bold></td>
<td><bold>45183.23</bold></td>
<td><bold>778.38</bold></td>
<td><bold>25551.1</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.57</bold></td>
<td><bold>0.56</bold></td>
<td><bold>&#x2013;2.66</bold></td>
<td><bold>&#x2013;0.49</bold></td>
<td><bold>&#x2013;0.08</bold></td>
<td><bold>0.06</bold></td>
<td><bold>&#x2013;0.08</bold></td>
<td><bold>0.6</bold></td>
<td><bold>&#x2013;0.08</bold></td>
<td><bold>0.292</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td>105.56</td>
<td>4262</td>
<td>490.45</td>
<td>20762.39</td>
<td>4.33</td>
<td>218.99</td>
<td>50.40</td>
<td>2157.73</td>
<td>50.40</td>
<td>18587.69</td>
<td>50.40</td>
<td>11489.77</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>264.5</td>
<td>7141.54</td>
<td>652.4</td>
<td>20256.7</td>
<td>8.017</td>
<td>284.40</td>
<td>129.97</td>
<td>3845.5</td>
<td>129.97</td>
<td>28673.89</td>
<td>129.97</td>
<td>11829</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>481.06</td>
<td>9917.79</td>
<td>1785.58</td>
<td>43896.77</td>
<td>19.02</td>
<td>699.58</td>
<td>354.49</td>
<td>7929.36</td>
<td>354.49</td>
<td>32101.63</td>
<td>354.49</td>
<td>16315.06</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>735.63</td>
<td>13379.96</td>
<td>2861.79</td>
<td>53352.1</td>
<td>29.28</td>
<td>1005.46</td>
<td>603.55</td>
<td>11938.82</td>
<td>603.55</td>
<td>38242.58</td>
<td>603.55</td>
<td>87175.66</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>812.28</td>
<td>13960.04</td>
<td>2837.58</td>
<td>54955.41</td>
<td>154.06</td>
<td>9768.59</td>
<td>647.49</td>
<td>13497.92</td>
<td>647.49</td>
<td>40577.24</td>
<td>647.5</td>
<td>18928.93</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>479.81</bold></td>
<td><bold>9732.27</bold></td>
<td><bold>1725.56</bold></td>
<td><bold>38644.85</bold></td>
<td><bold>42.94</bold></td>
<td><bold>2395.40</bold></td>
<td><bold>357.18</bold></td>
<td><bold>7873.87</bold></td>
<td><bold>357.18</bold></td>
<td><bold>31636.61</bold></td>
<td><bold>357.18</bold></td>
<td><bold>29147.68</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.72</bold></td>
<td><bold>0.75</bold></td>
<td><bold>&#x2013;10.17</bold></td>
<td><bold>&#x2013;3.063</bold></td>
<td><bold>&#x2013;0.34</bold></td>
<td><bold>&#x2013;0.24</bold></td>
<td><bold>&#x2013;0.34</bold></td>
<td><bold>0.69</bold></td>
<td><bold>&#x2013;0.34</bold></td>
<td><bold>0.67</bold></td>
</tr><tr>
<td><bold>Mech</bold></td>
<td><bold>% missing value</bold></td>
<td colspan="12"><bold>R</bold><sup><bold>2</bold></sup><bold>score</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.989</td>
<td colspan="2">0.985</td>
<td colspan="2">0.969</td>
<td colspan="2">0.965</td>
<td colspan="2">0.955</td>
<td colspan="2">0.974</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.984</td>
<td colspan="2">0.980</td>
<td colspan="2">0.966</td>
<td colspan="2">0.947</td>
<td colspan="2">0.933</td>
<td colspan="2">0.923</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.977</td>
<td colspan="2">0.971</td>
<td colspan="2">0.926</td>
<td colspan="2">0.923</td>
<td colspan="2">0.907</td>
<td colspan="2">0.913</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.978</td>
<td colspan="2">0.962</td>
<td colspan="2">0.903</td>
<td colspan="2">0.912</td>
<td colspan="2">0.902</td>
<td colspan="2">0.907</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.968</td>
<td colspan="2">0.953</td>
<td colspan="2">0.899</td>
<td colspan="2">0.898</td>
<td colspan="2">0.886</td>
<td colspan="2">0.859</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.979</bold></td>
<td colspan="2"><bold>0.970</bold></td>
<td colspan="2"><bold>0.933</bold></td>
<td colspan="2"><bold>0.929</bold></td>
<td colspan="2"><bold>0.917</bold></td>
<td colspan="2"><bold>0.915</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.009</bold></td>
<td colspan="2"><bold>0.050</bold></td>
<td colspan="2"><bold>0.054</bold></td>
<td colspan="2"><bold>0.068</bold></td>
<td colspan="2"><bold>0.070</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.995</td>
<td colspan="2">0.989</td>
<td colspan="2">0.979</td>
<td colspan="2">0.987</td>
<td colspan="2">0.980</td>
<td colspan="2">0.962</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.994</td>
<td colspan="2">0.985</td>
<td colspan="2">0.970</td>
<td colspan="2">0.978</td>
<td colspan="2">0.965</td>
<td colspan="2">0.953</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.987</td>
<td colspan="2">0.972</td>
<td colspan="2">0.941</td>
<td colspan="2">0.958</td>
<td colspan="2">0.943</td>
<td colspan="2">0.902</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.991</td>
<td colspan="2">0.979</td>
<td colspan="2">0.947</td>
<td colspan="2">0.960</td>
<td colspan="2">0.936</td>
<td colspan="2">0.899</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.983</td>
<td colspan="2">0.956</td>
<td colspan="2">0.926</td>
<td colspan="2">0.932</td>
<td colspan="2">0.906</td>
<td colspan="2">0.859</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.990</bold></td>
<td colspan="2"><bold>0.976</bold></td>
<td colspan="2"><bold>0.953</bold></td>
<td colspan="2"><bold>0.963</bold></td>
<td colspan="2"><bold>0.946</bold></td>
<td colspan="2"><bold>0.915</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.014</bold></td>
<td colspan="2"><bold>0.039</bold></td>
<td colspan="2"><bold>0.028</bold></td>
<td colspan="2"><bold>0.046</bold></td>
<td colspan="2"><bold>0.082</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.990</td>
<td colspan="2">0.985</td>
<td colspan="2">0.966</td>
<td colspan="2">0.953</td>
<td colspan="2">0.952</td>
<td colspan="2">0.973</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.976</td>
<td colspan="2">0.971</td>
<td colspan="2">0.932</td>
<td colspan="2">0.925</td>
<td colspan="2">0.917</td>
<td colspan="2">0.947</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.970</td>
<td colspan="2">0.958</td>
<td colspan="2">0.917</td>
<td colspan="2">0.901</td>
<td colspan="2">0.898</td>
<td colspan="2">0.909</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.973</td>
<td colspan="2">0.953</td>
<td colspan="2">0.913</td>
<td colspan="2">0.903</td>
<td colspan="2">0.886</td>
<td colspan="2">0.884</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.972</td>
<td colspan="2">0.959</td>
<td colspan="2">0.889</td>
<td colspan="2">0.882</td>
<td colspan="2">0.870</td>
<td colspan="2">0.860</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.976</bold></td>
<td colspan="2"><bold>0.965</bold></td>
<td colspan="2"><bold>0.924</bold></td>
<td colspan="2"><bold>0.913</bold></td>
<td colspan="2"><bold>0.904</bold></td>
<td colspan="2"><bold>0.915</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.012</bold></td>
<td colspan="2"><bold>0.057</bold></td>
<td colspan="2"><bold>0.069</bold></td>
<td colspan="2"><bold>0.079</bold></td>
<td colspan="2"><bold>0.067</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-4"><label>Table 4</label>
<caption>
<title>MAE, RMSE and R<sup>2</sup> score of the proposed algorithm against the compared algorithms (dermatology dataset)</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="14"><bold>Dermatology</bold></th>
</tr>
<tr>
<td rowspan="2"><bold>Mech</bold></td>
<td rowspan="2"><bold>% missing value</bold></td>
<td colspan="2"><bold>Proposed algorithm</bold></td>
<td colspan="2"><bold>Stochastic</bold></td>
<td colspan="2"><bold>NOCB</bold></td>
<td colspan="2"><bold>Median</bold></td>
<td colspan="2"><bold>EMI</bold></td>
<td colspan="2"><bold>Random</bold></td>
</tr>
<tr>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td>0.265</td>
<td>0.005</td>
<td>0.223</td>
<td>0.004</td>
<td>0.196</td>
<td>0.004</td>
<td>0.247</td>
<td>0.005</td>
<td>0.269</td>
<td>0.005</td>
<td>0.247</td>
<td>0.005</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>0.166</td>
<td>0.003</td>
<td>0.371</td>
<td>0.006</td>
<td>0.320</td>
<td>0.006</td>
<td>0.208</td>
<td>0.004</td>
<td>0.490</td>
<td>0.004</td>
<td>0.218</td>
<td>0.004</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>0.237</td>
<td>0.006</td>
<td>0.499</td>
<td>0.011</td>
<td>0.529</td>
<td>0.013</td>
<td>0.265</td>
<td>0.009</td>
<td>0.198</td>
<td>0.009</td>
<td>0.443</td>
<td>0.009</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.194</td>
<td>0.006</td>
<td>0.252</td>
<td>0.009</td>
<td>0.287</td>
<td>0.012</td>
<td>0.250</td>
<td>0.009</td>
<td>0.350</td>
<td>0.009</td>
<td>0.194</td>
<td>0.009</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.509</td>
<td>0.015</td>
<td>0.708</td>
<td>0.021</td>
<td>0.623</td>
<td>0.020</td>
<td>0.531</td>
<td>0.017</td>
<td>0.863</td>
<td>0.017</td>
<td>0.639</td>
<td>0.017</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.274</bold></td>
<td><bold>0.007</bold></td>
<td><bold>0.411</bold></td>
<td><bold>0.010</bold></td>
<td><bold>0.391</bold></td>
<td><bold>0.011</bold></td>
<td><bold>0.300</bold></td>
<td><bold>0.009</bold></td>
<td><bold>0.434</bold></td>
<td><bold>0.009</bold></td>
<td><bold>0.348</bold></td>
<td><bold>0.009</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.332</bold></td>
<td><bold>0.302</bold></td>
<td><bold>0.298</bold></td>
<td><bold>0.357</bold></td>
<td><bold>0.085</bold></td>
<td><bold>0.207</bold></td>
<td><bold>0.368</bold></td>
<td><bold>0.207</bold></td>
<td><bold>0.212</bold></td>
<td><bold>0.207</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td>0.325</td>
<td>0.005</td>
<td>0.277</td>
<td>0.005</td>
<td>0.430</td>
<td>0.007</td>
<td>0.331</td>
<td>0.006</td>
<td>0.413</td>
<td>0.006</td>
<td>0.519</td>
<td>0.006</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>0.166</td>
<td>0.004</td>
<td>0.487</td>
<td>0.010</td>
<td>0.366</td>
<td>0.008</td>
<td>0.269</td>
<td>0.007</td>
<td>0.399</td>
<td>0.007</td>
<td>0.311</td>
<td>0.007</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>0.237</td>
<td>0.006</td>
<td>0.319</td>
<td>0.008</td>
<td>0.282</td>
<td>0.009</td>
<td>0.226</td>
<td>0.007</td>
<td>0.342</td>
<td>0.007</td>
<td>0.288</td>
<td>0.007</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.155</td>
<td>0.006</td>
<td>0.138</td>
<td>0.008</td>
<td>0.242</td>
<td>0.010</td>
<td>0.188</td>
<td>0.008</td>
<td>0.434</td>
<td>0.008</td>
<td>0.193</td>
<td>0.008</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.182</td>
<td>0.008</td>
<td>0.265</td>
<td>0.010</td>
<td>0.394</td>
<td>0.014</td>
<td>0.207</td>
<td>0.010</td>
<td>0.221</td>
<td>0.010</td>
<td>0.224</td>
<td>0.010</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.213</bold></td>
<td><bold>0.006</bold></td>
<td><bold>0.297</bold></td>
<td><bold>0.008</bold></td>
<td><bold>0.343</bold></td>
<td><bold>0.010</bold></td>
<td><bold>0.244</bold></td>
<td><bold>0.008</bold></td>
<td><bold>0.362</bold></td>
<td><bold>0.008</bold></td>
<td><bold>0.307</bold></td>
<td><bold>0.008</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.284</bold></td>
<td><bold>0.277</bold></td>
<td><bold>0.379</bold></td>
<td><bold>0.368</bold></td>
<td><bold>0.129</bold></td>
<td><bold>0.218</bold></td>
<td><bold>0.412</bold></td>
<td><bold>0.218</bold></td>
<td><bold>0.306</bold></td>
<td><bold>0.218</bold></td>
</tr>
<tr>
<td rowspan="6"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td>0.158</td>
<td>0.004</td>
<td>0.173</td>
<td>0.005</td>
<td>0.236</td>
<td>0.006</td>
<td>0.180</td>
<td>0.005</td>
<td>0.310</td>
<td>0.005</td>
<td>0.331</td>
<td>0.005</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>0.373</td>
<td>0.008</td>
<td>0.573</td>
<td>0.012</td>
<td>0.465</td>
<td>0.013</td>
<td>0.386</td>
<td>0.012</td>
<td>0.423</td>
<td>0.012</td>
<td>0.604</td>
<td>0.012</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>0.271</td>
<td>0.008</td>
<td>0.349</td>
<td>0.010</td>
<td>0.412</td>
<td>0.011</td>
<td>0.329</td>
<td>0.012</td>
<td>0.409</td>
<td>0.012</td>
<td>0.335</td>
<td>0.012</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.325</td>
<td>0.012</td>
<td>0.327</td>
<td>0.012</td>
<td>0.538</td>
<td>0.019</td>
<td>0.350</td>
<td>0.017</td>
<td>0.304</td>
<td>0.017</td>
<td>0.548</td>
<td>0.017</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.372</td>
<td>0.014</td>
<td>0.303</td>
<td>0.013</td>
<td>0.367</td>
<td>0.016</td>
<td>0.456</td>
<td>0.019</td>
<td>0.923</td>
<td>0.019</td>
<td>0.485</td>
<td>0.019</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.300</bold></td>
<td><bold>0.009</bold></td>
<td><bold>0.345</bold></td>
<td><bold>0.010</bold></td>
<td><bold>0.403</bold></td>
<td><bold>0.013</bold></td>
<td><bold>0.340</bold></td>
<td><bold>0.013</bold></td>
<td><bold>0.474</bold></td>
<td><bold>0.013</bold></td>
<td><bold>0.461</bold></td>
<td><bold>0.013</bold></td>
</tr><tr>
<td></td>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.131</bold></td>
<td><bold>0.119</bold></td>
<td><bold>0.257</bold></td>
<td><bold>0.307</bold></td>
<td><bold>0.119</bold></td>
<td><bold>0.296</bold></td>
<td><bold>0.367</bold></td>
<td><bold>0.296</bold></td>
<td><bold>0.349</bold></td>
<td><bold>0.296</bold></td>
</tr><tr>
<td rowspan="1"><bold>Mech</bold></td>
<td rowspan="1"><bold>% missing value</bold></td>
<td colspan="12"><bold>R</bold><sup><bold>2</bold></sup><bold>score</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.999</td>
<td colspan="2">0.998</td>
<td colspan="2">0.995</td>
<td colspan="2">0.998</td>
<td colspan="2">0.997</td>
<td colspan="2">0.993</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.998</td>
<td colspan="2">0.996</td>
<td colspan="2">0.993</td>
<td colspan="2">0.994</td>
<td colspan="2">0.988</td>
<td colspan="2">0.981</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.995</td>
<td colspan="2">0.992</td>
<td colspan="2">0.985</td>
<td colspan="2">0.985</td>
<td colspan="2">0.981</td>
<td colspan="2">0.965</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.993</td>
<td colspan="2">0.987</td>
<td colspan="2">0.974</td>
<td colspan="2">0.982</td>
<td colspan="2">0.977</td>
<td colspan="2">0.952</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.992</td>
<td colspan="2">0.987</td>
<td colspan="2">0.978</td>
<td colspan="2">0.979</td>
<td colspan="2">0.971</td>
<td colspan="2">0.937</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.995</bold></td>
<td colspan="2"><bold>0.992</bold></td>
<td colspan="2"><bold>0.985</bold></td>
<td colspan="2"><bold>0.988</bold></td>
<td colspan="2"><bold>0.983</bold></td>
<td colspan="2"><bold>0.966</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.003</bold></td>
<td colspan="2"><bold>0.010</bold></td>
<td colspan="2"><bold>0.008</bold></td>
<td colspan="2"><bold>0.013</bold></td>
<td colspan="2"><bold>0.031</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.998</td>
<td colspan="2">0.997</td>
<td colspan="2">0.994</td>
<td colspan="2">0.994</td>
<td colspan="2">0.992</td>
<td colspan="2">0.983</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.998</td>
<td colspan="2">0.995</td>
<td colspan="2">0.992</td>
<td colspan="2">0.992</td>
<td colspan="2">0.989</td>
<td colspan="2">0.972</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.995</td>
<td colspan="2">0.993</td>
<td colspan="2">0.984</td>
<td colspan="2">0.988</td>
<td colspan="2">0.980</td>
<td colspan="2">0.967</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.993</td>
<td colspan="2">0.991</td>
<td colspan="2">0.982</td>
<td colspan="2">0.981</td>
<td colspan="2">0.977</td>
<td colspan="2">0.944</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.991</td>
<td colspan="2">0.989</td>
<td colspan="2">0.971</td>
<td colspan="2">0.979</td>
<td colspan="2">0.972</td>
<td colspan="2">0.926</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.995</bold></td>
<td colspan="2"><bold>0.993</bold></td>
<td colspan="2"><bold>0.985</bold></td>
<td colspan="2"><bold>0.987</bold></td>
<td colspan="2"><bold>0.982</bold></td>
<td colspan="2"><bold>0.958</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.002</bold></td>
<td colspan="2"><bold>0.010</bold></td>
<td colspan="2"><bold>0.008</bold></td>
<td colspan="2"><bold>0.013</bold></td>
<td colspan="2"><bold>0.038</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.996</td>
<td colspan="2">0.997</td>
<td colspan="2">0.987</td>
<td colspan="2">0.988</td>
<td colspan="2">0.986</td>
<td colspan="2">0.985</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.993</td>
<td colspan="2">0.991</td>
<td colspan="2">0.984</td>
<td colspan="2">0.980</td>
<td colspan="2">0.978</td>
<td colspan="2">0.971</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.991</td>
<td colspan="2">0.989</td>
<td colspan="2">0.980</td>
<td colspan="2">0.978</td>
<td colspan="2">0.976</td>
<td colspan="2">0.967</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.988</td>
<td colspan="2">0.989</td>
<td colspan="2">0.974</td>
<td colspan="2">0.968</td>
<td colspan="2">0.962</td>
<td colspan="2">0.954</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.991</td>
<td colspan="2">0.987</td>
<td colspan="2">0.975</td>
<td colspan="2">0.974</td>
<td colspan="2">0.966</td>
<td colspan="2">0.950</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.992</bold></td>
<td colspan="2"><bold>0.991</bold></td>
<td colspan="2"><bold>0.980</bold></td>
<td colspan="2"><bold>0.977</bold></td>
<td colspan="2"><bold>0.974</bold></td>
<td colspan="2"><bold>0.965</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.001</bold></td>
<td colspan="2"><bold>0.012</bold></td>
<td colspan="2"><bold>0.015</bold></td>
<td colspan="2"><bold>0.019</bold></td>
<td colspan="2"><bold>0.027</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-5"><label>Table 5</label>
<caption>
<title>MAE, RMSE and R<sup>2</sup> score of the proposed algorithm against the compared algorithms (parkinsons dataset)</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="14"><bold>Parkinsons</bold></th>
</tr>
<tr>
<td rowspan="2"><bold>Mech</bold></td>
<td rowspan="2"><bold>% missing value</bold></td>
<td colspan="2"><bold>Proposed algorithm</bold></td>
<td colspan="2"><bold>Stochastic</bold></td>
<td colspan="2"><bold>NOCB</bold></td>
<td colspan="2"><bold>Median</bold></td>
<td colspan="2"><bold>EMI</bold></td>
<td colspan="2"><bold>Random</bold></td>
</tr>
<tr>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td>1.267</td>
<td>0.035</td>
<td>2.070</td>
<td>0.043</td>
<td>0.662</td>
<td>0.014</td>
<td>0.831</td>
<td>0.021</td>
<td>1.828</td>
<td>0.021</td>
<td>1.596</td>
<td>0.021</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>1.774</td>
<td>0.064</td>
<td>1.972</td>
<td>0.057</td>
<td>6.857</td>
<td>0.137</td>
<td>1.884</td>
<td>0.070</td>
<td>1.927</td>
<td>0.070</td>
<td>2.103</td>
<td>0.070</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>1.974</td>
<td>0.072</td>
<td>2.305</td>
<td>0.084</td>
<td>0.385</td>
<td>0.016</td>
<td>1.934</td>
<td>0.069</td>
<td>4.167</td>
<td>0.069</td>
<td>2.300</td>
<td>0.069</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>2.656</td>
<td>0.115</td>
<td>6.924</td>
<td>0.278</td>
<td>1.850</td>
<td>0.066</td>
<td>2.178</td>
<td>0.103</td>
<td>3.131</td>
<td>0.103</td>
<td>7.423</td>
<td>0.103</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>2.110</td>
<td>0.099</td>
<td>2.868</td>
<td>0.140</td>
<td>2.425</td>
<td>0.064</td>
<td>2.904</td>
<td>0.140</td>
<td>3.492</td>
<td>0.140</td>
<td>3.516</td>
<td>0.140</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>1.956</bold></td>
<td><bold>0.077</bold></td>
<td><bold>3.228</bold></td>
<td><bold>0.120</bold></td>
<td><bold>2.436</bold></td>
<td><bold>0.059</bold></td>
<td><bold>1.946</bold></td>
<td><bold>0.081</bold></td>
<td><bold>2.909</bold></td>
<td><bold>0.081</bold></td>
<td><bold>3.388</bold></td>
<td><bold>0.081</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.394</bold></td>
<td><bold>0.361</bold></td>
<td><bold>0.197</bold></td>
<td><bold>&#x2013;0.297</bold></td>
<td><bold>&#x2013;0.005</bold></td>
<td><bold>0.046</bold></td>
<td><bold>0.328</bold></td>
<td><bold>0.046</bold></td>
<td><bold>0.423</bold></td>
<td><bold>0.046</bold></td>
</tr>
<tr>
<td rowspan="6"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td>1.085</td>
<td>0.024</td>
<td>1.734</td>
<td>0.042</td>
<td>0.223</td>
<td>0.005</td>
<td>0.951</td>
<td>0.025</td>
<td>2.464</td>
<td>0.025</td>
<td>2.018</td>
<td>0.025</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>5.895</td>
<td>0.109</td>
<td>5.047</td>
<td>0.113</td>
<td>6.452</td>
<td>0.102</td>
<td>6.323</td>
<td>0.129</td>
<td>7.252</td>
<td>0.129</td>
<td>5.956</td>
<td>0.129</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>1.493</td>
<td>0.056</td>
<td>3.553</td>
<td>0.127</td>
<td>0.904</td>
<td>0.022</td>
<td>1.766</td>
<td>0.076</td>
<td>2.639</td>
<td>0.076</td>
<td>6.722</td>
<td>0.076</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>1.794</td>
<td>0.071</td>
<td>1.449</td>
<td>0.048</td>
<td>2.606</td>
<td>0.070</td>
<td>2.071</td>
<td>0.079</td>
<td>2.628</td>
<td>0.079</td>
<td>2.130</td>
<td>0.079</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>7.332</td>
<td>0.240</td>
<td>7.272</td>
<td>0.268</td>
<td>9.779</td>
<td>0.306</td>
<td>7.027</td>
<td>0.254</td>
<td>8.483</td>
<td>0.254</td>
<td>8.404</td>
<td>0.254</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>3.520</bold></td>
<td><bold>0.100</bold></td>
<td><bold>3.811</bold></td>
<td><bold>0.119</bold></td>
<td><bold>3.993</bold></td>
<td><bold>0.101</bold></td>
<td><bold>3.628</bold></td>
<td><bold>0.113</bold></td>
<td><bold>4.693</bold></td>
<td><bold>0.113</bold></td>
<td><bold>5.046</bold></td>
<td><bold>0.113</bold></td>
</tr>
<tr>
<td></td>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.076</bold></td>
<td><bold>0.163</bold></td>
<td><bold>0.119</bold></td>
<td><bold>0.011</bold></td>
<td><bold>0.030</bold></td>
<td><bold>0.111</bold></td>
<td><bold>0.250</bold></td>
<td><bold>0.111</bold></td>
<td><bold>0.302</bold></td>
<td><bold>0.111</bold></td>
</tr>
<tr>
<td rowspan="6"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td>1.060</td>
<td>0.022</td>
<td>0.766</td>
<td>0.013</td>
<td>1.164</td>
<td>0.018</td>
<td>1.339</td>
<td>0.025</td>
<td>1.629</td>
<td>0.025</td>
<td>1.722</td>
<td>0.025</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>1.042</td>
<td>0.027</td>
<td>1.801</td>
<td>0.047</td>
<td>0.443</td>
<td>0.012</td>
<td>2.130</td>
<td>0.061</td>
<td>1.701</td>
<td>0.061</td>
<td>3.160</td>
<td>0.061</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>2.922</td>
<td>0.099</td>
<td>2.797</td>
<td>0.091</td>
<td>5.653</td>
<td>0.140</td>
<td>3.725</td>
<td>0.143</td>
<td>5.184</td>
<td>0.143</td>
<td>4.305</td>
<td>0.143</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>4.320</td>
<td>0.165</td>
<td>4.111</td>
<td>0.142</td>
<td>2.008</td>
<td>0.065</td>
<td>4.957</td>
<td>0.180</td>
<td>5.823</td>
<td>0.180</td>
<td>6.201</td>
<td>0.180</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>1.885</td>
<td>0.088</td>
<td>1.692</td>
<td>0.077</td>
<td>1.601</td>
<td>0.055</td>
<td>3.837</td>
<td>0.186</td>
<td>3.676</td>
<td>0.186</td>
<td>3.674</td>
<td>0.186</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>2.246</bold></td>
<td><bold>0.080</bold></td>
<td><bold>2.234</bold></td>
<td><bold>0.074</bold></td>
<td><bold>2.174</bold></td>
<td><bold>0.058</bold></td>
<td><bold>3.197</bold></td>
<td><bold>0.119</bold></td>
<td><bold>3.602</bold></td>
<td><bold>0.119</bold></td>
<td><bold>3.813</bold></td>
<td><bold>0.119</bold></td>
</tr><tr>
<td></td>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>&#x2013;0.005</bold></td>
<td><bold>&#x2013;0.085</bold></td>
<td><bold>&#x2013;0.033</bold></td>
<td><bold>&#x2013;0.381</bold></td>
<td><bold>0.298</bold></td>
<td><bold>0.325</bold></td>
<td><bold>0.377</bold></td>
<td><bold>0.325</bold></td>
<td><bold>0.411</bold></td>
<td><bold>0.325</bold></td>
</tr><tr>
<td rowspan="1"><bold>Mech</bold></td>
<td rowspan="1"><bold>% missing value</bold></td>
<td colspan="12"><bold>R</bold><sup><bold>2</bold></sup><bold>score</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.997</td>
<td colspan="2">0.998</td>
<td colspan="2">0.997</td>
<td colspan="2">0.990</td>
<td colspan="2">0.988</td>
<td colspan="2">0.985</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.992</td>
<td colspan="2">0.997</td>
<td colspan="2">0.976</td>
<td colspan="2">0.970</td>
<td colspan="2">0.966</td>
<td colspan="2">0.970</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.981</td>
<td colspan="2">0.991</td>
<td colspan="2">0.988</td>
<td colspan="2">0.966</td>
<td colspan="2">0.950</td>
<td colspan="2">0.973</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.996</td>
<td colspan="2">0.985</td>
<td colspan="2">0.984</td>
<td colspan="2">0.983</td>
<td colspan="2">0.968</td>
<td colspan="2">0.948</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.992</td>
<td colspan="2">0.993</td>
<td colspan="2">0.984</td>
<td colspan="2">0.956</td>
<td colspan="2">0.949</td>
<td colspan="2">0.931</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.992</bold></td>
<td colspan="2"><bold>0.993</bold></td>
<td colspan="2"><bold>0.986</bold></td>
<td colspan="2"><bold>0.973</bold></td>
<td colspan="2"><bold>0.964</bold></td>
<td colspan="2"><bold>0.961</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>&#x2212;0.001</bold></td>
<td colspan="2"><bold>0.005</bold></td>
<td colspan="2"><bold>0.019</bold></td>
<td colspan="2"><bold>0.028</bold></td>
<td colspan="2"><bold>0.031</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">1.000</td>
<td colspan="2">0.999</td>
<td colspan="2">0.999</td>
<td colspan="2">0.997</td>
<td colspan="2">0.994</td>
<td colspan="2">0.994</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.994</td>
<td colspan="2">0.995</td>
<td colspan="2">0.989</td>
<td colspan="2">0.987</td>
<td colspan="2">0.981</td>
<td colspan="2">0.974</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.996</td>
<td colspan="2">0.997</td>
<td colspan="2">0.990</td>
<td colspan="2">0.984</td>
<td colspan="2">0.974</td>
<td colspan="2">0.973</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.996</td>
<td colspan="2">0.997</td>
<td colspan="2">0.992</td>
<td colspan="2">0.989</td>
<td colspan="2">0.975</td>
<td colspan="2">0.980</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.987</td>
<td colspan="2">0.986</td>
<td colspan="2">0.967</td>
<td colspan="2">0.968</td>
<td colspan="2">0.950</td>
<td colspan="2">0.940</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.994</bold></td>
<td colspan="2"><bold>0.995</bold></td>
<td colspan="2"><bold>0.987</bold></td>
<td colspan="2"><bold>0.985</bold></td>
<td colspan="2"><bold>0.975</bold></td>
<td colspan="2"><bold>0.972</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>&#x2212;0.001</bold></td>
<td colspan="2"><bold>0.007</bold></td>
<td colspan="2"><bold>0.010</bold></td>
<td colspan="2"><bold>0.020</bold></td>
<td colspan="2"><bold>0.023</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.999</td>
<td colspan="2">0.999</td>
<td colspan="2">0.995</td>
<td colspan="2">0.990</td>
<td colspan="2">0.986</td>
<td colspan="2">0.988</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.996</td>
<td colspan="2">0.997</td>
<td colspan="2">0.996</td>
<td colspan="2">0.983</td>
<td colspan="2">0.973</td>
<td colspan="2">0.976</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.995</td>
<td colspan="2">0.996</td>
<td colspan="2">0.988</td>
<td colspan="2">0.974</td>
<td colspan="2">0.960</td>
<td colspan="2">0.968</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.990</td>
<td colspan="2">0.993</td>
<td colspan="2">0.979</td>
<td colspan="2">0.968</td>
<td colspan="2">0.965</td>
<td colspan="2">0.947</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.992</td>
<td colspan="2">0.994</td>
<td colspan="2">0.988</td>
<td colspan="2">0.970</td>
<td colspan="2">0.951</td>
<td colspan="2">0.962</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.995</bold></td>
<td colspan="2"><bold>0.996</bold></td>
<td colspan="2"><bold>0.989</bold></td>
<td colspan="2"><bold>0.977</bold></td>
<td colspan="2"><bold>0.967</bold></td>
<td colspan="2"><bold>0.968</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>&#x2212;0.001</bold></td>
<td colspan="2"><bold>0.005</bold></td>
<td colspan="2"><bold>0.018</bold></td>
<td colspan="2"><bold>0.028</bold></td>
<td colspan="2"><bold>0.027</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-6"><label>Table 6</label>
<caption>
<title>MAE, RMSE and R<sup>2</sup> score of the proposed algorithm against the compared algorithms (Pima Indians Diabetes dataset)</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="14"><bold>Pima Indians Diabetes</bold></th>
</tr>
<tr>
<td rowspan="2"><bold>Mech</bold></td>
<td rowspan="2"><bold>% missing value</bold></td>
<td colspan="2"><bold>Proposed algorithm</bold></td>
<td colspan="2"><bold>Stochastic</bold></td>
<td colspan="2"><bold>NOCB</bold></td>
<td colspan="2"><bold>Median</bold></td>
<td colspan="2"><bold>EMI</bold></td>
<td colspan="2"><bold>Random</bold></td>
</tr>
<tr>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
<td><bold>RMSE</bold></td>
<td><bold>MAE</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td>0.379</td>
<td>9.700</td>
<td>0.332</td>
<td>7.291</td>
<td>0.455</td>
<td>11.393</td>
<td>0.448</td>
<td>11.316</td>
<td>0.448</td>
<td>11.753</td>
<td>0.448</td>
<td>10.818</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>0.510</td>
<td>6.424</td>
<td>0.573</td>
<td>6.635</td>
<td>0.760</td>
<td>9.864</td>
<td>0.511</td>
<td>7.373</td>
<td>0.511</td>
<td>9.472</td>
<td>0.511</td>
<td>10.740</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>0.558</td>
<td>7.566</td>
<td>0.735</td>
<td>9.264</td>
<td>0.746</td>
<td>8.699</td>
<td>0.644</td>
<td>8.909</td>
<td>0.644</td>
<td>9.968</td>
<td>0.644</td>
<td>10.516</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.878</td>
<td>9.244</td>
<td>1.038</td>
<td>9.677</td>
<td>1.343</td>
<td>15.416</td>
<td>0.990</td>
<td>11.254</td>
<td>0.990</td>
<td>13.486</td>
<td>0.990</td>
<td>14.768</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.959</td>
<td>9.039</td>
<td>1.088</td>
<td>9.502</td>
<td>1.640</td>
<td>15.887</td>
<td>1.185</td>
<td>11.715</td>
<td>1.185</td>
<td>12.472</td>
<td>1.185</td>
<td>14.980</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.656</bold></td>
<td><bold>8.395</bold></td>
<td><bold>0.753</bold></td>
<td><bold>8.474</bold></td>
<td><bold>0.989</bold></td>
<td><bold>12.252</bold></td>
<td><bold>0.756</bold></td>
<td><bold>10.113</bold></td>
<td><bold>0.756</bold></td>
<td><bold>11.430</bold></td>
<td><bold>0.756</bold></td>
<td><bold>12.364</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.128</bold></td>
<td><bold>0.009</bold></td>
<td><bold>0.336</bold></td>
<td><bold>0.315</bold></td>
<td><bold>0.131</bold></td>
<td><bold>0.170</bold></td>
<td><bold>0.131</bold></td>
<td><bold>0.266</bold></td>
<td><bold>0.131</bold></td>
<td><bold>0.321</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td>0.849</td>
<td>3.821</td>
<td>0.473</td>
<td>7.773</td>
<td>0.352</td>
<td>5.153</td>
<td>0.266</td>
<td>3.798</td>
<td>0.247</td>
<td>4.384</td>
<td>0.247</td>
<td>11.382</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>0.874</td>
<td>7.494</td>
<td>0.688</td>
<td>9.948</td>
<td>0.670</td>
<td>10.193</td>
<td>0.430</td>
<td>7.684</td>
<td>0.504</td>
<td>11.153</td>
<td>0.504</td>
<td>13.846</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>1.021</td>
<td>10.058</td>
<td>0.953</td>
<td>10.781</td>
<td>1.084</td>
<td>12.833</td>
<td>0.697</td>
<td>12.217</td>
<td>0.947</td>
<td>14.416</td>
<td>0.947</td>
<td>16.033</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.845</td>
<td>7.480</td>
<td>1.012</td>
<td>9.752</td>
<td>1.130</td>
<td>12.879</td>
<td>0.702</td>
<td>7.826</td>
<td>0.702</td>
<td>9.056</td>
<td>0.702</td>
<td>18.766</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.939</td>
<td>8.057</td>
<td>1.531</td>
<td>12.185</td>
<td>1.818</td>
<td>16.443</td>
<td>1.147</td>
<td>9.919</td>
<td>1.147</td>
<td>15.752</td>
<td>1.147</td>
<td>22.155</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.906</bold></td>
<td><bold>7.382</bold></td>
<td><bold>0.932</bold></td>
<td><bold>10.088</bold></td>
<td><bold>1.011</bold></td>
<td><bold>11.500</bold></td>
<td><bold>0.649</bold></td>
<td><bold>8.289</bold></td>
<td><bold>0.709</bold></td>
<td><bold>10.952</bold></td>
<td><bold>0.709</bold></td>
<td><bold>16.436</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.028</bold></td>
<td><bold>0.268</bold></td>
<td><bold>0.104</bold></td>
<td><bold>0.358</bold></td>
<td><bold>&#x2013;0.396</bold></td>
<td><bold>0.109</bold></td>
<td><bold>&#x2013;0.277</bold></td>
<td><bold>0.326</bold></td>
<td><bold>&#x2013;0.277</bold></td>
<td><bold>0.551</bold></td>
</tr>
<tr>
<td rowspan="6"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td>1.070</td>
<td>8.684</td>
<td>0.484</td>
<td>9.689</td>
<td>0.469</td>
<td>9.466</td>
<td>0.502</td>
<td>10.700</td>
<td>0.502</td>
<td>9.781</td>
<td>0.502</td>
<td>9.367</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td>1.112</td>
<td>14.583</td>
<td>0.696</td>
<td>13.146</td>
<td>0.883</td>
<td>15.400</td>
<td>0.790</td>
<td>16.248</td>
<td>0.790</td>
<td>14.527</td>
<td>0.790</td>
<td>13.169</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td>0.86</td>
<td>7.75</td>
<td>1.06</td>
<td>10.36</td>
<td>1.16</td>
<td>11.70</td>
<td>1.06</td>
<td>10.98</td>
<td>1.06</td>
<td>11.56</td>
<td>1.06</td>
<td>18.88</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td>0.84</td>
<td>14.04</td>
<td>1.33</td>
<td>12.99</td>
<td>1.38</td>
<td>16.51</td>
<td>1.32</td>
<td>16.58</td>
<td>1.32</td>
<td>19.15</td>
<td>1.32</td>
<td>16.53</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td>0.84</td>
<td>12.07</td>
<td>1.66</td>
<td>13.74</td>
<td>1.77</td>
<td>15.70</td>
<td>1.66</td>
<td>15.82</td>
<td>1.66</td>
<td>14.25</td>
<td>1.66</td>
<td>18.23</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td><bold>0.94</bold></td>
<td><bold>11.43</bold></td>
<td><bold>1.05</bold></td>
<td><bold>11.99</bold></td>
<td><bold>1.13</bold></td>
<td><bold>13.76</bold></td>
<td><bold>1.07</bold></td>
<td><bold>14.07</bold></td>
<td><bold>1.07</bold></td>
<td><bold>13.85</bold></td>
<td><bold>1.07</bold></td>
<td><bold>15.24</bold></td>
</tr><tr>
<td></td>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td><bold>0.10</bold></td>
<td><bold>0.05</bold></td>
<td><bold>0.17</bold></td>
<td><bold>0.17</bold></td>
<td><bold>0.12</bold></td>
<td><bold>0.19</bold></td>
<td><bold>0.12</bold></td>
<td><bold>0.18</bold></td>
<td><bold>0.12</bold></td>
<td><bold>0.25</bold></td>
</tr><tr>
<td rowspan="1"><bold>Mech</bold></td>
<td rowspan="1"><bold>% missing value</bold></td>
<td colspan="12"><bold>R</bold><sup><bold>2</bold></sup><bold>score</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.98</td>
<td colspan="2">0.98</td>
<td colspan="2">0.97</td>
<td colspan="2">0.98</td>
<td colspan="2">0.96</td>
<td colspan="2">0.96</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.98</td>
<td colspan="2">0.97</td>
<td colspan="2">0.95</td>
<td colspan="2">0.97</td>
<td colspan="2">0.96</td>
<td colspan="2">0.94</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.97</td>
<td colspan="2">0.94</td>
<td colspan="2">0.93</td>
<td colspan="2">0.96</td>
<td colspan="2">0.94</td>
<td colspan="2">0.92</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.95</td>
<td colspan="2">0.91</td>
<td colspan="2">0.88</td>
<td colspan="2">0.92</td>
<td colspan="2">0.88</td>
<td colspan="2">0.86</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.95</td>
<td colspan="2">0.92</td>
<td colspan="2">0.88</td>
<td colspan="2">0.93</td>
<td colspan="2">0.88</td>
<td colspan="2">0.84</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.97</bold></td>
<td colspan="2"><bold>0.95</bold></td>
<td colspan="2"><bold>0.92</bold></td>
<td colspan="2"><bold>0.95</bold></td>
<td colspan="2"><bold>0.92</bold></td>
<td colspan="2"><bold>0.90</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.02</bold></td>
<td colspan="2"><bold>0.05</bold></td>
<td colspan="2"><bold>0.01</bold></td>
<td colspan="2"><bold>0.04</bold></td>
<td colspan="2"><bold>0.07</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MCAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.99</td>
<td colspan="2">0.97</td>
<td colspan="2">0.97</td>
<td colspan="2">0.99</td>
<td colspan="2">0.98</td>
<td colspan="2">0.96</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.98</td>
<td colspan="2">0.96</td>
<td colspan="2">0.96</td>
<td colspan="2">0.98</td>
<td colspan="2">0.96</td>
<td colspan="2">0.93</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.96</td>
<td colspan="2">0.95</td>
<td colspan="2">0.92</td>
<td colspan="2">0.95</td>
<td colspan="2">0.92</td>
<td colspan="2">0.86</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.96</td>
<td colspan="2">0.92</td>
<td colspan="2">0.90</td>
<td colspan="2">0.95</td>
<td colspan="2">0.93</td>
<td colspan="2">0.85</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.96</td>
<td colspan="2">0.92</td>
<td colspan="2">0.92</td>
<td colspan="2">0.92</td>
<td colspan="2">0.92</td>
<td colspan="2">0.85</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.97</bold></td>
<td colspan="2"><bold>0.94</bold></td>
<td colspan="2"><bold>0.93</bold></td>
<td colspan="2"><bold>0.96</bold></td>
<td colspan="2"><bold>0.94</bold></td>
<td colspan="2"><bold>0.89</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.03</bold></td>
<td colspan="2"><bold>0.04</bold></td>
<td colspan="2"><bold>0.01</bold></td>
<td colspan="2"><bold>0.03</bold></td>
<td colspan="2"><bold>0.09</bold></td>
</tr>
<tr>
<td rowspan="7"><bold>MNAR</bold></td>
<td><bold>10</bold></td>
<td colspan="2">0.98</td>
<td colspan="2">0.97</td>
<td colspan="2">0.96</td>
<td colspan="2">0.96</td>
<td colspan="2">0.95</td>
<td colspan="2">0.96</td>
</tr>
<tr>
<td><bold>20</bold></td>
<td colspan="2">0.95</td>
<td colspan="2">0.94</td>
<td colspan="2">0.93</td>
<td colspan="2">0.94</td>
<td colspan="2">0.93</td>
<td colspan="2">0.94</td>
</tr>
<tr>
<td><bold>30</bold></td>
<td colspan="2">0.95</td>
<td colspan="2">0.93</td>
<td colspan="2">0.91</td>
<td colspan="2">0.93</td>
<td colspan="2">0.88</td>
<td colspan="2">0.87</td>
</tr>
<tr>
<td><bold>40</bold></td>
<td colspan="2">0.94</td>
<td colspan="2">0.91</td>
<td colspan="2">0.90</td>
<td colspan="2">0.91</td>
<td colspan="2">0.87</td>
<td colspan="2">0.88</td>
</tr>
<tr>
<td><bold>50</bold></td>
<td colspan="2">0.93</td>
<td colspan="2">0.89</td>
<td colspan="2">0.87</td>
<td colspan="2">0.90</td>
<td colspan="2">0.85</td>
<td colspan="2">0.82</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td colspan="2"><bold>0.95</bold></td>
<td colspan="2"><bold>0.93</bold></td>
<td colspan="2"><bold>0.91</bold></td>
<td colspan="2"><bold>0.93</bold></td>
<td colspan="2"><bold>0.90</bold></td>
<td colspan="2"><bold>0.90</bold></td>
</tr>
<tr>
<td></td>
<td colspan="2"><bold>Improvement</bold></td>
<td colspan="2"><bold>0.02</bold></td>
<td colspan="2"><bold>0.04</bold></td>
<td colspan="2"><bold>0.02</bold></td>
<td colspan="2"><bold>0.06</bold></td>
<td colspan="2"><bold>0.06</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>MVs are considered a critical problem in pattern recognition, ML and data mining applications. Many extensive studies have been performed for manipulating the problem of MVs especially in medical data. In addition to MVs, FS is the data preprocessing strategy which has been considered to be efficient when preparing data (specifically large volume data) in ML. It has been confirmed to be efficient and effective in handling high-dimensional data for any data dependent tool. Reducing the computational cost of modeling and the number of input features helps in improving the performance of the model.</p>
<p>In this paper, novel algorithm was proposed to manipulate M<italic>Vs</italic>. The proposed algorithm depends on FS of similarity classifier with Parkash&#x2019;s fuzzy entropy measure to select the candidate feature and the BRR model to predict MVs in the selected feature. Hence, the proposed algorithm consists mainly of two phases. In the first phase, the FS of similarity classifier with Parkash&#x2019;s fuzzy entropy is used to select features to be imputed one after one. In the second phase, the MVs in the selected feature are predicted using the BRR model. The first and second phases are repeated until the imputation of the whole dataset. The proposed algorithm is easy to implement and can deal with all MVs from any missingness mechanism. Furthermore, the proposed algorithm exhibits a good performance against the compared algorithms.</p>
<p>In future research, the proposed algorithm will be implemented on new medical datasets like pulmonary embolism data and cardiovascular disease. Furthermore, additional performance metrics will be taken into consideration such as the normalized root mean square error (NRMSE), statistical tests and predictive accuracy (PAC).</p>
</sec>
</body>
<back>
<ack>
<p>This project was funded by the Deanship of Scientific Research (DSR) at King Abdulaziz University (KAU), Jeddah, Saudi Arabia, under grant No. (PH: 13-130-1442). The authors, therefore acknowledge with thanks DSR for technical and financial support.</p>
</ack><fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> This project was funded by the Deanship of Scientific Research (DSR) at King Abdulaziz University (KAU), Jeddah, Saudi Arabia, under grant No. (PH: 13-130-1442).</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. S.</given-names> <surname>Osman</surname></string-name>, <string-name><given-names>A. M.</given-names> <surname>Abu-Mahfouz</surname></string-name> and <string-name><given-names>P. R.</given-names> <surname>Page</surname></string-name></person-group>, &#x201C;<article-title>A Survey on data imputation techniques: Water distribution system as a use case</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>6</volume>, pp. <fpage>63279</fpage>&#x2013;<lpage>63291</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Bouras</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>Hybrid missing value imputation algorithms using fuzzy c-means and vaguely quantified rough set</chapter-title>,&#x201D; in <source>IEEE Transactions on Fuzzy Systems</source>. <publisher-name>Early Acce</publisher-name>, <year>2021</year>. DOI <pub-id pub-id-type="doi">10.1109/TFUZZ.2021.3058643</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Mostafa</surname></string-name>, <string-name><given-names>A. S.</given-names> <surname>Eladimy</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Hamad</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Amano</surname></string-name></person-group>, &#x201C;<article-title>CBRG: A novel algorithm for handling missing data using bayesian ridge regression and feature selection based on gain ratio</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>216969</fpage>&#x2013;<lpage>216985</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>An efficient ensemble method for missing value imputation in microarray gene expression data</article-title>,&#x201D; <source>BMC Bioinformatics</source>, vol. <volume>22</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>25</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. I.</given-names> <surname>Lewin</surname></string-name></person-group>, &#x201C;<article-title>Getting clinical about neural networks</article-title>,&#x201D; <source>IEEE Intelligent Systems and their Applications</source>, vol. <volume>15</volume>, no. <issue>1</issue>, pp. <fpage>2</fpage>&#x2013;<lpage>5</lpage>, <year>2000</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. N.</given-names> <surname>Baraldi</surname></string-name> and <string-name><given-names>C. K.</given-names> <surname>Enders</surname></string-name></person-group>, &#x201C;<article-title>An introduction to modern missing data analyses</article-title>,&#x201D; <source>Journal of School Psychology</source>, vol. <volume>48</volume>, no. <issue>1</issue>, pp. <fpage>5</fpage>&#x2013;<lpage>37</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Doquire</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Verleysen</surname></string-name></person-group>, &#x201C;<article-title>Feature selection with missing data using mutual information estimators</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>90</volume>, no. <issue>2</issue>, pp. <fpage>3</fpage>&#x2013;<lpage>11</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Mostafa</surname></string-name></person-group>, &#x201C;<article-title>Missing data imputation by the aid of features similarities</article-title>,&#x201D; <source>Int. Journal of Big Data Management</source>, vol. <volume>1</volume>, no. <issue>1</issue>, pp. <fpage>81</fpage>&#x2013;<lpage>103</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Mostafa</surname></string-name></person-group>, &#x201C;<article-title>Imputing missing values using cumulative linear regression</article-title>,&#x201D; <source>CAAI Transactions on Intelligence Technology</source>, vol. <volume>4</volume>, no. <issue>3</issue>, pp. <fpage>182</fpage>&#x2013;<lpage>200</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. L.</given-names> <surname>Yadav</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Roychoudhury</surname></string-name></person-group>, &#x201C;<article-title>Handling missing values: A study of popular imputation packages in R</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>160</volume>, no. <issue>9</issue>, pp. <fpage>104</fpage>&#x2013;<lpage>118</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. C.</given-names> <surname>Acock</surname></string-name></person-group>, &#x201C;<article-title>Working with missing values</article-title>,&#x201D; <source>Journal of Marriage and Family</source>, vol. <volume>67</volume>, no. <issue>4</issue>, pp. <fpage>1012</fpage>&#x2013;<lpage>1028</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. B.</given-names> <surname>Albayati</surname></string-name> and <string-name><given-names>A. M.</given-names> <surname>Altamimi</surname></string-name></person-group>, &#x201C;<article-title>An empirical study for detecting fake facebook profiles using supervised mining techniques</article-title>,&#x201D; <source>Informatica</source>, vol. <volume>43</volume>, no. <issue>1</issue>, pp. <fpage>77</fpage>&#x2013;<lpage>86</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Madley-Dowd</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Hughes</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Tilling</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Heron</surname></string-name></person-group>, &#x201C;<article-title>The proportion of missing data should not be used to guide decisions on multiple imputation</article-title>,&#x201D; <source>Journal of Clinical Epidemiology</source>, vol. <volume>110</volume>, no. <issue>1</issue>, pp. <fpage>63</fpage>&#x2013;<lpage>73</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Mostafa</surname></string-name>, <string-name><given-names>A. S.</given-names> <surname>Eladimy</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Hamad</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Amano</surname></string-name></person-group>, &#x201C;<article-title>CBRL and CBRC: Novel algorithms for improving missing value imputation accuracy based on bayesian ridge regression</article-title>,&#x201D; <source>Symmetry (Basel)</source>, vol. <volume>12</volume>, no. <issue>10</issue>, pp. <fpage>1594</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Varoquaux</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Buitinck</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Louppe</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Grisel</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Pedregosa</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Scikit-learn: Machine learning without learning the machinery</article-title>,&#x201D; <source>GetMobile: Mobile Computing and Communications</source>, vol. <volume>19</volume>, no. <issue>1</issue>, pp. <fpage>29</fpage>&#x2013;<lpage>33</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P. L.</given-names> <surname>Roth</surname></string-name></person-group>, &#x201C;<article-title>Missing data: A conceptual review for applied psychologists</article-title>,&#x201D; <source>Personnel Psychology</source>, vol. <volume>47</volume>, no. <issue>3</issue>, pp. <fpage>537</fpage>&#x2013;<lpage>560</lpage>, <year>1994</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P. J.</given-names> <surname>Garc&#x00ED;a-Laencina</surname></string-name>, <string-name><given-names>J. L.</given-names> <surname>Sancho-G&#x00F3;mez</surname></string-name> and <string-name><given-names>A. R.</given-names> <surname>Figueiras-Vidal</surname></string-name></person-group>, &#x201C;<article-title>Pattern classification with missing data: A review</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>19</volume>, no. <issue>2</issue>, pp. <fpage>263</fpage>&#x2013;<lpage>282</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. M.</given-names> <surname>Hamer</surname></string-name> and <string-name><given-names>P. M.</given-names> <surname>Simpson</surname></string-name></person-group>, &#x201C;<article-title>Last observation carried forward versus mixed models in the analysis of psychiatric clinical trials</article-title>,&#x201D; <source>American Journal of Psychiatry</source>, vol. <volume>166</volume>, no. <issue>6</issue>, pp. <fpage>639</fpage>&#x2013;<lpage>641</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>R. J.</given-names> <surname>Little</surname></string-name> and <string-name><given-names>D. B.</given-names> <surname>Rubin</surname></string-name></person-group>, &#x201C;<chapter-title>Maximum likelihood for general patterns of missing data: Introduction and theory with ignorable nonresponse</chapter-title>,&#x201D; in <source>Statistical Analysis with Missing Data</source>. <publisher-name>Second Edition, John Wiley &#x0026; Sons, Inc.</publisher-name>, <publisher-loc>Hoboken, New Jersey, USA</publisher-loc>, pp. <fpage>164</fpage>&#x2013;<lpage>189</lpage>, <year>2002</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. M.</given-names> <surname>Lang</surname></string-name> and <string-name><given-names>T. D.</given-names> <surname>Little</surname></string-name></person-group>, &#x201C;<article-title>Principled missing data treatments</article-title>,&#x201D; <source>Prevention Science</source>, vol. <volume>19</volume>, no. <issue>3</issue>, pp. <fpage>284</fpage>&#x2013;<lpage>294</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Kang</surname></string-name></person-group>, &#x201C;<article-title>The prevention and handling of the missing data</article-title>,&#x201D; <source>Korean Journal of Anesthesiology</source>, vol. <volume>64</volume>, no. <issue>5</issue>, pp. <fpage>402</fpage>&#x2013;<lpage>406</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. P.</given-names> <surname>Dempster</surname></string-name>, <string-name><given-names>N. M.</given-names> <surname>Laird</surname></string-name> and <string-name><given-names>D. B.</given-names> <surname>Rubin</surname></string-name></person-group>, &#x201C;<article-title>Maximum likelihood from incomplete data via the EM algorithm</article-title>,&#x201D; <source>Journal of the Royal Statistical Society: Series B (Methodological)</source>, vol. <volume>39</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>22</lpage>, <year>1977</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. Van</given-names> <surname>Hulse</surname></string-name> and <string-name><given-names>T. M.</given-names> <surname>Khoshgoftaar</surname></string-name></person-group>, &#x201C;<article-title>Incomplete-case nearest neighbor imputation in software measurement data</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>259</volume>, no. <issue>2</issue>, pp. <fpage>596</fpage>&#x2013;<lpage>610</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Williams</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xue</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Carin</surname></string-name></person-group>, &#x201C;<article-title>Incomplete-data classification using logistic regression</article-title>,&#x201D; in <conf-name>Proc. of the 22nd Int. Conf. on Machine Learning</conf-name>, <publisher-loc>Association for Computing Machinery, New York, NY, United States</publisher-loc>, pp. <fpage>972</fpage>&#x2013;<lpage>979</lpage>, <year>2005</year>. </mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Wagstaff</surname></string-name></person-group>, &#x201C;<chapter-title>Clustering with missing values: No imputation required</chapter-title>,&#x201D; in <source>Classification, Clustering, and Data Mining Applications</source>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>649</fpage>&#x2013;<lpage>658</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. L.</given-names> <surname>Silva-Ram&#x00ED;rez</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Pino-Mej&#x00ED;as</surname></string-name>, <string-name><given-names>M.</given-names> <surname>L&#x00F3;pez-Coello</surname></string-name> and <string-name><given-names>M. D.</given-names> <surname>Cubiles-de-la-Vega</surname></string-name></person-group>, &#x201C;<article-title>Missing value imputation on missing completely at random data using multilayer perceptrons</article-title>,&#x201D; <source>Neural Networks</source>, vol. <volume>24</volume>, no. <issue>1</issue>, pp. <fpage>121</fpage>&#x2013;<lpage>129</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Nikfalazar</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>Yeh</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bedingfield</surname></string-name> and <string-name><given-names>H. A.</given-names> <surname>Khorshidi</surname></string-name></person-group>, &#x201C;<article-title>Missing data imputation using decision trees and fuzzy clustering with iterative learning</article-title>,&#x201D; <source>Knowledge and Information Systems</source>, vol. <volume>62</volume>, no. <issue>6</issue>, pp. <fpage>2419</fpage>&#x2013;<lpage>2437</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Luukka</surname></string-name></person-group>, &#x201C;<article-title>Feature selection using fuzzy entropy measures with similarity classifier</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>38</volume>, no. <issue>4</issue>, pp. <fpage>4600</fpage>&#x2013;<lpage>4607</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Rianne</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Peter</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Jaap</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Gerko</surname></string-name></person-group>, &#x201C;<article-title>Generate missing values with ampute</article-title>,&#x201D; <year>2017</year>, [Online]. Available: <uri>https://rianneschouten.github.io/mice_ampute/vignette/ampute.html</uri>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>M. D. Nilsel</given-names> <surname>Ilter</surname></string-name> and <string-name><given-names>H. A.</given-names> <surname>Guvenir</surname></string-name></person-group>, &#x201C;<article-title>Dermatology</article-title>,&#x201D; <comment>[Online]. 2021. Available: </comment> <uri>https://archive.ics.uci.edu/ml/datasets/dermatology</uri>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Wi. H.</given-names> <surname>Wolberg</surname></string-name></person-group>, &#x201C;<article-title>Breast cancer wisconsin</article-title>,&#x201D; <comment>[Online]. 2021. Available: </comment> <uri>https://archive.ics.uci.edu/ml/datasets/breast&#x002B;cancer&#x002B;wisconsin&#x002B;(original)</uri>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Max</given-names> <surname>Little</surname></string-name></person-group>, &#x201C;<article-title>Parkinsons data set</article-title>,&#x201D; <comment>[Online]. 2021. Available: </comment> <uri>https://archive.ics.uci.edu/ml/datasets/parkinsons</uri>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>R. A.</given-names> <surname>Rossi</surname></string-name> and <string-name><given-names>Nesreen K.</given-names> <surname>Ahmed</surname></string-name></person-group>, &#x201C;<article-title>Pima Indians Diabetes</article-title>,&#x201D; <comment>[Online]. 2021. Available: </comment> <uri>http://networkrepository.com/pima-indians-diabetes.php</uri>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Kearney</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Barkat</surname></string-name></person-group>, &#x201C;<article-title>Autoimpute</article-title>,&#x201D; <comment>[Online]. 2021. Available: </comment> <uri>https://autoimpute.readthedocs.io/en/latest/</uri>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Law</surname></string-name></person-group>, &#x201C;<article-title>Impyute</article-title>,&#x201D; <comment>[Online]. 2021. Available:</comment> <uri>https://impyute.readthedocs.io/en/latest/</uri>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. J.</given-names> <surname>Willmott</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Matsuura</surname></string-name></person-group>, &#x201C;<article-title>Advantages of the mean absolute error (MAE) over the root mean square error (RMSE) in assessing average model performance</article-title>,&#x201D; <source>Climate Research</source>, vol. <volume>30</volume>, no. <issue>1</issue>, pp. <fpage>79</fpage>&#x2013;<lpage>82</lpage>, <year>2005</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>